A Method, System, Device and Medium for Dynamic Scheduling of Visual Recognition Algorithm Containers

Through real-time monitoring and dynamic scheduling of visual recognition algorithm containers, the limitations of existing systems in resource allocation, task queue management and performance optimization are solved, and efficient resource utilization and task processing are achieved.

CN120107759BActive Publication Date: 2025-07-01QIANXUN TECH (SHENZHEN) CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510586443.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-07-01
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

Existing visual identification systems have limitations in resource allocation strategies, task queue management, real-time monitoring and performance optimization, resulting in low resource utilization, slow task processing speed and low recognition efficiency.

Method used

A dynamic scheduling method for containers of visual recognition algorithm is proposed. By monitoring the running status and performance indicators of the algorithm container in real time, evaluating the startup ranking of the algorithm to be started, dynamically adjusting resource allocation, optimizing task queue management, and dynamically adjusting algorithm parameters based on real-time data.

Benefits of technology

It realizes efficient management and optimized execution of algorithm containers, improves the processing speed of visual recognition tasks and the overall performance of the system, and improves resource utilization and recognition efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107759B_ABST
    Figure CN120107759B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system, device and medium for dynamic scheduling of vision recognition algorithm containers. The method includes: obtaining a collection of algorithms to be started; obtaining a start ranking list of the algorithms to be started; sequentially taking out the algorithms to be started from the start ranking list in the order of start ranking, and determining whether the corresponding recognition tasks to be recognized can be completed within a preset time; if so, it means that only one algorithm container needs to be started; if not, it means that multiple algorithm containers need to be started, and calculate the number of algorithm containers to be started; based on the number of algorithm containers to be started, determine whether the GPU resources are sufficient; determine whether all the started algorithms are among the top few in the start ranking list; The present invention realizes the intelligent scheduling and dynamic scaling of algorithm containers by dynamically adjusting the resource allocation of algorithm containers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method, system, device and medium for dynamic scheduling of visual recognition algorithm containers, and belongs to the technical field of cloud computing resource management. Background Art

[0002] Currently, visual recognition technology has become an indispensable part of the modern technology ecosystem. Through functions such as image recognition, object detection, and scene understanding, it provides users with an unprecedented interactive experience and decision-making support. Moreover, driven by artificial intelligence technology, machine learning and deep learning algorithms have made remarkable progress in the field of visual recognition, enabling computer vision systems to process complex image and video data and be applied to multiple industries such as autonomous driving, medical diagnosis, and security monitoring.

[0003] To cope with the rapidly changing market demands, algorithm containerization technology has emerged. It allows algorithms to be quickly deployed in different computing environments in the form of containers, improving the portability and scalability of algorithms. Although containerization technology brings convenience, it also introduces new resource management problems, which are specifically introduced as follows.

[0004] 1) Limitations of resource allocation strategies: Existing visual recognition systems usually adopt polling or static resource allocation strategies. These strategies lack flexibility in resource allocation and cannot dynamically adjust resources according to the actual needs of tasks. As a result, resources are often insufficient during peak task loads and cannot meet the processing requirements of high-concurrency tasks, while a large amount of resources are wasted during low task loads, reducing the overall efficiency and economic benefits of the system.

[0005] 2) Limitations of task queue management and scheduling decisions: Existing visual recognition systems lack effective mechanisms to manage and optimize task queues and often cannot achieve priority scheduling of tasks. As a result, they cannot respond quickly under high load, while resources are wasted under low load, making it difficult for the system to make reasonable scheduling decisions based on the urgency of tasks and resource requirements. This not only affects the processing speed of visual recognition tasks but also limits the dynamic scaling ability of algorithm containers in a multi-task environment and cannot achieve optimal allocation and utilization of resources.

[0006] 3) Lack of real-time monitoring: When existing visual recognition systems execute visual recognition tasks, they often neglect the real-time monitoring of the performance of algorithm containers. This neglect limits the refined management of resources by the system, resulting in low recognition efficiency, inability to fully utilize the computing power of algorithm containers, and difficulty in ensuring the accuracy and reliability of recognition results.

[0007] 4) Lack of performance optimization: Existing visual recognition systems lack effective performance monitoring and optimization mechanisms, making it difficult to dynamically adjust algorithm parameters according to real-time data during the execution of visual recognition tasks. This is not conducive to improving recognition efficiency and accuracy. Especially in application scenarios that require high efficiency and high precision, this shortcoming is particularly prominent, seriously affecting the overall performance of the system and the user experience.

[0008] As can be seen from the above, how to effectively schedule and manage algorithm containers with limited resources has become an urgent problem to be solved in the field of cloud computing. Therefore, the present invention urgently needs to develop a dynamic, efficient, and intelligent method and system for scheduling and executing visual recognition algorithm containers. Summary of the Invention

[0009] In view of the above existing technical problems, the present invention provides a method, system, device, and medium for dynamically scheduling visual recognition algorithm containers to solve the problems of the efficiency and resource utilization rate of algorithm container scheduling and execution tasks in existing visual recognition systems, and achieve the technical objectives of realizing the efficient management and optimized execution of algorithm containers, and improving the processing speed of visual recognition tasks and the overall performance of the system.

[0010] To achieve the above technical objectives, first, the present invention provides a method for dynamically scheduling visual recognition algorithm containers, including an algorithm container scheduling stage; the algorithm container scheduling stage includes the following steps:

[0011] Based on the recognition task table, obtain the algorithm set corresponding to the current task to be recognized; based on the algorithm list, obtain the algorithm set of currently started algorithms; take the union of the algorithm set corresponding to the task to be recognized and the algorithm set of started algorithms to obtain the algorithm set to be started.

[0012] Real-time monitor the running status and performance metrics of algorithm containers, and evaluate the start ranking of each algorithm to be started in the algorithm set to be started according to the video memory occupancy, inference speed, actual throughput, unrecognized task volume, and task waiting duration of the algorithm containers, to obtain the start ranking table of algorithms to be started.

[0013] In the order of the start ranking, take out the algorithms to be started from the start ranking table one by one, and judge whether the corresponding tasks to be recognized can be completed within the preset time; if so, it means that only one algorithm container needs to be started; if not, it means that multiple algorithm containers need to be started, and calculate the number of algorithm containers to be started.

[0014] Based on the number of algorithm containers to be started, judge whether the GPU resources are sufficient; if sufficient, start one algorithm container; if not, judge whether there is a started algorithm with a lower ranking than the current algorithm to be started. If so, close the algorithm container of the started algorithm, and repeat the judgment of whether the GPU resources are sufficient. If not, exit the current algorithm container scheduling.

[0015] Determine whether all the started algorithms are among the top several in the start ranking list; if so, end the current algorithm container scheduling, and if not, return to sequentially retrieve the algorithms to be started from the start ranking list in the order of start ranking until the start ranking list is traversed.

[0016] Specifically, for the method of the present invention, the start rankings of each algorithm to be started in the set of algorithms to be started are evaluated to obtain a start ranking list of the algorithms to be started, including the following sub-steps:

[0017] Obtain the video memory occupancy ratio p1, inference speed ratio p2, actual throughput ratio p3, unrecognized task volume ratio p4, and task waiting duration ratio p5 of each algorithm to be started;

[0018] Among them, the video memory occupancy ratio p1, the calculation formula is: p1 = algorithm video memory occupancy / sum of total algorithm video memory occupancies.

[0019] The inference speed ratio p2, the calculation formula is: p2 = algorithm inference speed / sum of total algorithm inference speeds.

[0020] The actual throughput ratio p3, the calculation formula is: p3 = algorithm actual throughput / total algorithm throughput.

[0021] The ratio p4 of the unrecognized task volume, the calculation formula is: p4 = number of unrecognized tasks of the algorithm / total number of unrecognized tasks.

[0022] The task waiting duration ratio p5, the calculation formula is: p5 = algorithm task waiting duration / sum of total algorithm task waiting durations.

[0023] According to the preset weights a, b, c, d, e, calculate the scores of each algorithm to be started, the calculation formula is: score = p2 × a + p3 × b + p4 × c + p5 × d - p1 × e.

[0024] Arrange all the algorithms to be started in descending order of scores to obtain a start ranking list of the algorithms to be started.

[0025] Specifically, for the method of the present invention, the determination of whether the corresponding unrecognized task can be completed within the preset time includes the following sub-steps:

[0026] Calculate the number of tasks that the algorithm to be started can complete, the calculation formula is: number of tasks that can be completed = algorithm throughput capacity × preset time.

[0027] Determine whether the number of tasks that can be completed is greater than the total number of unrecognized tasks;

[0028] If so, it means that it can be completed within the preset time; if not, it means that it cannot be completed within the preset time.

[0029] Specifically, for the method of the present invention, the formula for calculating the number of algorithm containers to be started is: the number of algorithm containers = the number of tasks to be recognized currently ÷ the number of tasks that can be completed within the preset time.

[0030] Further, for the method of the present invention, before the algorithm container scheduling stage, there is also an identification request storage stage; the identification request storage stage includes the following steps:

[0031] The business system service invokes the algorithm identification request and sends the algorithm identification request including the data to be recognized, request parameters, and unique identification identifier to the intelligent identification system service.

[0032] The intelligent identification system service performs parameter verification on the algorithm identification request; if the verification is successful, the algorithm identification request is recorded in the database, its status is marked as to be recognized, and at the same time, a postgres table record including the unique identification identifier and the algorithm primary key identifier is generated; if the verification fails, the algorithm identification request is recorded in the database, and its status is marked as parameter verification failed.

[0033] All algorithm identification requests with the status marked as to be recognized are stored in the identification task table to form a to-be-recognized task queue.

[0034] Further, for the method of the present invention, after the algorithm container scheduling stage, there is also an identification task execution stage; the identification task execution stage includes the following steps:

[0035] The currently started algorithm container obtains the corresponding to-be-recognized task from the identification task table, marks its status as being recognized, and performs identification processing on it.

[0036] When all tasks with the status marked as being recognized are recognized, the algorithm container enters the sleep state.

[0037] Even further, for the method of the present invention, the identification task execution stage further includes the following steps:

[0038] Once an algorithm identification request with the status marked as to be recognized is stored in the identification task table, an algorithm new task notification signal is triggered to wake up the corresponding algorithm container in the sleep state.

[0039] The awakened algorithm container fetches tasks and determines whether there is a corresponding to-be-recognized task in the task list;

[0040] If not, it determines whether no task has been fetched after looping a preset number of times; if not, the loop count + 1, and it returns to the awakened algorithm container to fetch tasks; if so, an algorithm sleep notification signal is triggered to make the awakened algorithm container enter the sleep state;

[0041] If there is, after the awakened algorithm container obtains the task to be recognized, it marks its status as being recognized, performs recognition processing on it, obtains the recognition result and determines whether it is normal; if it is normal, it indicates that the callback recognition is successful; if it is abnormal, it indicates that the callback recognition fails; then it determines whether the retry count of the callback recognition failure is greater than the preset count. If so, it modifies the recorded result to callback recognition failure; if not, it modifies the recorded result to callback recognition success.

[0042] Second, the present invention also provides a dynamic scheduling system for visual recognition algorithm containers, including an algorithm container scheduling module; the algorithm container scheduling module includes an algorithm collection obtaining unit, an evaluation startup ranking unit, a container quantity calculating unit, a GPU resource releasing unit, and a startup ranking checking unit;

[0043] The algorithm collection obtaining unit is used to: based on the recognition task table, obtain the algorithm collection corresponding to the current task to be recognized; based on the algorithm list, obtain the currently started algorithm collection; take the union of the algorithm collection corresponding to the task to be recognized and the currently started algorithm collection to obtain the algorithm collection to be started;

[0044] The evaluation startup ranking unit is used to: monitor the running status and performance metrics of the algorithm containers in real time, and evaluate the startup ranking of each to-be-started algorithm in the to-be-started algorithm collection according to the video memory occupancy, inference speed, actual throughput, unrecognized task quantity, and task waiting duration of the algorithm containers, to obtain the startup ranking table of the to-be-started algorithms;

[0045] The container quantity calculating unit is used to: sequentially take out the to-be-started algorithms from the startup ranking table in the order of the startup ranking, and determine whether the corresponding tasks to be recognized can be completed within the preset time; if so, it means that only one algorithm container needs to be started; if not, it means that multiple algorithm containers need to be started, and calculate the number of algorithm containers to be started;

[0046] The GPU resource releasing unit is used to: based on the number of algorithm containers to be started, determine whether the GPU resources are sufficient; if sufficient, start one algorithm container; if not, determine whether there is a started algorithm with a lower ranking than the currently to-be-started algorithm. If so, close the algorithm container of the started algorithm, and repeatedly determine whether the GPU resources are sufficient. If not, exit the current algorithm container scheduling;

[0047] The startup ranking checking unit is used to: determine whether all the started algorithms are among the top several in the startup ranking table; if so, end the current algorithm container scheduling. If not, return to sequentially take out the to-be-started algorithms from the startup ranking table in the order of the startup ranking until the startup ranking table is traversed.

[0048] Thirdly, the present invention further provides an electronic device, including a processor and a memory coupled to the processor; the memory has computer program instructions stored therein; when the computer program instructions are executed by the processor, the electronic device executes the above-mentioned method for dynamically scheduling a visual recognition algorithm container.

[0049] Fourthly, the present invention further provides a computer-readable storage medium, including computer program instructions stored therein; when the computer program instructions are executed, the above-mentioned method for dynamically scheduling a visual recognition algorithm container is implemented.

[0050] In summary, the present invention provides a dynamic, efficient, and intelligent method for scheduling and executing a visual recognition algorithm container. This method can dynamically adjust the resource allocation of the algorithm container according to the actual requirements of the visual recognition task, optimize the task queue management, and achieve the intelligent scheduling and dynamic scaling of the algorithm container. At the same time, this method can also monitor the algorithm performance in real time, dynamically adjust the algorithm parameters according to the monitoring results, so as to improve the processing speed and accuracy of the recognition task, thereby enhancing the overall performance of the system and user satisfaction.

[0051] Moreover, the present invention is particularly suitable for the dynamic scheduling and execution of artificial intelligence visual recognition algorithms. Through an intelligent scheduling system, the efficient management and optimized execution of algorithm containers are realized, the processing speed of visual recognition tasks and the overall performance of the system are improved. Especially in terms of resource allocation, task scheduling, and performance monitoring, an innovative system is provided, and the specific technical advantages are as follows:

[0052] 1. Real-time monitoring and dynamic scheduling: The system can monitor the performance indicators of algorithm containers in real time and dynamically adjust the resource allocation according to the monitoring data. This real-time monitoring and dynamic adjustment mechanism enables the system to quickly respond to changes in resource requirements and improve resource utilization.

[0053] 2. Intelligent management of task queues: Manage the task queue through intelligent algorithms, optimize the task processing order and resource allocation. This intelligent management mechanism ensures that high-priority tasks can be processed preferentially and improves the efficiency of task processing.

[0054] 3. Multi-dimensional scheduling decision-making: Consider multiple dimensions such as task urgency, resource utilization, and algorithm performance comprehensively to achieve optimal scheduling decisions. This multi-dimensional consideration makes the scheduling decisions more comprehensive and accurate, and improves the scheduling performance of the system.

[0055] 4. Real-time monitoring of algorithm performance: During the execution of recognition tasks, monitor the algorithm performance in real time, dynamically adjust the algorithm parameters, and improve the recognition efficiency. This real-time monitoring and dynamic adjustment mechanism ensures that the algorithm can run in the best state and improves the accuracy and efficiency of recognition tasks. Description of the Drawings

[0056] Figure 1 This is the flowchart of the recognition request storage phase when implementing the method of the present invention;

[0057] Figure 2 This is the flowchart of the algorithm container scheduling phase when implementing the method of the present invention;

[0058] Figure 3 This is the flowchart of the recognition task execution phase when implementing the method of the present invention;

[0059] Figure 4 This is the algorithm list interface diagram of S2-2 in the algorithm container scheduling phase when implementing the method of the present invention;

[0060] Figure 5 This is the sub-step flowchart of S2-4 in the algorithm container scheduling phase when implementing the method of the present invention;

[0061] Figure 6 This is the sub-step flowchart of S2-5 in the algorithm container scheduling phase when implementing the method of the present invention;

[0062] Figure 7 This is the sub-step flowchart of S3-3 in the recognition task execution phase when implementing the method of the present invention;

[0063] Figure 8 This is the principle block diagram when implementing the system of the present invention;

[0064] Figure 9 This is the workflow diagram of using different namespaces and containers in Kubernetes (K8S) when implementing the system of the present invention;

[0065] Figure 10 This is the principle block diagram when implementing the device of the present invention. Detailed implementation manners

[0066] Hereinafter, embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0067] In the description of the embodiments of the present disclosure, the term "including" and its similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects, unless otherwise specified. There may also be other explicit and implicit definitions hereinafter.

[0068] Embodiment 1: The dynamic scheduling method for the visual recognition algorithm container of the present invention.

[0069] This embodiment provides a dynamic scheduling method for a visual recognition algorithm container, including the recognition request storage stage, the algorithm container scheduling stage, and the recognition task execution stage. As Figure 1 、 Figure 2 、 Figure 3 shown, the logic of resource allocation and scheduling decisions in each stage is described in detail, showing how to dynamically adjust the resource allocation and the number of replicas of the algorithm container according to the real-time monitoring data and the status of the task queue, which is specifically introduced as follows.

[0070] S1. Recognition request storage stage: As Figure 1 shown, the recognition request storage stage includes initiating an algorithm recognition request, parameter verification, and storing it in the recognition task table. The specific steps are as follows.

[0071] S1-1. The business system service (including but not limited to the client or the upstream system that needs to call the algorithm recognition) initiates an algorithm recognition request and sends the algorithm recognition request to the intelligent recognition system service. The algorithm recognition request contains all the information required for algorithm recognition, including the data to be recognized (such as images, videos, etc.), the requested parameters (such as the type of recognition algorithm, the accuracy requirement for recognition, etc.), and the unique recognition identifier used to uniquely identify the request.

[0072] S1-2. After receiving the algorithm recognition request, the intelligent recognition system service first performs parameter verification on it. The verification includes authentication information verification and necessary parameter verification, and judges whether the verification is successful.

[0073] If the verification fails, it returns failure, records the algorithm recognition request in the database, and marks its status as parameter verification failed.

[0074] If the verification is successful, the algorithm recognition request is also recorded in the database, and its status is marked as to be recognized. At the same time, a postgres table record is generated in the database.

[0075] In specific implementation, the postgres table record includes the unique identification mark brought by the business system service and the algorithm primary key identification. It should be noted that each time the business system service calls the algorithm recognition request, it needs to bring the unique identification mark as the unique identification of the current algorithm recognition request. Moreover, since there are multiple recognition algorithms (such as face recognition algorithm, meter recognition algorithm, etc.), the algorithm primary key identification is needed to represent information such as which algorithm is used for recognition.

[0076] S1-3. Store all algorithm recognition requests with the status marked as to-be-recognized into the recognition task table to form a to-be-recognized task queue. Each to-be-recognized task in this to-be-recognized task queue corresponds to a to-be-recognized algorithm recognition request and a to-be-started algorithm (specified by the algorithm primary key identification).

[0077] In specific implementation, the recognition task table is a postgres database table used to store the algorithm recognition request information to be recognized. When the status of an algorithm recognition request is marked as "to-be-recognized" and stored into the recognition task table, it becomes a to-be-recognized task, waiting for the system to allocate an algorithm and perform the recognition work. This step centrally manages all algorithm recognition requests with the status of "to-be-recognized" to form a to-be-recognized task queue. The system will take tasks from this to-be-recognized task queue according to a certain strategy (such as first-in-first-out, priority, etc.) and execute the corresponding algorithm to complete the recognition work.

[0078] S2. Algorithm container scheduling stage: It is mainly divided into an algorithm ranking sub-stage and an algorithm start / stop sub-stage, and the scheduling is automatically performed every 10 minutes, which is introduced as follows.

[0079] Firstly, the algorithm ranking sub-stage: Obtain the algorithms of all to-be-recognized tasks from the recognition task table, and obtain all the started algorithms from the algorithm list, so as to obtain all the to-be-started algorithms, and evaluate the start ranking of each to-be-started algorithm. The factors for evaluating the start ranking include the video memory occupancy, inference speed, actual throughput, unrecognized task volume, and task waiting duration of the algorithm container, etc., so as to obtain the start ranking table of the to-be-started algorithms.

[0080] Secondly, the algorithm start / stop sub-stage: The core of this stage is to start the algorithms ranked ahead as much as possible according to the start ranking table of the to-be-started algorithms under limited computing power resources until the resources are insufficient. Specifically, first start the to-be-started algorithm ranked at the top of the start ranking table, and calculate how many algorithm containers need to be started according to the throughput capacity of this to-be-started algorithm to recognize all the to-be-recognized tasks corresponding to this algorithm within the preset time. Then call the api of k8s to start the algorithm container, and so on, until the resources are insufficient, then end this stage.

[0081] As Figure 2 shown, the specific steps in the algorithm container scheduling stage are as follows.

[0082] S2-1. Based on the recognition task list, obtain the algorithm collection corresponding to the currently pending recognition task.

[0083] In specific implementation, traverse the recognition task list, and according to the algorithm primary key identifier in the algorithm recognition request, extract all unique algorithms to form the algorithm collection for the pending recognition task. This refers to the algorithm collection that can be used to process the pending recognition request in the intelligent recognition system service. These algorithms may include face recognition algorithms, meter recognition algorithms, etc. Each algorithm has its specific application scenario and recognition ability. In addition, after the status of the algorithm recognition request is marked as "pending recognition", the intelligent recognition system service will select the corresponding algorithm from the algorithm collection according to the algorithm primary key identifier specified in the algorithm recognition request to process the request.

[0084] S2-2. Based on the algorithm list, obtain the collection of currently started algorithms.

[0085] In specific implementation, as Figure 4 shown, multiple algorithms are maintained in the algorithm list, such as: face recognition algorithm, meter recognition algorithm, personnel detection algorithm, etc. Access the algorithm list and filter out the algorithms with the status of "started" to form the collection of started algorithms.

[0086] S2-3. Take the union of the above two algorithm collections to obtain the collection of algorithms to be started.

[0087] In specific implementation, since the algorithms of the pending recognition task and the started algorithms need to be scheduled, the algorithm collection of the pending recognition task and the collection of started algorithms are merged into a new collection, and duplicate items are removed to form the collection of algorithms to be started. This collection contains all the algorithms that appear in the algorithm collection of the pending recognition task or the collection of started algorithms, but each algorithm will only appear once (that is, duplicate elements are removed). In practical applications, this means that all the algorithms of the pending recognition task and all the started algorithms need to be scheduled, but each algorithm only needs to be scheduled once. This operation is usually used in scenarios such as resource allocation and task scheduling to ensure that all relevant algorithms can be properly processed.

[0088] S2-4. Real-time monitor the running status and performance metrics of the algorithm container, and based on the video memory occupancy, inference speed, actual throughput, unrecognized task volume, and task waiting duration of the algorithm container, evaluate the start ranking of each algorithm to be started in the collection of algorithms to be started, and obtain the start ranking table of the algorithms to be started.

[0089] During specific implementation, the running status and performance metrics of the algorithm containers are monitored in real time, such as video memory occupancy, inference speed, throughput, etc., so as to adjust the startup strategy and resource configuration of the algorithm containers in a timely manner. Then, based on the algorithm collection to be started, the startup priority ranking of each algorithm to be started is evaluated. The ranking basis includes but is not limited to the proportion of video card storage occupancy of the algorithm container, the proportion of inference speed, the proportion of actual throughput, the proportion of unrecognized task volume, and the proportion of task waiting duration. Next, by comprehensively evaluating these factors, a startup priority ranking list of the algorithms to be started is obtained.

[0090] As Figure 5 shown, the steps for evaluating the startup ranking of each algorithm to be started in the algorithm collection to be started and obtaining the startup ranking list of the algorithms to be started include the following sub-steps:

[0091] Obtain the proportion p1 of the video memory occupancy of each algorithm to be started. The calculation formula is: p1 = algorithm video memory occupancy / total algorithm video memory occupancy sum. Where p1 represents the ratio of the video memory currently occupied by a certain algorithm to be started to the total video memory occupancy of all algorithms. Algorithm video memory occupancy represents the video memory amount of this algorithm to be started, and the unit can be MB, GB, etc. Total algorithm video memory occupancy sum represents the sum of the current video memory occupancies of all algorithms in the algorithm collection to be started.

[0092] Obtain the proportion p2 of the inference speed of each algorithm to be started. The calculation formula is: p2 = algorithm inference speed / total algorithm inference speed sum. Where p2 represents the ratio (or relative speed) of the inference speed of a certain algorithm to be started to the total inference speed of all algorithms. Algorithm inference speed represents the number of tasks or data volume that this algorithm to be started can process per second, and the unit can be tasks / s, images / s, etc. Total algorithm inference speed sum represents the sum of the inference speeds of all algorithms in the algorithm collection to be started.

[0093] Obtain the proportion p3 of the actual throughput of each algorithm to be started. The calculation formula is: p3 = algorithm actual throughput / total algorithm throughput. And, the calculation formula for the actual throughput of a single algorithm is as follows: for the algorithm to be started, its actual throughput is 0; for the started algorithm, its actual throughput = the amount of tasks done during the startup period / startup duration. Where p3 represents the ratio of the number of tasks actually completed by a certain algorithm to be started within a period of time to the total number of tasks actually completed by all algorithms. Algorithm actual throughput represents the number of tasks completed by this algorithm to be started within the specified time window. Total algorithm throughput represents the sum of the number of tasks completed by all algorithms in the algorithm collection to be started within the same time window.

[0094] Obtain the proportion p4 of the unrecognized task volume of each algorithm to be started. The calculation formula is: p4 = the number of tasks to be recognized by the algorithm / the total number of tasks to be recognized by all algorithms. Among them, p4 represents the ratio of the number of tasks currently to be recognized by a certain algorithm to be started to the total number of tasks to be recognized by all algorithms. The number of unrecognized tasks of the algorithm represents the number of tasks currently queuing for recognition of this algorithm to be started. The total number of tasks to be recognized by all algorithms represents the total number of tasks currently queuing for recognition of all algorithms in the set of algorithms to be started.

[0095] Obtain the proportion p5 of the task waiting duration of each algorithm to be started. The calculation formula is: p5 = the task waiting duration of the algorithm / the total task waiting duration of all algorithms. Among them, p5 represents the ratio (or relative length) of the average waiting time of a certain algorithm to be started to the total average waiting time of all algorithms. The task waiting duration of the algorithm represents the average waiting time for the task of this algorithm to be started to be processed, and the unit can be seconds, minutes, etc. The total task waiting duration of all algorithms represents the total average waiting time of all algorithms in the set of algorithms to be started.

[0096] Obtain the score of each algorithm to be started. The calculation formula is: score = p2×a + p3×b + p4×c + p5×d - p1×e. Among them, a, b, c, d, and e refer to weights, and specific values are obtained based on experience. And the setting of the weights a, b, c, d, and e can be adjusted according to the actual situation of the system and business requirements. For example, if the system has high requirements for real-time performance, the weight of the inference speed (p2) can be set relatively high; if it is desired to balance resource usage, the weight of the video memory occupancy (p1) can also be increased accordingly.

[0097] Sort all the algorithms to be started in descending order of scores to obtain the start ranking list of the algorithms to be started. This is a ranking list based on various performance indicators of the algorithms (such as the proportion of video card storage occupancy, the proportion of inference speed, the proportion of actual throughput, the proportion of unrecognized task volume, the proportion of task waiting duration, etc., and the scores obtained through weighted calculation). The higher the ranking of the algorithm to be started, the better its comprehensive performance may be, or it may be more in line with the scheduling requirements of the current system.

[0098] S2-5, as Figure 6 shown, start and stop the algorithms according to the start ranking list, including the following sub-steps:

[0099] S2-51. Take out the algorithms to be started one by one from the start ranking list in the order of the start ranking.

[0100] In specific implementation, under limited computing power resources, according to the start ranking list of the algorithms to be started, try to start the algorithm containers with higher rankings in turn.

[0101] S2-52. Determine whether the task to be recognized corresponding to the algorithm to be started can be completed within a preset time. If it can, it means that only one algorithm container needs to be started. If not, it means that multiple algorithm containers need to be started, and calculate the number of algorithm containers to be started.

[0102] Specifically, when determining whether the task to be recognized corresponding to the algorithm to be started can be completed within a preset time, first calculate the number of tasks that the algorithm to be started can complete. The calculation formula is: the number of tasks that can be completed = algorithm throughput × preset time (such as 10 minutes). Among them, the algorithm throughput refers to the total amount of tasks that can be processed per unit time, reflecting the overall processing efficiency; 10 minutes is the single-cycle waiting time of the scheduler, and can also be adjusted to other times according to the actual situation. Then determine whether the number of tasks that can be completed is greater than the total number of tasks to be recognized. If it is greater, it means that it can be completed within the preset time. If it is not greater, it means that it cannot be completed within the preset time.

[0103] If it can be completed within the preset time, it means that the number of algorithm containers to be started is 1.

[0104] If it cannot be completed within the preset time, it means that more than one algorithm container needs to be started, and calculate the specific number of algorithm containers to be started. The calculation formula is: the number of algorithm containers = the number of tasks to be recognized currently ÷ the number of tasks that can be completed within the preset time (such as 10 minutes).

[0105] Specifically, for each algorithm to be started, according to its throughput, calculate the number of algorithm containers to be started to ensure that all tasks to be recognized corresponding to the algorithm can be recognized within the next 10 minutes. And subsequently, start the algorithm containers according to the number of algorithm containers to be started.

[0106] S2-53. Based on the number of algorithm containers to be started, determine whether the GPU resources are sufficient. If they are sufficient, start one algorithm container. If not, determine whether there is a started algorithm with a lower rank than the currently to-be-started algorithm. If there is, close the algorithm container of the started algorithm with a lower rank to release the GPU resources, and repeat this step to continue determining whether the GPU resources are sufficient. If not, it means that more algorithm containers corresponding to algorithms with lower ranks cannot be closed, so exit the current algorithm container scheduling.

[0107] Specifically, when it is found that the GPU resources are not enough to start a new algorithm container, check whether there is an algorithm with a lower rank than the to-be-started algorithm among the started algorithms corresponding to the currently started algorithm containers. Here, "lower rank" means that these started algorithms are more backward in the start ranking list, perhaps because their comprehensive performance is poor, or because their current task load is light and the waiting time is long, etc.

[0108] If it is found that the algorithm containers corresponding to algorithms with lower rankings have been started, then in order to release GPU resources to algorithms that are more in need (i.e., have higher rankings), the algorithm containers corresponding to these algorithms with lower rankings will be selected to be closed. This is done to optimize resource utilization and ensure that high-performance or high-priority algorithms to be started can obtain sufficient resources to execute tasks.

[0109] After closing the algorithm containers corresponding to algorithms with lower rankings, check again whether the GPU resources are sufficient to start new algorithm containers. If they are still insufficient, this process may be repeated until sufficient GPU resources are found or there are no more algorithm containers corresponding to algorithms with lower rankings that can be closed. If new algorithm containers still cannot be started ultimately, then exit the current algorithm scheduling process.

[0110] S2-54. Determine whether all started algorithms are among the top several in the start ranking list; if so, end the current algorithm container scheduling; if not, return to S2-51, and take out the algorithms to be started one by one from the start ranking list in the order of start ranking until the start ranking list is traversed, and then end the current scheduling.

[0111] The purpose of this step is to check whether the currently started algorithms are all algorithms with high rankings in the start ranking list. The "top several" here does not specifically specify the quantity, but is judged based on whether the GPU resources are sufficient and can be set according to the actual situation.

[0112] In specific implementation, first view all the started algorithms and compare their rankings in the start ranking list. If all the started algorithms are ranked very high in the ranking list (i.e., they have high scores and good comprehensive performance), and at this time the GPU resources are not sufficient to start more algorithm containers, then it is considered that the optimal algorithm combination has been started currently and no further scheduling is required.

[0113] If the above conditions are met, directly end the current scheduling and no longer attempt to start other algorithms to be started or adjust the current algorithm combination. If the conditions are not met (i.e., there are still algorithms with higher rankings that have not been started, or although the started algorithms have high rankings, the GPU resources are still sufficient to start more algorithms), then continue to loop, traverse the algorithms to be started in the start ranking list, and make decisions on starting and stopping the algorithms according to the steps of S2-51.

[0114] S2-6. After ending the current scheduling, enter the sleep state and return to S2-1.

[0115] This step indicates the end of an algorithm container scheduling process, and then enters the sleep state (such as 10 minutes), and then starts a new algorithm container scheduling process again.

[0116] During specific implementation, after completing one round of algorithm container scheduling (whether because the optimal algorithm combination has been started or because the start-up ranking list has been traversed and all possible start / stop decisions have been made), the current algorithm container scheduling process ends. To reduce system load, save resources, and give the system a buffer and preparation time, a 10-minute sleep period is entered. During this period, no algorithm container scheduling or task processing is performed. After the sleep period ends, a new algorithm container scheduling process starts, begins to execute from step S2-1, re-evaluates the current task requirements, algorithm performance, GPU resources, etc., and makes new scheduling decisions.

[0117] According to the above technical solution, the algorithm task scheduling method of the intelligent recognition system of the present invention realizes the centralized management of tasks to be recognized and the intelligent scheduling of algorithm containers through the cooperation of the recognition request storage phase and the algorithm container scheduling phase. In the algorithm container scheduling phase, by comprehensively considering multiple factors such as the video memory occupancy, inference speed, actual throughput, unrecognized task volume, and task waiting duration of the algorithm container, a start-up ranking list of algorithms to be started is calculated, so as to realize the intelligent start and stop of algorithm containers under limited computing power resources, improve the execution efficiency of recognition tasks and the utilization rate of system resources.

[0118] S3. Recognition task execution phase:

[0119] After the scheduled algorithm container is started, it automatically enters the recognition task execution phase. Specifically: the currently started algorithm container obtains the tasks to be recognized assigned to the corresponding algorithm from the task list and marks the status of these tasks as "being recognized". Subsequently, the currently started algorithm container performs recognition processing on these tasks being recognized through the corresponding algorithm. When the algorithm container has recognized all the tasks being recognized assigned to it, it will enter the sleep state. This is to save resources and prevent the algorithm container from still occupying system resources when there are no tasks.

[0120] However, as Figure 3 shown, once an algorithm recognition request with a status of "to be recognized" is stored in the recognition task table, an algorithm new task notification signal will be triggered. If the algorithm container receives the new task notification signal during the sleep period, it will be awakened and continue to execute tasks, including the following sub-steps:

[0121] S3-1. When an algorithm recognition request with a status mark of "to be recognized" is stored in the recognition task table, an algorithm new task notification signal is triggered, and it is only triggered once within the pre-trial time to wake up the corresponding algorithm container in the sleep state.

[0122] During specific implementation, to avoid resource waste caused by repeated triggering of task notification signals within a short period, the system is set to trigger a new task notification signal only once within 5 seconds (this time can be configured as needed).

[0123] S3-2. The awakened algorithm container retrieves the task and determines whether there is a corresponding task to be recognized in the task list.

[0124] If there is no task to be recognized, it is determined whether no task has been retrieved after three loops (this number of times can be configured as needed). If not, the loop count is incremented by 1, and the process returns to the awakened algorithm container to retrieve the task, repeating this step S3-2; if so, the sleep mechanism is triggered, and a notification signal for the algorithm to enter the sleep state is sent to this algorithm container, causing it to enter the sleep state and no longer enter the algorithm container scheduling stage.

[0125] If there is a task to be recognized, after the awakened algorithm container retrieves the task to be recognized, it changes the status of the task to be recognized to "being recognized". Also, to ensure data consistency, transaction processing is required for the operations of reading and retrieving tasks.

[0126] In some embodiments, when retrieving tasks, preference is given to retrieving synchronous recognition tasks that are being recognized but whose execution time has exceeded half an hour (this time can also be configured).

[0127] Specifically, the algorithm recognition initiated by the business system is divided into synchronous recognition tasks and asynchronous recognition tasks. Asynchronous recognition tasks do not require timeliness, and the recognition results are notified to the business system through callbacks. Synchronous recognition tasks are given priority by the system because they require timeliness.

[0128] S3-3. The awakened algorithm container performs recognition processing on the task being recognized, as Figure 7 shown, including the following sub-steps: Before recognition, download the pictures stored in MinIO and construct recognition parameters. During recognition, call the corresponding algorithm through HTTP to recognize this task. After recognition, obtain the recognition result of this task by the awakened algorithm container.

[0129] S3-4. Obtain the recognition result and determine whether the returned recognition result is normal. If the recognition result is normal, it indicates that the callback recognition is successful; if the recognition result is abnormal, it indicates that the callback recognition fails.

[0130] During specific implementation, callback recognition refers to a mechanism in which during the recognition process, when the recognition result is generated, the result is returned to the developer through a callback. In this way, the developer can process or display the recognition result in the callback function, realize interaction with the recognition task, and improve the user experience.

[0131] S3-5. Determine whether the number of retries for callback recognition failure is greater than three (this number can be configured as needed). If the number of retries exceeds three, modify the recorded result to callback recognition failure; if the number of retries does not exceed three, modify the recorded result to callback recognition success.

[0132] S3-6. End the loop in the current execution stage of the recognition task and enter the next loop.

[0133] As can be seen from the above technical solutions, the process of the current execution stage of this recognition task ensures the efficient, stable, and reliable operation of the recognition task execution stage through reasonable task scheduling, priority processing, sleep mechanism, callback notification, and other strategies.

[0134] In summary, the method of the present invention has the following technical effects:

[0135] 1) The real-time performance monitoring technology of the algorithm container can accurately capture the resource usage and performance metrics of the algorithm container.

[0136] 2) The dynamic resource scheduling algorithm can achieve the optimal allocation of resources according to real-time data and historical trend prediction.

[0137] 3) The intelligent management mechanism of the task queue improves the task processing efficiency through priority sorting and task scheduling strategies.

[0138] 4) The multi-dimensional comprehensive evaluation system for scheduling decisions ensures the comprehensiveness and accuracy of scheduling decisions.

[0139] 5) The real-time monitoring and optimization technology of algorithm performance improves the execution efficiency and accuracy of the recognition task.

[0140] Embodiment 2: The dynamic scheduling system of the visual recognition algorithm container of the present invention.

[0141] As Figure 8 shown, this embodiment provides a dynamic scheduling system for a visual recognition algorithm container, including an algorithm container scheduling module; the algorithm container scheduling module includes an algorithm collection obtaining unit, an evaluation startup ranking unit, a container number calculating unit, a GPU resource releasing unit, and a startup ranking checking unit, which are specifically introduced as follows.

[0142] Among them, the algorithm collection obtaining unit is used to: obtain the algorithm collection of the current task to be recognized based on the recognition task table; obtain the algorithm collection of the currently started algorithms based on the algorithm list; take the union of the algorithm collection of the task to be recognized and the algorithm collection of the started algorithms to obtain the algorithm collection to be started.

[0143] The evaluation and startup ranking unit is used to: monitor the running status and performance metrics of the algorithm containers in real time, and evaluate the startup ranking of each to-be-started algorithm in the to-be-started algorithm set according to the video memory occupancy, inference speed, actual throughput, unrecognized task volume, and task waiting duration of the algorithm containers, so as to obtain the startup ranking table of the to-be-started algorithms.

[0144] The computing container number unit is used to: sequentially take out the to-be-started algorithms from the startup ranking table in the order of the startup ranking, and determine whether the corresponding to-be-recognized tasks can be completed within a preset time; if yes, it means that only one algorithm container needs to be started; if not, it means that multiple algorithm containers need to be started, and calculate the number of algorithm containers to be started.

[0145] The GPU resource release unit is used to: judge whether the GPU resources are sufficient based on the number of algorithm containers to be started; if sufficient, start one algorithm container; if not, judge whether there is a started algorithm with a lower ranking than the current to-be-started algorithm, if so, close the algorithm container of the started algorithm, and repeatedly judge whether the GPU resources are sufficient, if not, exit the current algorithm container scheduling.

[0146] The startup ranking check unit is used to: judge whether all the started algorithms are among the top several in the startup ranking table; if so, end the current algorithm container scheduling, if not, return to sequentially take out the to-be-started algorithms from the startup ranking table in the order of the startup ranking until the startup ranking table is traversed.

[0147] In other embodiments, the system of the present invention further includes an identification request storage module; the identification request storage module includes a request initiation unit, a parameter verification unit, and a storage unit, which are specifically introduced as follows.

[0148] Among them, the request initiation unit is used to: the business system service invokes the algorithm identification request, and sends the algorithm identification request including the to-be-identified data, request parameters, and unique identification identifier to the intelligent identification system service.

[0149] The parameter verification unit is used to: the intelligent identification system service performs parameter verification on the algorithm identification request; if the verification is successful, record the algorithm identification request in the database, mark it as to-be-identified, and generate a postgres table record including the unique identification identifier and the algorithm primary key identifier; if the verification fails, record the algorithm identification request in the database and mark it as parameter verification failed.

[0150] The storage unit is used to: store all the algorithm identification requests marked as to-be-identified into the identification task table to form a to-be-identified task queue.

[0151] In other embodiments, the system of the present invention further includes an identification task execution module; the identification task execution module includes an algorithm identification processing unit, an algorithm container wake-up unit, and a task fetching unit, which are specifically introduced as follows.

[0152] Among them, the algorithm identification processing unit is used for: the currently started algorithm container obtains the corresponding task to be identified from the identification task table, marks its status as being identified, and performs identification processing on it; when all tasks marked as being identified are identified, the algorithm container enters the sleep state.

[0153] The algorithm container wake-up unit is used for: once an algorithm identification request marked as to be identified is stored in the identification task table, it triggers an algorithm new task notification signal to wake up the corresponding algorithm container in the sleep state.

[0154] The task fetching unit is used for: the awakened algorithm container fetches tasks and determines whether there are corresponding tasks to be identified in the task list; if not, it determines whether the task has not been fetched after cycling a preset number of times; if not, the cycle count is incremented by 1, and it returns to the awakened algorithm container to fetch tasks; if so, it triggers an algorithm sleep notification signal to make the awakened algorithm container enter the sleep state; if there are tasks, after the awakened algorithm container fetches the task to be identified, it marks its status as being identified and performs identification processing on it to obtain an identification result and determines whether it is normal; if it is normal, it indicates that the callback identification is successful; if it is abnormal, it indicates that the callback identification fails; then it determines whether the retry count of the callback identification failure is greater than the preset number of times; if so, the recorded result is modified to the callback identification failure; if not, the recorded result is modified to the callback identification success.

[0155] In specific implementation, the system of the present invention uses different namespaces and containers in Kubernetes (K8S) to implement the work process. As Figure 9 shown, the entire process shows the process from task generation, storage, to script injection and algorithm container startup, involving the collaborative work of multiple namespaces and containers, which are specifically introduced as follows.

[0156] The K8S-default-cloud namespace is a default namespace in Kubernetes for storing and managing various resource objects. The ai-web pod container, the ai-service intelligent computing pod container, and the ai-service-algorithm-schedule-scripts algorithm scheduling pod container are Pod containers running in the K8S-default-cloud namespace, which are used for Web services, intelligent computing, and algorithm scheduling respectively. Moreover, the ai-web pod container initiates an algorithm recognition request to the ai-service intelligent computing pod container.

[0157] Redis and Postgres are basic services on the host for data storage. Moreover, Redis is used for data storage and access of task reading and execution. The ai-service intelligent computing pod container stores all algorithm recognition requests marked with pending status into the recognition task table of Postgres to form a pending recognition task queue. Once a recognition task is monitored, the ai-service-algorithm-schedule-scripts algorithm scheduling pod container fetches the pending recognition task from the recognition task table of Postgres and performs algorithm container scheduling.

[0158] The K8S-(ai-model)-(ai-model) namespace is another Kubernetes namespace dedicated to storing and managing resource objects related to ai-model. In the ai-model namespace, there are multiple algorithm containers (such as algorithm 1 pod container, algorithm 2 pod container,..., algorithm N pod container), which are started after injecting scripts and are used to execute specific algorithm tasks. Moreover, multiple algorithm containers are started after injecting scripts. The ai-inject-service pod container is used to ensure that the algorithm containers can load these scripts when starting from the PVC directory of the injection script to the ai-model namespace.

[0159] Moreover, the recognition request storage module includes the ai-web pod container, the ai-service intelligent computing pod container, Redis, and Postgres, and has the following functions: managing the queue of visual recognition tasks, including task reception, queuing, and status update. This module adopts an advanced queue management algorithm to intelligently schedule tasks according to the urgency of tasks and resource requirements, ensuring that high-priority tasks can be processed quickly.

[0160] The algorithm container scheduling module includes the ai-service-algorithm-schedule-scripts algorithm scheduling pod container, which has the following functions: (1) Resource scheduling decision: Dynamically calculate the resource allocation and replica number adjustment strategy according to the running status of the algorithm container and the task queue. And use machine learning algorithms to predict resource requirements and task loads, and automatically adjust the replica number of the algorithm container to adapt to the changing resource requirements. (2) Container scheduling execution: Execute the start, stop, and resource adjustment operations of the algorithm container. And through the interface with the container management system, achieve precise control of the algorithm container to ensure that the resource scheduling decision can be executed quickly and accurately. (3) Scheduling log recording: Record the key information during the scheduling process for system monitoring and troubleshooting. And detailedly record the detailed information of each scheduling operation, including scheduling time, resource allocation situation, task processing results, etc., to provide data support for system maintenance and optimization.

[0161] The recognition task execution module includes multiple algorithm containers, the ai-inject-service pod container, and has the following functions: (1) Responsible for executing visual recognition tasks and monitoring algorithm performance in real time. This module integrates high-performance visual recognition algorithms and can monitor performance indicators such as recognition speed and accuracy of the algorithm in real time while executing tasks to ensure the efficient completion of recognition tasks.

[0162] Embodiment 3: The electronic device of the present invention.

[0163] As Figure 10 shown, this embodiment provides an electronic device, including a processor and a memory coupled to the processor; wherein, the memory has computer program instructions stored therein; when the computer program instructions are executed by the processor, the electronic device executes a visual recognition algorithm container dynamic scheduling method as described in Embodiment 1.

[0164] Specifically, the electronic device includes a central processing unit (CPU) and / or a graphics processing unit (GPU), which can execute various appropriate actions and processes according to the computer program instructions stored in the read-only memory (ROM) or the computer program instructions loaded from the storage unit to the random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored. The CPU / GPU, ROM, and RAM are connected to each other through a bus. The input / output (I / O) interface is also connected to the bus. In addition, the electronic device may further include a coprocessor.

[0165] Moreover, multiple components in the electronic device are connected to the I / O interface, including: an input unit, such as a keyboard, a mouse, etc.; an output unit, such as various types of displays, speakers, etc.; a storage unit, such as a disk, an optical disc, etc.; and a communication unit, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit allows the device to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0166] Each of the methods or processes described above can be executed by a CPU / GPU. For example, in some embodiments, the method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the CPU / GPU, one or more steps or actions of the methods or processes described above can be performed.

[0167] Embodiment 4: The computer-readable storage medium of the present invention.

[0168] This embodiment provides a computer-readable storage medium, including computer program instructions stored therein, and when the computer program instructions are executed, a visual recognition algorithm container dynamic scheduling method described in Embodiment 1 is implemented.

[0169] Specifically, a computer-readable storage medium can be a tangible device that can hold and store computer program instructions used by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. More specific examples (a non-exhaustive list) of a computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punched card or raised structures in grooves storing computer program instructions thereon, and any suitable combination of the above. The computer-readable storage medium used herein is not construed as an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0170] The computer program instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer program instructions from the network and forwards the computer program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0171] Moreover, the computer program instructions for performing the operations of the present disclosure can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages and conventional procedural programming languages. The computer program instructions can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., using an Internet service provider to connect through the Internet). In some embodiments, by using the state information of the computer program instructions to customize an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer program instructions to implement various aspects of the present disclosure.

[0172] These computer program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is produced that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other devices to work in a specific manner. Thus, the computer-readable medium storing the instructions includes a manufactured article that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0173] In addition, computer program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices, so that a series of operation steps are executed on the computer, other programmable data processing apparatus, or other devices to generate a computer-implemented process, thereby enabling the instructions executed on the computer, other programmable data processing apparatus, or other devices to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0174] Moreover, the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of methods, systems, and devices according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, and the module, segment of a program, or part of an instruction contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the boxes may occur in a different order from that marked in the accompanying drawings. For example, two consecutive boxes may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as combinations of boxes in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.

[0175] The various embodiments of the present disclosure have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements to the technologies in the market, or to enable other ordinary skill in the art in the technical field to understand the embodiments disclosed herein.

Claims

1. A method for dynamically scheduling a visual recognition algorithm container, characterized in that: Includes the algorithm container scheduling stage; The algorithm container scheduling phase includes the following steps: Based on the recognition task table, obtain the algorithm collection corresponding to the current task to be recognized; based on the algorithm list, obtain the currently started algorithm collection; take the union of the algorithm collection corresponding to the task to be recognized and the started algorithm collection to obtain the algorithm collection to be started; Monitor the running status and performance indicators of the algorithm container in real time, and evaluate the startup ranking of each algorithm to be started in the set of algorithms to be started based on the memory usage, inference speed, actual throughput, number of unidentified tasks, and task waiting time of the algorithm container, and obtain a startup ranking table of the algorithms to be started; According to the order of the startup ranking, the algorithms to be started are taken out one by one from the startup ranking table to determine whether the corresponding tasks to be identified can be completed within the preset time; if yes, it means that only one algorithm container needs to be started; if not, it means that multiple algorithm containers need to be started, and the number of algorithm containers that need to be started is calculated; Based on the number of algorithm containers that need to be started, determine whether GPU resources are sufficient; if so, start an algorithm container; if not, determine whether there is an algorithm that has been started with a lower ranking than the current algorithm to be started. If so, close the algorithm container of the started algorithm and repeat the determination of whether GPU resources are sufficient. If not, exit the current algorithm container scheduling; Determine whether all started algorithms are in the first few in the startup ranking table; if so, end the algorithm container scheduling; if not, return to the order of startup ranking, and take out the algorithms to be started one by one from the startup ranking table until the startup ranking table is traversed.

2. According to claim 1, a method for dynamically scheduling a visual recognition algorithm container is characterized in that: The step of evaluating the startup ranking of each algorithm to be started in the set of algorithms to be started to obtain a startup ranking table of the algorithms to be started includes the following sub-steps: Get the proportion of video memory usage p1, inference speed p2, actual throughput p3, unrecognized task volume p4, and task waiting time p5 of each algorithm to be started; Among them, the proportion of video memory usage is p1, and the calculation formula is: p1 = algorithm video memory usage / total algorithm video memory usage; The reasoning speed accounts for p2, and the calculation formula is: p2 = algorithm reasoning speed / total algorithm reasoning speed; The actual throughput accounts for p3, calculated as follows: p3 = actual algorithm throughput / total algorithm throughput; The proportion of tasks to be identified is p4, and the calculation formula is: p4 = number of tasks to be identified by the algorithm / total number of tasks to be identified; The task waiting time accounts for p5, and the calculation formula is: p5 = algorithm task waiting time / total algorithm task waiting time; According to the preset weights a, b, c, d, e, the score of each algorithm to be started is calculated. The calculation formula is: score = p2×a+p3×b+p4×c+p5×d-p1×e; All algorithms to be started are arranged in descending order according to their scores to obtain a startup ranking list of the algorithms to be started.

3. The method for dynamically scheduling a visual recognition algorithm container according to claim 1, characterized in that: The step of judging whether the corresponding task to be identified can be completed within the preset time includes the following sub-steps: Calculate the number of tasks that can be completed by the algorithm to be started. The calculation formula is: number of tasks that can be completed = algorithm throughput × preset time; Determine whether the number of tasks that can be executed is greater than the total number of tasks to be identified; If yes, it means that the task can be completed within the preset time; if no, it means that the task cannot be completed within the preset time.

4. The method for dynamically scheduling a visual recognition algorithm container according to claim 1, characterized in that: The calculation is performed to determine the number of algorithm containers that need to be started, and the calculation formula is: number of algorithm containers = number of tasks to be identified currently / number of tasks that can be executed within a preset time.

5. A method for dynamically scheduling visual recognition algorithm containers according to any one of claims 1 to 4, characterized in that: Before the algorithm container scheduling stage, there is also an identification request storage stage; the identification request storage stage includes the following steps: The business system service calls the algorithm identification request and sends the algorithm identification request including the data to be identified, request parameters and unique identification identifier to the intelligent identification system service; The intelligent identification system service performs parameter verification on the algorithm identification request; if the verification succeeds, the algorithm identification request is recorded in the database, and its status is marked as pending identification, and a postgres table record containing a unique identification identifier and an algorithm primary key identifier is generated; if the verification fails, the algorithm identification request is recorded in the database, and its status is marked as parameter verification failure; All algorithm recognition requests marked as pending recognition are stored in the recognition task table to form a pending recognition task queue.

6. A method for dynamically scheduling visual recognition algorithm containers according to any one of claims 1 to 4, characterized in that: After the algorithm container scheduling phase, an identification task execution phase is also included; the identification task execution phase includes the following steps: The currently started algorithm container obtains the corresponding task to be identified from the identification task table, marks its status as being identified, and performs identification processing on it; When all tasks marked as being identified are identified, the algorithm container enters a dormant state.

7. A method for dynamically scheduling visual recognition algorithm containers according to claim 6, characterized in that: The identification task execution phase also includes the following steps: Once an algorithm recognition request with a status marked as pending recognition is stored in the recognition task table, the algorithm new task notification signal is triggered to wake up the corresponding algorithm container in the dormant state; The awakened algorithm container takes the task and determines whether there is a corresponding task to be identified in the task list; If not, determine whether the task is not obtained after the preset number of cycles; if not, increase the number of cycles by 1 and return to the awakened algorithm container to obtain the task; if so, trigger the algorithm to enter the sleep notification signal, so that the awakened algorithm container enters the sleep state; If yes, the awakened algorithm container will take the task to be identified, mark its status as being identified, perform identification processing on it, get the identification result and judge whether it is normal; if normal, it means that the callback identification is successful; if abnormal, it means that the callback identification fails; then judge whether the number of retries of the callback identification failure is greater than the preset number, if so, modify the record result to callback identification failure; if not, modify the record result to callback identification success.

8. A visual recognition algorithm container dynamic scheduling system, characterized in that: It includes an algorithm container scheduling module; the algorithm container scheduling module includes an algorithm collection acquisition unit, a startup ranking evaluation unit, a container quantity calculation unit, a GPU resource release unit, and a startup ranking check unit; The algorithm collection acquisition unit is used to: acquire the algorithm collection corresponding to the current task to be identified based on the identification task table; acquire the currently started algorithm collection based on the algorithm list; and obtain the algorithm collection corresponding to the task to be identified and the started algorithm collection by taking the union of the algorithm collection corresponding to the task to be identified and the started algorithm collection to obtain the algorithm collection to be started; The evaluation startup ranking unit is used to: monitor the running status and performance indicators of the algorithm container in real time, and evaluate the startup ranking of each algorithm to be started in the set of algorithms to be started according to the video memory occupancy, inference speed, actual throughput, unrecognized task amount and task waiting time of the algorithm container, so as to obtain a startup ranking table of the algorithms to be started; The container quantity calculation unit is used to: take out the algorithms to be started one by one from the startup ranking table according to the startup ranking order, and determine whether the corresponding tasks to be identified can be completed within the preset time; If yes, it means that only one algorithm container needs to be started; if no, it means that multiple algorithm containers need to be started and the number of algorithm containers that need to be started needs to be calculated; The GPU resource release unit is used to: determine whether the GPU resources are sufficient based on the number of algorithm containers that need to be started; if sufficient, start an algorithm container; if insufficient, determine whether there is a started algorithm with a lower ranking than the current algorithm to be started, if yes, close the algorithm container of the started algorithm, and repeatedly determine whether the GPU resources are sufficient, if no, exit the current algorithm container scheduling; The startup ranking checking unit is used to determine whether all started algorithms are in the first few positions in the startup ranking table; if so, end the algorithm container scheduling; if not, return to the order of startup ranking, and take out the algorithms to be started one by one from the startup ranking table until the startup ranking table is traversed.

9. An electronic device, characterized in that: It includes a processor and a memory coupled to the processor; the memory has computer program instructions stored therein; when the computer program instructions are executed by the processor, the electronic device executes a visual recognition algorithm container dynamic scheduling method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The invention comprises computer program instructions stored therein; when the computer program instructions are executed, the method for dynamically scheduling a visual recognition algorithm container as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Artificial intelligence computer vision reasoning method

    CN114298313A

  • GPU resource scheduling method and system based on dynamic weight calculation

    CN119415245A