Visual identification algorithm container dynamic scheduling method, system, equipment and medium

Through real-time monitoring and dynamic scheduling of visual recognition algorithm containers, the limitations of existing systems in resource allocation, task queue management and performance optimization are solved, and efficient algorithm container management and visual recognition task processing are realized.

CN120107759AActive Publication Date: 2025-06-06QIANXUN TECH (SHENZHEN) CO LTD +1

Patent Information

Application Number
CN202510586443.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-06-06
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

The existing visual identification system has limitations in resource allocation strategy, task queue management, real-time monitoring and performance optimization, resulting in waste of resources, low task processing speed and low recognition efficiency.

Method used

A dynamic scheduling method for containers of visual recognition algorithm is proposed. By monitoring the running status and performance indicators of the algorithm container in real time, evaluating the startup ranking of the algorithm to be started, dynamically adjusting resource allocation, optimizing task queue management, and dynamically adjusting algorithm parameters based on real-time data.

Benefits of technology

It realizes efficient management and optimized execution of algorithm containers, improves the processing speed of visual recognition tasks and the overall performance of the system, and improves resource utilization and accuracy of identification tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107759A_ABST
    Figure CN120107759A_ABST
Patent Text Reader

Abstract

The invention discloses a visual recognition algorithm container dynamic scheduling method, system and device and a medium. The method comprises the steps that a to-be-started algorithm set is obtained; obtaining a starting ranking table of the to-be-started algorithm; according to the sequence of the starting ranking, the to-be-started algorithms are taken out from the starting ranking table one by one, and whether the to-be-recognized tasks corresponding to the to-be-started algorithms can be executed within preset time or not is judged; if yes, only one algorithm container needs to be started; if not, a plurality of algorithm containers need to be started, and the number of the algorithm containers needing to be started is calculated; based on the number of the algorithm containers needing to be started, whether GPU resources are sufficient or not is judged; judging whether all started algorithms are the first few bits in a starting ranking table or not; according to the method, the resource allocation of the algorithm container is dynamically adjusted, so that intelligent scheduling and dynamic capacity expansion and contraction of the algorithm container are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method, system, device and medium for dynamically scheduling a visual recognition algorithm container, and belongs to the technical field of cloud computing resource management. Background Art

[0002] At present, visual recognition technology has become an indispensable part of the modern technology ecosystem. It provides users with unprecedented interactive experience and decision support through functions such as image recognition, object detection and scene understanding. In addition, driven by artificial intelligence technology, machine learning and deep learning algorithms have made significant progress in the field of visual recognition, enabling computer vision systems to process complex image and video data and be applied to multiple industries such as autonomous driving, medical diagnosis, and security monitoring.

[0003] In order to cope with the rapidly changing market demands, algorithm containerization technology has emerged, which allows algorithms to be quickly deployed in different computing environments in the form of containers, improving the portability and scalability of algorithms. Although containerization technology brings convenience, it also introduces new resource management issues, which are described below.

[0004] 1) Limitations of resource allocation strategies: Existing visual recognition systems usually adopt polling or static resource allocation strategies. These strategies lack flexibility in resource allocation and cannot dynamically adjust resources according to the actual needs of the task. As a result, resources are often insufficient during peak task loads and cannot meet the processing requirements of high-concurrency tasks. During low task loads, a large amount of resources are wasted, reducing the overall efficiency and economic benefits of the system.

[0005] 2) Limitations of task queue management and scheduling decisions: Existing visual recognition systems lack an effective mechanism to manage and optimize task queues, and often fail to implement priority scheduling of tasks, resulting in an inability to respond quickly under high load conditions and a waste of resources under low load conditions, making it difficult for the system to make reasonable scheduling decisions based on the urgency of the task and resource requirements. This not only affects the processing speed of visual recognition tasks, but also limits the dynamic expansion and contraction capabilities of the algorithm container in a multi-tasking environment, making it impossible to achieve optimal allocation and utilization of resources.

[0006] 3) Lack of real-time monitoring: Existing visual recognition systems often ignore the real-time monitoring of the algorithm container performance when performing visual recognition tasks. This neglect limits the system's refined management of resources, resulting in low recognition efficiency, inability to fully utilize the computing power of the algorithm container, and difficulty in ensuring the accuracy and reliability of the recognition results.

[0007] 4) Lack of performance optimization: Existing visual recognition systems lack effective performance monitoring and optimization mechanisms, making it difficult to dynamically adjust algorithm parameters based on real-time data during the execution of visual recognition tasks. This is not conducive to improving recognition efficiency and accuracy. This shortcoming is particularly prominent in application scenarios that require high efficiency and high precision, which seriously affects the overall performance of the system and user experience.

[0008] From the above, it can be seen that how to effectively schedule and manage algorithm containers under limited resources has become an urgent problem to be solved in the field of cloud computing. Therefore, the present invention urgently needs to develop a dynamic, efficient, and intelligent visual recognition algorithm container scheduling and execution method and system. Summary of the invention

[0009] In response to the above-mentioned existing technical problems, the present invention provides a method, system, device and medium for dynamic scheduling of visual recognition algorithm containers to solve the problems of efficiency and resource utilization in algorithm container scheduling and execution tasks in existing visual recognition systems, so as to achieve the technical purpose of realizing efficient management and optimized execution of algorithm containers, and improving the processing speed of visual recognition tasks and the overall performance of the system.

[0010] To achieve the above technical objectives, firstly, the present invention provides a method for dynamically scheduling a visual recognition algorithm container, including an algorithm container scheduling stage; the algorithm container scheduling stage includes the following steps: Based on the recognition task table, obtain the algorithm collection corresponding to the current task to be recognized; based on the algorithm list, obtain the currently started algorithm collection; take the union of the algorithm collection corresponding to the task to be recognized and the started algorithm collection to obtain the algorithm collection to be started.

[0011] Monitor the running status and performance indicators of the algorithm container in real time, and evaluate the startup ranking of each algorithm to be started in the set of algorithms to be started based on the memory usage, inference speed, actual throughput, number of unidentified tasks, and task waiting time of the algorithm container, and obtain the startup ranking table of the algorithms to be started.

[0012] According to the order of startup ranking, the algorithms to be started are taken out one by one from the startup ranking table to determine whether the corresponding tasks to be identified can be completed within the preset time; if so, it means that only one algorithm container needs to be started; if not, it means that multiple algorithm containers need to be started, and the number of algorithm containers that need to be started is calculated.

[0013] Based on the number of algorithm containers that need to be started, determine whether the GPU resources are sufficient; if sufficient, start an algorithm container; if not, determine whether there is an already started algorithm with a lower ranking than the current algorithm to be started. If so, close the algorithm container of the started algorithm and repeat the determination of whether the GPU resources are sufficient. If not, exit this algorithm container scheduling.

[0014] Determine whether all started algorithms are in the first few in the startup ranking table; if so, end the algorithm container scheduling; if not, return to the order of startup ranking, and take out the algorithms to be started one by one from the startup ranking table until the startup ranking table is traversed.

[0015] Specifically, the method of the present invention includes evaluating the startup ranking of each algorithm to be started in the set of algorithms to be started to obtain a startup ranking table of the algorithms to be started, including the following sub-steps: Get the proportion of video memory usage p1, inference speed p2, actual throughput p3, unrecognized task volume p4, and task waiting time p5 of each algorithm to be started; Among them, the proportion of video memory usage is p1, and the calculation formula is: p1 = algorithm video memory usage / total algorithm video memory usage.

[0016] The reasoning speed accounts for p2, and the calculation formula is: p2 = algorithm reasoning speed / total algorithm reasoning speed.

[0017] The actual throughput accounts for p3, calculated as follows: p3 = actual algorithm throughput / total algorithm throughput.

[0018] The proportion of tasks to be identified is p4, and the calculation formula is: p4 = number of tasks to be identified by the algorithm / total number of tasks to be identified.

[0019] The task waiting time accounts for p5, and the calculation formula is: p5 = algorithm task waiting time / total algorithm task waiting time.

[0020] According to the preset weights a, b, c, d, and e, the score of each algorithm to be started is calculated. The calculation formula is: score = p2×a + p3×b + p4×c + p5×d - p1×e.

[0021] All algorithms to be started are arranged in descending order according to their scores to obtain a startup ranking list of the algorithms to be started.

[0022] Specifically, the method of the present invention determines whether the corresponding task to be identified can be completed within a preset time, including the following sub-steps: Calculate the number of tasks that can be executed by the algorithm to be started. The calculation formula is: number of tasks that can be executed = algorithm throughput × preset time.

[0023] Determine whether the number of tasks that can be executed is greater than the total number of tasks to be identified; If yes, it means that the task can be completed within the preset time; if no, it means that the task cannot be completed within the preset time.

[0024] The method of the present invention is specific, the calculation of the number of algorithm containers that need to be started is calculated by the following formula: number of algorithm containers = number of tasks to be identified currently / number of tasks that can be executed within a preset time.

[0025] The method of the present invention further includes a phase of identifying and placing requests in a warehouse before the algorithm container scheduling phase; the phase of identifying and placing requests in a warehouse includes the following steps: The business system service calls the algorithm recognition request and sends the algorithm recognition request containing the data to be recognized, request parameters and a unique identification identifier to the intelligent recognition system service.

[0026] The intelligent identification system service performs parameter verification on the algorithm identification request; if the verification is successful, the algorithm identification request is recorded in the database, and its status is marked as pending identification, and a postgres table record containing a unique identification identifier and an algorithm primary key identifier is generated; if the verification fails, the algorithm identification request is recorded in the database, and its status is marked as parameter verification failure.

[0027] All algorithm recognition requests marked as pending recognition are stored in the recognition task table to form a pending recognition task queue.

[0028] The method of the present invention further includes, after the algorithm container scheduling stage, an identification task execution stage, wherein the identification task execution stage includes the following steps: The currently started algorithm container obtains the corresponding task to be identified from the identification task table, marks its status as being identified, and performs identification processing on it.

[0029] When all tasks marked as being identified are identified, the algorithm container enters a dormant state.

[0030] Furthermore, the method of the present invention further comprises the following steps: Once an algorithm recognition request with a status marked as pending recognition is stored in the recognition task table, a new algorithm task notification signal is triggered to wake up the corresponding algorithm container in a dormant state.

[0031] The awakened algorithm container takes the task and determines whether there is a corresponding task to be identified in the task list; If not, determine whether the task is not obtained after the preset number of cycles; if not, increase the number of cycles by 1 and return to the awakened algorithm container to obtain the task; if so, trigger the algorithm to enter the sleep notification signal, so that the awakened algorithm container enters the sleep state; If yes, the awakened algorithm container will take the task to be identified, mark its status as being identified, perform identification processing on it, get the identification result and judge whether it is normal; if normal, it means that the callback identification is successful; if abnormal, it means that the callback identification fails; then judge whether the number of retries of the callback identification failure is greater than the preset number, if so, modify the record result to callback identification failure; if not, modify the record result to callback identification success.

[0032] Secondly, the present invention also provides a visual recognition algorithm container dynamic scheduling system, including an algorithm container scheduling module; the algorithm container scheduling module includes an algorithm collection acquisition unit, an evaluation startup ranking unit, a container quantity calculation unit, a GPU resource release unit, and a startup ranking check unit; The algorithm collection acquisition unit is used to: acquire the algorithm collection corresponding to the current task to be identified based on the identification task table; acquire the currently started algorithm collection based on the algorithm list; and obtain the algorithm collection corresponding to the task to be identified and the started algorithm collection by taking the union of the algorithm collection corresponding to the task to be identified and the started algorithm collection to obtain the algorithm collection to be started; The evaluation startup ranking unit is used to: monitor the running status and performance indicators of the algorithm container in real time, and evaluate the startup ranking of each algorithm to be started in the set of algorithms to be started according to the video memory occupancy, inference speed, actual throughput, unrecognized task amount and task waiting time of the algorithm container, so as to obtain a startup ranking table of the algorithms to be started; The container quantity calculation unit is used to: take out the algorithms to be started one by one from the startup ranking table according to the startup ranking order, and judge whether the corresponding tasks to be identified can be completed within the preset time; if yes, it means that only one algorithm container needs to be started; if not, it means that multiple algorithm containers need to be started, and the number of algorithm containers to be started is calculated; The GPU resource release unit is used to: determine whether the GPU resources are sufficient based on the number of algorithm containers that need to be started; if sufficient, start an algorithm container; if insufficient, determine whether there is a started algorithm with a lower ranking than the current algorithm to be started, if yes, close the algorithm container of the started algorithm, and repeatedly determine whether the GPU resources are sufficient, if no, exit the current algorithm container scheduling; The startup ranking checking unit is used to determine whether all started algorithms are in the first few positions in the startup ranking table; if so, end the algorithm container scheduling; if not, return to the order of startup ranking, and take out the algorithms to be started one by one from the startup ranking table until the startup ranking table is traversed.

[0033] Thirdly, the present invention provides an electronic device, including a processor and a memory coupled to the processor; the memory has computer program instructions stored therein; when the computer program instructions are executed by the processor, the electronic device executes the method for dynamic scheduling of a visual recognition algorithm container.

[0034] Fourthly, the present invention further provides a computer-readable storage medium, including computer program instructions stored therein; when the computer program instructions are executed, the dynamic scheduling method of a visual recognition algorithm container is implemented.

[0035] In summary, the present invention provides a dynamic, efficient and intelligent visual recognition algorithm container scheduling and execution method. The method can dynamically adjust the resource allocation of the algorithm container according to the actual needs of the visual recognition task, optimize the task queue management, and realize the intelligent scheduling and dynamic expansion and contraction of the algorithm container. At the same time, the method can also monitor the algorithm performance in real time, and dynamically adjust the algorithm parameters according to the monitoring results to improve the processing speed and accuracy of the recognition task, thereby improving the overall performance of the system and user satisfaction.

[0036] Moreover, the present invention is particularly suitable for the dynamic scheduling and execution of artificial intelligence visual recognition algorithms. Through the intelligent scheduling system, efficient management and optimized execution of algorithm containers are achieved, the processing speed of visual recognition tasks and the overall performance of the system are improved, especially in terms of resource allocation, task scheduling and performance monitoring, an innovative system is provided, and the specific technical advantages are as follows: 1. Real-time monitoring and dynamic scheduling: The system can monitor the performance indicators of the algorithm container in real time and dynamically adjust resource allocation based on the monitoring data. This real-time monitoring and dynamic adjustment mechanism enables the system to quickly respond to changes in resource requirements and improve resource utilization.

[0037] 2. Intelligent management of task queues: Manage task queues through intelligent algorithms to optimize task processing order and resource allocation. This intelligent management mechanism ensures that high-priority tasks are processed first, improving the efficiency of task processing.

[0038] 3. Multi-dimensional scheduling decision: Comprehensively consider multiple dimensions such as task urgency, resource utilization, and algorithm performance to achieve the optimal scheduling decision. This multi-dimensional consideration makes the scheduling decision more comprehensive and accurate, and improves the scheduling performance of the system.

[0039] 4. Real-time monitoring of algorithm performance: When performing recognition tasks, the algorithm performance is monitored in real time, and the algorithm parameters are adjusted dynamically to improve recognition efficiency. This real-time monitoring and dynamic adjustment mechanism ensures that the algorithm can run in the best state and improves the accuracy and efficiency of the recognition task. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 A flowchart of the stage of identifying a request to be put into storage when the method of the present invention is implemented; Figure 2 It is a flow chart of the algorithm container scheduling phase when the method of the present invention is implemented; Figure 3 A flow chart showing the identification task execution phase when the method of the present invention is implemented; Figure 4 This is an algorithm list interface diagram of S2-2 in the algorithm container scheduling stage when the method of the present invention is implemented; Figure 5 It is a flow chart of sub-steps S2-4 in the algorithm container scheduling stage when the method of the present invention is implemented; Figure 6 It is a flow chart of sub-steps S2-5 in the algorithm container scheduling stage when the method of the present invention is implemented; Figure 7 It is a sub-step flow chart of S3-3 in the identification task execution phase when the method of the present invention is implemented; Figure 8 It is a principle block diagram when the system of the present invention is implemented; Fig. 9 A workflow diagram for using different namespaces and containers in Kubernetes (K8S) when implementing the system of the present invention; Fig.10 It is a principle block diagram when the device of the present invention is implemented. DETAILED DESCRIPTION

[0041] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0042] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects, unless explicitly stated. Other explicit and implicit definitions may also be included below.

[0043] Embodiment 1: The visual recognition algorithm container dynamic scheduling method of the present invention.

[0044] This embodiment provides a method for dynamically scheduling a visual recognition algorithm container, including a recognition request storage phase, an algorithm container scheduling phase, and a recognition task execution phase. Figure 1 , Figure 2 , Figure 3 As shown in the figure, the logic of resource allocation and scheduling decisions at each stage is described in detail, and how to dynamically adjust the resource allocation and number of copies of the algorithm container based on real-time monitoring data and task queue status is demonstrated. The details are as follows.

[0045] S1, identification request storage stage: Figure 1 As shown, the identification request storage stage includes initiating an algorithm identification request, parameter verification, and storage into the identification task table. The specific steps are as follows.

[0046] S1-1. The business system service (including but not limited to the client or the upstream system that needs to call algorithm identification) initiates an algorithm identification request and sends the algorithm identification request to the intelligent identification system service. The algorithm identification request contains all the information required for algorithm identification, including the data to be identified (such as images, videos, etc.), the requested parameters (such as the type of identification algorithm, the accuracy requirements of identification, etc.), and a unique identification identifier used to uniquely identify the request.

[0047] S1-2, after receiving the algorithm recognition request, the intelligent recognition system service first performs parameter verification on it, including authentication information verification and necessary parameter verification, and determines whether the verification is successful.

[0048] If the verification fails, a failure is returned, the algorithm identification request is recorded in the database, and its status is marked as parameter verification failure.

[0049] If the verification is successful, the algorithm recognition request is also recorded in the database and its status is marked as pending recognition. At the same time, a postgres table record is generated in the database.

[0050] In specific implementation, the postgres table record includes the unique identification identifier brought by the business system service and the algorithm primary key identifier. It should be noted that each time the business system service calls the algorithm identification request, it needs to carry a unique identification identifier as the unique identifier of the algorithm identification request. In addition, since there are multiple recognition algorithms (such as face recognition algorithms, meter recognition algorithms, etc.), the algorithm primary key identifier is required to represent information such as which algorithm is used for recognition.

[0051] S1-3. All algorithm identification requests marked as pending identification are stored in the identification task table to form a pending identification task queue. Each pending identification task in the pending identification task queue corresponds to an algorithm identification request to be identified and an algorithm to be started (specified by the algorithm primary key identifier).

[0052] In specific implementation, the recognition task table is a postgres database table, which is used to store the algorithm recognition request information to be recognized. When the status of an algorithm recognition request is marked as "to be recognized" and stored in the recognition task table, it becomes a task to be recognized, waiting for the system to assign an algorithm and perform recognition work. This step is to centrally manage all algorithm recognition requests with a status of "to be recognized" to form a queue of tasks to be recognized. The system will take out tasks from this queue of tasks to be recognized according to certain strategies (such as first-in-first-out, priority, etc.) and execute the corresponding algorithm to complete the recognition work.

[0053] S2, algorithm container scheduling phase: mainly divided into the algorithm ranking sub-phase and the algorithm start and stop sub-phase, and is automatically scheduled every 10 minutes, as described below.

[0054] First, the algorithm ranking sub-stage: obtain the algorithms of all tasks to be identified from the identification task table, obtain all started algorithms from the algorithm list, thereby obtaining all algorithms to be started, and evaluate the startup ranking of each algorithm to be started. The factors for evaluating the startup ranking include the memory usage of the algorithm container, inference speed, actual throughput, the number of unidentified tasks, and task waiting time, so as to obtain the startup ranking table of the algorithms to be started.

[0055] Second, the algorithm start and stop sub-stage: The core of this stage is to start the top-ranked algorithms as much as possible according to the startup ranking list of the algorithms to be started under limited computing resources until resources are insufficient. Specifically, first start the algorithm to be started that ranks highest in the startup ranking list, and calculate how many algorithm containers need to be started based on the throughput capacity of the algorithm to be started in order to identify all the tasks to be identified corresponding to the algorithm within the preset time. Then call the k8s API to start the algorithm container, and so on, until resources are insufficient, then end this stage.

[0056] like Figure 2 As shown, the specific steps of the container scheduling phase of this algorithm are as follows.

[0057] S2-1. Based on the recognition task table, obtain the algorithm collection corresponding to the current task to be recognized.

[0058] During the specific implementation, the recognition task table is traversed, and all unique algorithms are extracted according to the algorithm primary key identifier in the algorithm recognition request to form a collection of algorithms for the task to be recognized. This refers to the set of algorithms that can be used to process requests to be recognized in the intelligent recognition system service. These algorithms may include face recognition algorithms, meter recognition algorithms, etc., and each algorithm has its specific application scenarios and recognition capabilities. In addition, after the status of the algorithm recognition request is marked as "to be recognized", the intelligent recognition system service will select the corresponding algorithm from the algorithm collection to process the request according to the algorithm primary key identifier specified in the algorithm recognition request.

[0059] S2-2. Based on the algorithm list, obtain the currently started algorithm collection.

[0060] When implementing it, Figure 4 As shown, the algorithm list maintains multiple algorithms, such as face recognition algorithm, meter recognition algorithm, personnel detection algorithm, etc. Access the algorithm list and filter out algorithms with a status of "started" to form a collection of started algorithms.

[0061] S2-3. Take the union of the above two algorithm sets to obtain the algorithm set to be started.

[0062] In specific implementation, since the algorithms of the tasks to be identified and the algorithms that have been started need to be scheduled, the algorithm set of the tasks to be identified and the set of algorithms that have been started are merged into a new set, and duplicates are removed to form a collection of algorithms to be started. This set contains all the algorithms that appear in the algorithm set of the tasks to be identified or the set of algorithms that have been started, but each algorithm only appears once (that is, duplicate elements are removed). In practical applications, this means that all algorithms of the tasks to be identified and all algorithms that have been started need to be scheduled, but each algorithm only needs to be scheduled once. This operation is usually used in scenarios such as resource allocation and task scheduling to ensure that all related algorithms are properly processed.

[0063] S2-4. Monitor the running status and performance indicators of the algorithm container in real time, and evaluate the startup ranking of each algorithm to be started in the set of algorithms to be started based on the memory usage, inference speed, actual throughput, amount of unidentified tasks and task waiting time of the algorithm container, and obtain a startup ranking table of the algorithms to be started.

[0064] During the specific implementation, the running status and performance indicators of the algorithm container, such as video memory occupancy, inference speed, throughput, etc., are monitored in real time to adjust the startup strategy and resource allocation of the algorithm container in a timely manner. Then, based on the collection of algorithms to be started, the startup priority ranking of each algorithm to be started is evaluated. The ranking basis includes but is not limited to the percentage of graphics card storage occupancy, inference speed, actual throughput, unidentified tasks, and task waiting time of the algorithm container. Then, by comprehensively evaluating these factors, the startup priority ranking list of the algorithms to be started is obtained.

[0065] like Figure 5 As shown, the evaluation of the startup ranking of each algorithm to be started in the set of algorithms to be started to obtain a startup ranking table of the algorithms to be started includes the following sub-steps: Get the proportion of video memory usage p1 of each algorithm to be started, calculated as follows: p1 = algorithm video memory usage / total algorithm video memory usage. Where p1 represents: the ratio of the video memory currently occupied by a certain algorithm to be started to the total video memory usage of all algorithms. Algorithm video memory usage represents: the video memory usage of the algorithm to be started, in MB or GB, etc. The total algorithm video memory usage represents: the sum of the current video memory usage of all algorithms in the set of algorithms to be started.

[0066] Get the ratio of the inference speed of each algorithm to be started, p2, calculated as follows: p2 = algorithm inference speed / total algorithm inference speed. Among them, p2 represents: the ratio (or relative speed) of the inference speed of a certain algorithm to be started to the sum of the inference speeds of all algorithms. The algorithm inference speed represents: the number of tasks or data that the algorithm to be started can process per second, and the unit can be tasks / s, images / s, etc. The total algorithm inference speed represents: the sum of the inference speeds of all algorithms in the set of algorithms to be started.

[0067] Get the actual throughput ratio p3 of each algorithm to be started, calculated as follows: p3 = actual algorithm throughput / total algorithm throughput. In addition, the calculation formula for the actual throughput of a single algorithm is as follows: for algorithms to be started, their actual throughput is 0; for algorithms that have been started, their actual throughput = the number of tasks done during the startup time period / startup duration. Among them, p3 represents: the ratio of the number of tasks actually completed by a certain algorithm to be started within a period of time to the total number of tasks actually completed by all algorithms. The actual algorithm throughput represents: the number of tasks completed by the algorithm to be started within a specified time window. The total algorithm throughput represents: the sum of the number of tasks completed by all algorithms in the same time window in the set of algorithms to be started.

[0068] Get the proportion of unrecognized tasks p4 of each algorithm to be started, calculated as follows: p4 = number of tasks to be identified by the algorithm / number of tasks to be identified by the total number of tasks to be identified by the algorithm. Among them, p4 represents: the ratio of the number of tasks to be identified by a certain algorithm to be started to the total number of tasks to be identified by all algorithms. The number of unrecognized tasks by the algorithm represents: the number of tasks currently queued for identification by the algorithm to be started. The total number of tasks to be identified by the algorithm represents: the sum of the number of tasks currently queued for identification by all algorithms in the set of algorithms to be started.

[0069] Get the task waiting time ratio p5 of each algorithm to be started, calculated as: p5 = algorithm task waiting time / total algorithm task waiting time. Among them, p5 represents: the ratio (or relative length) of the average waiting time of a task of a certain algorithm to be started to the sum of the average waiting time of all algorithm tasks. The algorithm task waiting time represents: the average waiting time for the algorithm task to be started to be processed, which can be in seconds, minutes, etc. The total algorithm task waiting time represents: the sum of the average waiting time of all algorithm tasks in the set of algorithms to be started.

[0070] Get the score of each algorithm to be started, and the calculation formula is: score = p2×a + p3×b + p4×c + p5×d - p1×e. Among them, a, b, c, d, and e refer to weights, and the specific values ​​are obtained based on experience. In addition, the settings of weights a, b, c, d, and e need to be adjusted according to the actual system conditions and business needs. For example, if the system has high requirements for real-time performance, the weight of inference speed (p2) can be set higher; if you want to balance resource usage, the weight of video memory occupancy (p1) can also be increased accordingly.

[0071] All algorithms to be started are sorted in descending order according to their scores to obtain a startup ranking list of algorithms to be started. This is a list that is sorted based on various performance indicators of the algorithm (such as the percentage of graphics card storage usage, the percentage of inference speed, the percentage of actual throughput, the percentage of unrecognized tasks, the percentage of task waiting time, etc., and the scores obtained by weighted calculation). The higher the ranking of the algorithm to be started, the better its overall performance may be, or it may be more in line with the scheduling requirements of the current system.

[0072] S2-5, such as Figure 6 As shown, starting and stopping the algorithm according to the startup ranking table includes the following sub-steps: S2-51. According to the order of the startup ranking, the algorithms to be started are taken out one by one from the startup ranking table.

[0073] In specific implementation, under limited computing resources, according to the startup ranking list of the algorithms to be started, try to start the algorithm containers with the highest ranking in turn.

[0074] S2-52. Determine whether the task to be identified corresponding to the algorithm to be started can be completed within the preset time; if so, it means that only one algorithm container needs to be started; if not, it means that multiple algorithm containers need to be started, and the number of algorithm containers that need to be started is calculated.

[0075] In specific implementation, determine whether the tasks to be identified corresponding to the algorithm to be started can be completed within the preset time. First, calculate the number of tasks that can be executed by the algorithm to be started. The calculation formula is: number of tasks that can be executed = algorithm throughput × preset time (such as 10 minutes). Among them, the algorithm throughput refers to the total number of tasks that can be processed per unit time, reflecting the overall processing efficiency; 10 minutes is the waiting time for a single cycle of the scheduler, and it can also be adjusted to other times according to actual conditions. Then determine whether the number of tasks that can be executed is greater than the total number of tasks to be identified. If it is greater, it means that it can be executed within the preset time; if it is not greater, it means that it cannot be executed within the preset time.

[0076] If the execution can be completed within the preset time, it means that the number of algorithm containers that need to be started is 1.

[0077] If the task cannot be completed within the preset time, it means that more than one algorithm container needs to be started, and the specific number of algorithm containers that need to be started is calculated. The calculation formula is: number of algorithm containers = number of tasks to be identified ÷ number of tasks that can be executed within the preset time (such as 10 minutes).

[0078] Specifically, for each algorithm to be started, the number of algorithm containers that need to be started is calculated based on its throughput capacity to ensure that all tasks to be identified corresponding to the algorithm can be identified within the next 10 minutes. In addition, the algorithm containers are subsequently started based on the number of algorithm containers that need to be started.

[0079] S2-53, based on the number of algorithm containers that need to be started, determine whether the GPU resources are sufficient. If sufficient, start an algorithm container; if not, determine whether there is an algorithm that has been started with a lower ranking than the current algorithm to be started. If so, close the algorithm container of the algorithm that has been started with a lower ranking to release GPU resources, and repeat this step to continue to determine whether the GPU resources are sufficient. If not, it means that it is impossible to close more algorithm containers corresponding to algorithms with lower rankings, so exit this algorithm container scheduling.

[0080] Specifically, when it is found that the GPU resources are insufficient to start a new algorithm container, it checks whether there are any algorithms that are ranked lower than the algorithm to be started among the started algorithms corresponding to the currently started algorithm container. Here, "low ranking" means that these started algorithms are at the bottom of the startup ranking list, which may be due to their poor overall performance, or because their current task load is light, waiting time is long, etc.

[0081] If it is found that the algorithm containers corresponding to algorithms with lower rankings have been started, in order to release GPU resources for algorithms that need them more (i.e., higher rankings), the algorithm containers corresponding to these algorithms with lower rankings will be closed. This is done to optimize resource utilization and ensure that high-performance or high-priority algorithms to be started can get enough resources to perform tasks.

[0082] After closing the algorithm container corresponding to the lower-ranked algorithm, recheck whether the GPU resources are sufficient to start a new algorithm container. If it is still insufficient, this process may be repeated until sufficient GPU resources are found or there are no more algorithm containers corresponding to the lower-ranked algorithms that can be closed. If a new algorithm container still cannot be started in the end, then exit this algorithm scheduling process.

[0083] S2-54, determine whether all started algorithms are in the first few in the startup ranking table; if so, end this algorithm container scheduling, if not, return to S2-51, and take out the algorithms to be started from the startup ranking table one by one in the order of startup ranking, until the startup ranking table is traversed, and then end this scheduling.

[0084] The purpose of this step is to check whether the currently started algorithms are all the algorithms that rank high on the startup ranking list. The "top few" here does not specify the number, but is based on whether the GPU resources are sufficient. You can set it according to the actual situation.

[0085] In the specific implementation, first check all the started algorithms and compare their rankings on the startup ranking table. If all the started algorithms are ranked very high on the ranking list (that is, they have high scores and good overall performance), and the GPU resources are insufficient to start more algorithm containers, then it is considered that the optimal algorithm combination has been started and no further scheduling is required.

[0086] If the above conditions are met, the current scheduling is terminated directly, and no attempt is made to start other algorithms to be started or adjust the current algorithm combination. If the conditions are not met (i.e., there are algorithms with higher rankings that have not been started, or although the started algorithms are ranked higher, the GPU resources are still sufficient to start more algorithms), the loop continues, traversing the algorithms to be started on the startup ranking list, and making algorithm start and stop decisions according to steps S2-51.

[0087] S2-6. After completing this scheduling, enter the sleep state and return to S2-1.

[0088] This step indicates the end of an algorithm container scheduling process, and then it will enter a dormant state (such as 10 minutes), and then restart a new algorithm container scheduling process.

[0089] In specific implementation, after completing an algorithm container scheduling (whether because the optimal algorithm combination has been started or because the startup ranking table has been traversed and all possible start and stop decisions have been made), the algorithm container scheduling process ends. In order to reduce system load, save resources, and give the system a buffer and preparation time, it enters a 10-minute dormancy period. During this period, no algorithm container scheduling or task processing is performed. After the dormancy period ends, a new algorithm container scheduling process is restarted, starting from step S2-1, re-evaluating the current task requirements, algorithm performance, GPU resources, etc., and making new scheduling decisions.

[0090] According to the above technical solution, the algorithm task scheduling method of the intelligent recognition system of the present invention realizes the centralized management of the recognition tasks and the intelligent scheduling of the algorithm container through the cooperation of the recognition request storage stage and the algorithm container scheduling stage. In the algorithm container scheduling stage, by comprehensively considering multiple factors such as the memory occupancy, inference speed, actual throughput, unrecognized task volume and task waiting time of the algorithm container, the startup ranking list of the algorithm to be started is calculated, thereby realizing the intelligent start and stop of the algorithm container under limited computing power resources, and improving the execution efficiency of the recognition task and the utilization rate of system resources.

[0091] S3, identification task execution phase: After the scheduled algorithm container is started, it automatically enters the recognition task execution phase. Specifically, the currently started algorithm container will obtain the tasks to be recognized assigned to the corresponding algorithm from the task list and mark the status of these tasks as "recognizing". Subsequently, the currently started algorithm container recognizes these tasks in recognition through the corresponding algorithm. When the algorithm container has recognized all the tasks assigned to it, it will enter the dormant state. This is to save resources and prevent the algorithm container from occupying system resources when there are no tasks.

[0092] However, if Figure 3 As shown in the figure, once an algorithm recognition request with a status of pending recognition is stored in the recognition task table, the algorithm new task notification signal will be triggered. If the algorithm container receives a new task notification signal during sleep, it will be awakened and continue to execute the task, including the following sub-steps: S3-1. When an algorithm recognition request with a status marked as pending recognition is stored in the recognition task table, the algorithm new task notification signal is triggered, and it is only triggered once within the pre-examination time to wake up the corresponding algorithm container in the dormant state.

[0093] During specific implementation, in order to avoid wasting resources caused by repeatedly triggering the task notification signal in a short period of time, the system is set to trigger a new task notification signal only once within 5 seconds (this time can be configured as needed).

[0094] S3-2, the awakened algorithm container takes the task and determines whether there is a corresponding task to be identified in the task list.

[0095] If there is no task to be identified, it is determined whether no task is obtained after three cycles (this number can be configured as needed). If not, the number of cycles is increased by 1, and the algorithm container that was awakened is returned to obtain the task, and step S3-2 is repeated; if yes, the sleep mechanism is triggered, and an algorithm sleep notification signal is sent to the algorithm container, so that it enters the sleep state and no longer enters the algorithm container scheduling stage.

[0096] If there is a task to be identified, the awakened algorithm container will change the status of the task to be identified to "identifying" after getting the task to be identified. In addition, in order to ensure data consistency, the operations of reading and getting tasks need to be processed by transactions.

[0097] In some embodiments, when taking tasks, priority is given to synchronous recognition tasks that are being recognized but have been executed for more than half an hour (this time is also configurable).

[0098] Specifically, the algorithm recognition initiated by the business system is divided into synchronous recognition tasks and asynchronous recognition tasks. Asynchronous recognition tasks do not require timeliness and notify the business system of the recognition results through callbacks. Synchronous recognition tasks are prioritized by the system because they require timeliness.

[0099] S3-3, the awakened algorithm container performs recognition processing on the task being recognized, such as Figure 7 As shown, it includes the following sub-steps: before recognition, download the image in MinIO storage and build recognition parameters. During recognition, call the corresponding algorithm through HTTP to recognize the task. After recognition, obtain the recognition result of the awakened algorithm container for the task.

[0100] S3-4, get the recognition result, and determine whether the returned recognition result is normal. If the recognition result is normal, it means that the callback recognition is successful; if the recognition result is abnormal, it means that the callback recognition fails.

[0101] In specific implementation, callback recognition refers to a mechanism in which, during the recognition process, when the recognition result is generated, the result is returned to the developer through a callback. In this way, the developer can process or display the recognition result in the callback function, realize the interaction with the recognition task, and improve the user experience.

[0102] S3-5. Determine whether the number of retries for callback recognition failure is greater than three times (this number can be configured as needed). If the number of retries exceeds three times, the record result is modified to callback recognition failure; if the number of retries does not exceed three times, the record result is modified to callback recognition success.

[0103] S3-6, end the cycle of this recognition task execution phase and enter the next cycle.

[0104] It can be seen from the above technical solution that the process of the recognition task execution phase ensures efficient, stable and reliable operation of the recognition task execution phase through reasonable task scheduling, priority processing, sleep mechanism and callback notification strategies.

[0105] In summary, the method of the present invention has the following technical effects: 1) Real-time performance monitoring technology for algorithm containers, which can accurately capture the resource usage and performance indicators of algorithm containers.

[0106] 2) Dynamic resource scheduling algorithm, which can achieve optimal resource allocation based on real-time data and historical trend predictions.

[0107] 3) Intelligent management mechanism of task queue improves task processing efficiency through priority sorting and task scheduling strategies.

[0108] 4) A multi-dimensional comprehensive evaluation system for dispatching decisions to ensure the comprehensiveness and accuracy of dispatching decisions.

[0109] 5) Real-time monitoring and optimization technology of algorithm performance to improve the execution efficiency and accuracy of recognition tasks.

[0110] Embodiment 2: The visual recognition algorithm container dynamic scheduling system of the present invention.

[0111] like Figure 8 As shown, this embodiment provides a visual recognition algorithm container dynamic scheduling system, including an algorithm container scheduling module; the algorithm container scheduling module includes an algorithm collection acquisition unit, a startup ranking evaluation unit, a container quantity calculation unit, a GPU resource release unit, and a startup ranking check unit, which are specifically introduced as follows.

[0112] Among them, the algorithm collection acquisition unit is used to: acquire the algorithm collection of the current task to be identified based on the identification task table; acquire the currently started algorithm collection based on the algorithm list; and take the union of the algorithm collection of the task to be identified and the started algorithm collection to obtain the algorithm collection to be started.

[0113] The evaluation startup ranking unit is used to: monitor the running status and performance indicators of the algorithm container in real time, and evaluate the startup ranking of each algorithm to be started in the set of algorithms to be started according to the video memory occupancy, inference speed, actual throughput, unidentified task volume and task waiting time of the algorithm container, so as to obtain a startup ranking table of the algorithms to be started.

[0114] The container quantity calculation unit is used to: take out the algorithms to be started one by one from the startup ranking table in the order of the startup ranking, and judge whether the corresponding tasks to be identified can be completed within the preset time; if so, it means that only one algorithm container needs to be started; if not, it means that multiple algorithm containers need to be started, and the number of algorithm containers that need to be started is calculated.

[0115] The GPU resource release unit is used to: determine whether the GPU resources are sufficient based on the number of algorithm containers that need to be started; if sufficient, start an algorithm container; if insufficient, determine whether there is a started algorithm with a lower ranking than the current algorithm to be started, if so, close the algorithm container of the started algorithm, and repeatedly determine whether the GPU resources are sufficient, if not, exit the current algorithm container scheduling.

[0116] The startup ranking checking unit is used to determine whether all started algorithms are in the first few positions in the startup ranking table; if so, end the algorithm container scheduling; if not, return to the order of startup ranking, and take out the algorithms to be started one by one from the startup ranking table until the startup ranking table is traversed.

[0117] In other embodiments, the system of the present invention further includes a module for identifying and requesting storage; the module for identifying and requesting storage includes a request initiating unit, a parameter checking unit, and a storage unit, which are specifically described as follows.

[0118] The request initiating unit is used for: the business system service calls the algorithm identification request, and sends the algorithm identification request including the data to be identified, the request parameters and the unique identification identifier to the intelligent identification system service.

[0119] The parameter verification unit is used to: the intelligent identification system service performs parameter verification on the algorithm identification request; if the verification is successful, the algorithm identification request is recorded in the database and marked as to be identified, and a postgres table record containing a unique identification identifier and an algorithm primary key identifier is generated; if the verification fails, the algorithm identification request is recorded in the database and marked as parameter verification failure.

[0120] The storage unit is used to store all algorithm recognition requests marked as to-be-recognized into a recognition task table to form a to-be-recognized task queue.

[0121] In other embodiments, the system of the present invention further includes an identification task execution module; the identification task execution module includes an algorithm identification processing unit, an algorithm container wake-up unit, and a task fetching unit, which are specifically described as follows.

[0122] Among them, the algorithm recognition processing unit is used to: the currently started algorithm container obtains the corresponding task to be recognized from the recognition task table, marks its status as being recognized, and performs recognition processing on it; when all tasks marked as being recognized are recognized, the algorithm container enters a dormant state.

[0123] The algorithm container awakening unit is used to trigger an algorithm new task notification signal to awaken the corresponding algorithm container in a dormant state once an algorithm recognition request with a status marked as to-be-recognized is stored in the recognition task table.

[0124] The task fetching unit is used for: the awakened algorithm container to fetch tasks, and to determine whether there are corresponding tasks to be identified in the task list; if not, to determine whether the task has not been fetched after a preset number of cycles; if not, the number of cycles is increased by 1, and the awakened algorithm container is returned to fetch tasks; if so, the algorithm is triggered to enter a sleep notification signal, so that the awakened algorithm container enters a sleep state; if so, after the awakened algorithm container fetches the task to be identified, its state is marked as being identified, and identification processing is performed on it to obtain the identification result and determine whether it is normal; if normal, it indicates that the callback identification is successful; if abnormal, it indicates that the callback identification fails; then determine whether the number of retries for the callback identification failure is greater than the preset number, if so, modify the record result to indicate that the callback identification failed; if not, modify the record result to indicate that the callback identification was successful.

[0125] In specific implementation, the system of the present invention uses different namespaces and containers in Kubernetes (K8S) to implement workflows. Fig. 9 As shown in the figure, the entire process shows the process from task generation and storage to script injection and algorithm container startup, involving the collaborative work of multiple namespaces and containers. The details are as follows.

[0126] The K8S-default-cloud namespace is a default namespace in Kubernetes, which is used to store and manage various resource objects. The ai-web pod container, ai-service intelligent computing pod container, and ai-service-algorithm- schedule- scripts algorithm scheduling pod container are Pod containers running in the K8S-default-cloud namespace, which are used for Web services, intelligent computing, and algorithm scheduling, respectively. In addition, the ai-web pod container initiates an algorithm recognition request to the ai-service intelligent computing pod container.

[0127] Redis and Postgres are basic services on the host machine for data storage. In addition, Redis is used for data storage and access for task reading and execution. The ai-service intelligent computing pod container stores all algorithm recognition requests marked as pending recognition into the Postgres recognition task table to form a queue of pending recognition tasks. Once a recognition task is detected, the ai-service-algorithm- schedule- scripts algorithm scheduling pod container obtains the pending recognition task from the Postgres recognition task table and schedules the algorithm container.

[0128] K8S-(ai-model)-(ai-model) namespace is another Kubernetes namespace dedicated to storing and managing resource objects related to ai-model. In the ai-model namespace, there are multiple algorithm containers (such as algorithm 1 pod container, algorithm 2 pod container, ..., algorithm N pod container), which are started after the script is injected to perform specific algorithm tasks. In addition, multiple algorithm containers are started after the script is injected. The ai-inject-service pod container is used to inject scripts into the PVC directory of the ai-model namespace to ensure that the algorithm container can load these scripts when it starts.

[0129] In addition, the recognition request storage module includes ‌ai-web pod container‌, ‌ai-service intelligent computing pod container‌, Redis and Postgres, and has the following functions: managing the queue of visual recognition tasks, including task reception, queuing and status update. This module uses advanced queue management algorithms to intelligently schedule tasks according to the urgency of the task and resource requirements, ensuring that high-priority tasks can be processed quickly.

[0130] The algorithm container scheduling module includes ai-service-algorithm- schedule- scripts algorithm scheduling pod container, which has the following functions: (1) Resource scheduling decision: Dynamically calculate resource allocation and replica number adjustment strategy according to the running status and task queue of the algorithm container. And use machine learning algorithm to predict resource demand and task load, and automatically adjust the number of replicas of the algorithm container to adapt to the changing resource demand. (2) Container scheduling execution: Execute the start, stop and resource adjustment operations of the algorithm container. And through the interface with the container management system, it realizes precise control of the algorithm container to ensure that resource scheduling decisions can be executed quickly and accurately. (3) Scheduling log recording: Record key information in the scheduling process for system monitoring and troubleshooting. And record detailed information of each scheduling operation, including scheduling time, resource allocation, task processing results, etc., to provide data support for system maintenance and optimization.

[0131] The recognition task execution module includes multiple algorithm containers and ai-inject-service pod containers, and has the following functions: (1) Responsible for executing visual recognition tasks and monitoring algorithm performance in real time. This module integrates a high-performance visual recognition algorithm and can monitor the algorithm's recognition speed, accuracy and other performance indicators in real time while executing tasks to ensure efficient completion of recognition tasks.

[0132] Embodiment 3: The electronic device of the present invention.

[0133] like Fig.10 As shown, this embodiment provides an electronic device, including a processor, and a memory coupled to the processor; wherein the memory has computer program instructions stored therein; when the computer program instructions are executed by the processor, the electronic device executes a visual recognition algorithm container dynamic scheduling method as described in Example 1.

[0134] Specifically, the electronic device includes a central processing unit (CPU) and / or a graphics processing unit (GPU), which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) or loaded from a storage unit into a random access memory (RAM). In RAM, various programs and data required for device operation can also be stored. CPU / GPU, ROM and RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus. In addition, the electronic device may also include a coprocessor.

[0135] Furthermore, a number of components in an electronic device are connected to the I / O interface, including: input units, such as keyboards, mice, etc.; output units, such as various types of displays, speakers, etc.; storage units, such as disks, optical disks, etc.; and communication units, such as network cards, modems, wireless communication transceivers, etc. The communication unit allows the device to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunication networks.

[0136] The various methods or processes described above may be performed by a CPU / GPU. For example, in some embodiments, the method may be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program may be loaded and / or installed on a device via a ROM and / or a communication unit. When the computer program is loaded into RAM and executed by the CPU / GPU, one or more steps or actions in the method or process described above may be performed.

[0137] Embodiment 4: Computer readable storage medium of the present invention.

[0138] This embodiment provides a computer-readable storage medium, including computer program instructions stored therein, and when the computer program instructions are executed, the method for dynamic scheduling of visual recognition algorithm containers described in Example 1 is implemented.

[0139] Specifically, a computer-readable storage medium can be a tangible device that can hold and store computer program instructions used by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, for example, a punch card or a convex structure in a groove on which a computer program instruction is stored, and any suitable combination thereof. The computer-readable storage medium used here is not interpreted as a transient signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated by a waveguide or other transmission medium (e.g., a light pulse by an optical fiber cable), or an electrical signal transmitted by a wire.

[0140] The computer program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer program instructions from the network and forwards the computer program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0141] Furthermore, the computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages, and conventional procedural programming languages. The computer program instructions may be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, by using the state information of the computer program instructions to personalize an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit may execute the computer program instructions, thereby implementing various aspects of the present disclosure.

[0142] These computer program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer program instructions can also be stored in a computer-readable storage medium, and these instructions make the computer, programmable data processing device and / or other equipment work in a specific way, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0143] In addition, computer program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operating steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0144] Moreover, the flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the methods, systems and devices according to the multiple embodiments of the present disclosure. In this regard, each frame in the flow chart or block diagram can represent a part of a module, a program segment or an instruction, and a part of the module, a program segment or an instruction includes one or more executable instructions for realizing the specified logical function. In some alternative implementations, the functions marked in the frame can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous frames can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each frame in the block diagram and / or flow chart, and the combination of frames in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0145] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A method for dynamically scheduling a visual recognition algorithm container, characterized in that: Includes the algorithm container scheduling stage; The algorithm container scheduling phase includes the following steps: Based on the recognition task table, obtain the algorithm collection corresponding to the current task to be recognized; based on the algorithm list, obtain the currently started algorithm collection; take the union of the algorithm collection corresponding to the task to be recognized and the started algorithm collection to obtain the algorithm collection to be started; Monitor the running status and performance indicators of the algorithm container in real time, and evaluate the startup ranking of each algorithm to be started in the set of algorithms to be started based on the memory usage, inference speed, actual throughput, number of unidentified tasks, and task waiting time of the algorithm container, and obtain a startup ranking table of the algorithms to be started; According to the order of the startup ranking, the algorithms to be started are taken out one by one from the startup ranking table to determine whether the corresponding tasks to be identified can be completed within the preset time; if yes, it means that only one algorithm container needs to be started; if not, it means that multiple algorithm containers need to be started, and the number of algorithm containers that need to be started is calculated; Based on the number of algorithm containers that need to be started, determine whether GPU resources are sufficient; if so, start an algorithm container; if not, determine whether there is an algorithm that has been started with a lower ranking than the current algorithm to be started. If so, close the algorithm container of the started algorithm and repeat the determination of whether GPU resources are sufficient. If not, exit the current algorithm container scheduling; Determine whether all started algorithms are in the first few in the startup ranking table; if so, end the algorithm container scheduling; if not, return to the order of startup ranking, and take out the algorithms to be started one by one from the startup ranking table until the startup ranking table is traversed.

2. According to claim 1, a method for dynamically scheduling a visual recognition algorithm container is characterized in that: The step of evaluating the startup ranking of each algorithm to be started in the set of algorithms to be started to obtain a startup ranking table of the algorithms to be started includes the following sub-steps: Get the proportion of video memory usage p1, inference speed p2, actual throughput p3, unrecognized task volume p4, and task waiting time p5 of each algorithm to be started; Among them, the proportion of video memory usage is p1, and the calculation formula is: p1 = algorithm video memory usage / total algorithm video memory usage; The reasoning speed accounts for p2, and the calculation formula is: p2 = algorithm reasoning speed / total algorithm reasoning speed; The actual throughput accounts for p3, calculated as follows: p3 = actual algorithm throughput / total algorithm throughput; The proportion of tasks to be identified is p4, and the calculation formula is: p4 = number of tasks to be identified by the algorithm / total number of tasks to be identified; The task waiting time accounts for p5, and the calculation formula is: p5 = algorithm task waiting time / total algorithm task waiting time; According to the preset weights a, b, c, d, e, the score of each algorithm to be started is calculated. The calculation formula is: score = p2×a+p3×b+p4×c+p5×d-p1×e; All algorithms to be started are arranged in descending order according to their scores to obtain a startup ranking list of the algorithms to be started.

3. The method for dynamically scheduling a visual recognition algorithm container according to claim 1, characterized in that: The step of judging whether the corresponding task to be identified can be completed within the preset time includes the following sub-steps: Calculate the number of tasks that can be completed by the algorithm to be started. The calculation formula is: number of tasks that can be completed = algorithm throughput × preset time; Determine whether the number of tasks that can be executed is greater than the total number of tasks to be identified; If yes, it means that the task can be completed within the preset time; if no, it means that the task cannot be completed within the preset time.

4. The method for dynamically scheduling a visual recognition algorithm container according to claim 1, characterized in that: The calculation is performed to determine the number of algorithm containers that need to be started, and the calculation formula is: number of algorithm containers = number of tasks to be identified currently / number of tasks that can be executed within a preset time.

5. A method for dynamically scheduling visual recognition algorithm containers according to any one of claims 1 to 4, characterized in that: Before the algorithm container scheduling stage, there is also an identification request storage stage; the identification request storage stage includes the following steps: The business system service calls the algorithm identification request and sends the algorithm identification request including the data to be identified, request parameters and unique identification identifier to the intelligent identification system service; The intelligent identification system service performs parameter verification on the algorithm identification request; if the verification succeeds, the algorithm identification request is recorded in the database, and its status is marked as pending identification, and a postgres table record containing a unique identification identifier and an algorithm primary key identifier is generated; if the verification fails, the algorithm identification request is recorded in the database, and its status is marked as parameter verification failure; All algorithm recognition requests marked as pending recognition are stored in the recognition task table to form a pending recognition task queue.

6. A method for dynamically scheduling visual recognition algorithm containers according to any one of claims 1 to 4, characterized in that: After the algorithm container scheduling phase, an identification task execution phase is also included; the identification task execution phase includes the following steps: The currently started algorithm container obtains the corresponding task to be identified from the identification task table, marks its status as being identified, and performs identification processing on it; When all tasks marked as being identified are identified, the algorithm container enters a dormant state.

7. A method for dynamically scheduling visual recognition algorithm containers according to claim 6, characterized in that: The identification task execution phase also includes the following steps: Once an algorithm recognition request with a status marked as pending recognition is stored in the recognition task table, the algorithm new task notification signal is triggered to wake up the corresponding algorithm container in the dormant state; The awakened algorithm container takes the task and determines whether there is a corresponding task to be identified in the task list; If not, determine whether the task is not obtained after the preset number of cycles; if not, increase the number of cycles by 1 and return to the awakened algorithm container to obtain the task; if so, trigger the algorithm to enter the sleep notification signal, so that the awakened algorithm container enters the sleep state; If yes, the awakened algorithm container will take the task to be identified, mark its status as being identified, perform identification processing on it, get the identification result and judge whether it is normal; if normal, it means that the callback identification is successful; if abnormal, it means that the callback identification fails; then judge whether the number of retries of callback identification failure is greater than the preset number, if so, modify the record result to callback identification failure; if not, modify the record result to callback identification success.

8. A visual recognition algorithm container dynamic scheduling system, characterized in that: It includes an algorithm container scheduling module; the algorithm container scheduling module includes an algorithm collection acquisition unit, a startup ranking evaluation unit, a container quantity calculation unit, a GPU resource release unit, and a startup ranking check unit; The algorithm collection acquisition unit is used to: acquire the algorithm collection corresponding to the current task to be identified based on the identification task table; acquire the currently started algorithm collection based on the algorithm list; and obtain the algorithm collection corresponding to the task to be identified and the started algorithm collection by taking the union of the algorithm collection corresponding to the task to be identified and the started algorithm collection to obtain the algorithm collection to be started; The evaluation startup ranking unit is used to: monitor the running status and performance indicators of the algorithm container in real time, and evaluate the startup ranking of each algorithm to be started in the set of algorithms to be started according to the video memory occupancy, inference speed, actual throughput, unrecognized task amount and task waiting time of the algorithm container, so as to obtain a startup ranking table of the algorithms to be started; The container quantity calculation unit is used to: take out the algorithms to be started one by one from the startup ranking table according to the startup ranking order, and determine whether the corresponding tasks to be identified can be completed within the preset time; If yes, it means that only one algorithm container needs to be started; if no, it means that multiple algorithm containers need to be started and the number of algorithm containers that need to be started needs to be calculated; The GPU resource release unit is used to: determine whether the GPU resources are sufficient based on the number of algorithm containers that need to be started; if sufficient, start an algorithm container; if insufficient, determine whether there is a started algorithm with a lower ranking than the current algorithm to be started, if yes, close the algorithm container of the started algorithm, and repeatedly determine whether the GPU resources are sufficient, if no, exit the current algorithm container scheduling; The startup ranking checking unit is used to determine whether all started algorithms are in the first few positions in the startup ranking table; if so, end the algorithm container scheduling; if not, return to the order of startup ranking, and take out the algorithms to be started one by one from the startup ranking table until the startup ranking table is traversed.

9. An electronic device, characterized in that: It includes a processor and a memory coupled to the processor; the memory has computer program instructions stored therein; when the computer program instructions are executed by the processor, the electronic device executes a visual recognition algorithm container dynamic scheduling method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The invention comprises computer program instructions stored therein; when the computer program instructions are executed, the method for dynamically scheduling a visual recognition algorithm container as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Artificial intelligence computer vision reasoning method

    CN114298313A

  • GPU resource scheduling method and system based on dynamic weight calculation

    CN119415245A

  • Computing power scheduling method and system based on containerization technology

    CN119862003A

Cited By

  • Process task starting control method and semiconductor cleaning equipment

    CN120821248A