Method for artificial intelligence calculation and calculation equipment
By reconstructing the AI algorithm task set as a converged task, sharing model calculation and business calculation, and optimizing the task execution order through directed acyclic graph, the development difficulty and resource reuse difficulties caused by AI chip diversity in the existing technology are solved, and efficient AI inference performance is achieved.
Patent Information
- Application Number
- CN202411884159.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-05-16
AI Technical Summary
There are many types of existing AI chips. Developing AI algorithms with superior performance requires familiarity with the characteristics and runtime of different chips, which makes it difficult to develop and requires a lot of manpower investment. Different algorithms may call the same model, resulting in difficulty in resource reuse.
By obtaining the algorithm task set, reconstructing it into fusion tasks, sharing model calculation and business calculation, scheduling computing resources to execute computing tasks, using the data structure provided by the predetermined framework to realize business calculation, and optimizing the task execution order through directed acyclic graphs, realizing resource reuse and performance optimization.
It blocks the differences in the underlying software and hardware, reduces duplicate development work, realizes resource reuse, standardizes algorithm code writing, and ensures the ultimate performance of AI inference.
Smart Images

Figure CN120011006A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and computing device for artificial intelligence computing. Background Art
[0002] As artificial intelligence technology continues to develop, machine learning and deep learning technologies continue to mature, and various AI chips dedicated to deep learning computing have been designed. They can be applied to various production and life scenarios, from health care to autonomous driving, from smart life to smart cities, AI chips play an important role in them.
[0003] There are many types of AI chips on the market. Each chip manufacturer will launch its own runtime for model reasoning. At the same time, different chips have different characteristics. For AI algorithm application development engineers, if they want to develop an AI algorithm with good enough performance, they must be familiar with the chip characteristics and how to use different runtimes. The requirements for algorithm engineers are very high. At the same time, if they need to connect to different AI chips, a lot of manpower investment is required, and an algorithm application requires multiple models to complete. Different algorithms may call the same model at the bottom.
[0004] Therefore, a technical solution is needed to shield the differences between the underlying software and hardware, reduce duplication of development work, achieve resource reuse, standardize algorithm code writing, and ensure the ultimate performance of AI reasoning. Summary of the invention
[0005] The present invention aims to provide a method and computing device for artificial intelligence computing, which can shield the differences between the underlying software and hardware, reduce repeated development work, realize resource reuse, standardize the algorithm code writing, and ensure the ultimate performance of AI reasoning.
[0006] According to one aspect of the present invention, a method for artificial intelligence computing is provided, the method comprising:
[0007] Acquire an algorithm task set for an artificial intelligence application, the algorithm task set comprising a plurality of algorithm tasks, each algorithm task comprising a plurality of model calculations and a plurality of business calculations related to the plurality of model calculations;
[0008] Reconstructing the algorithm task set into a fusion task, in which at least one model calculation is shared by a first algorithm task group and at least one business calculation is shared by a second algorithm task group, the first algorithm task group and the second algorithm task group each include at least two algorithm tasks, and the first algorithm task group and the second algorithm task group are the same or different;
[0009] The computing resources are scheduled to execute computing tasks according to the fusion task, and the computing resources include a central processing unit, memory, and artificial intelligence computing power.
[0010] According to some embodiments, the plurality of business calculations are implemented using a data structure provided by a predetermined framework, wherein the data structure includes an input data structure and an output data structure.
[0011] According to some embodiments, the input data structure includes a result threshold, an image focus area and / or an input source, and the output data structure includes an alarm identifier, an analysis category and / or a detection box.
[0012] According to some embodiments, the algorithm task set is restructured into a fusion task, including:
[0013] The multiple model calculations are constructed as nodes, the multiple business calculations are constructed as edges, and the nodes and edges form a directed acyclic graph. The same model calculations in the algorithm task set correspond to the same node in the directed acyclic graph, and the same business calculations in the algorithm task set correspond to the same edge in the directed acyclic graph.
[0014] According to some embodiments, reconstructing the algorithm task set into a fusion task further includes:
[0015] Graph optimization is performed on the directed acyclic graph.
[0016] According to some embodiments, scheduling computing resources to perform computing tasks according to the fusion task includes:
[0017] The data structure provided by the underlying platform is called to perform model calculations, and the data structure provided by the underlying platform is transparent to the developer.
[0018] According to some embodiments, scheduling computing resources to perform computing tasks according to the fusion task includes:
[0019] Based on the predetermined framework, a unified thread pool is used to schedule CPU resources to execute the multiple business calculations.
[0020] According to some embodiments, scheduling computing resources to perform computing tasks according to the fusion task further includes:
[0021] Based on the computing priority configuration, computing resources are allocated preferentially to model computing and business computing with high priority.
[0022] According to another aspect of the present invention, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the method as described above is implemented.
[0023] According to another aspect of the present invention, there is provided a computing device, comprising:
[0024] Processor; and
[0025] A memory stores a computer program, and when the computer program is executed by the processor, the method described in any one of the above items is implemented.
[0026] According to an embodiment of the present invention, an algorithm task set for artificial intelligence applications is first obtained, each algorithm task includes multiple model calculations and multiple business calculations related to the multiple model calculations, and the algorithm task set is reconstructed into a fusion task, in which at least one model calculation is shared by the first algorithm task group, and at least one business calculation is shared by the second algorithm task group, and finally the computing resources are scheduled according to the fusion task to execute the computing task. The present invention schedules computing resources according to the fusion task, can realize resource reuse, and ensure the performance of AI reasoning.
[0027] According to some embodiments, the dependencies between computing tasks can be identified through graph analysis, thereby optimizing the execution order of tasks, reducing unnecessary data transmission and calculation, and achieving efficient use of resources. The graph structure allows easy identification of which tasks can be executed in parallel, which is critical for fully utilizing the computing power of multi-core processors or multiple machines, and can greatly shorten the overall execution time of tasks.
[0028] According to some embodiments, the allocation and priority of tasks are dynamically adjusted according to the system load to ensure that high-priority or resource-intensive tasks can be processed in a timely manner, thereby improving the system's response speed and service quality. When new computing nodes or algorithm services need to be added, this can be easily achieved by adding corresponding nodes or edges to the graph without the need for large-scale reconstruction of the entire system.
[0029] According to some embodiments, the graph structure helps to quickly locate failed tasks and reschedule the execution of these tasks, or find alternative paths to complete the calculation, thereby improving the reliability and fault tolerance of the system. Displaying complex calculation processes in the form of graphs makes it easier for developers and operation and maintenance personnel to understand and monitor the operating status of the entire system, which helps to quickly diagnose and solve problems.
[0030] According to some embodiments, through in-depth analysis of the graph, bottlenecks can be discovered, providing clear directions for performance tuning, such as improving overall efficiency by adjusting the configuration of specific nodes.
[0031] It is to be understood that the foregoing general description and the following detailed description are exemplary only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for describing the embodiments are briefly introduced below.
[0033] Figure 1 A flow chart of a method for artificial intelligence computing according to an example embodiment is shown.
[0034] Figure 2 A schematic diagram showing a directed acyclic graph consisting of algorithmic business processing and model calculation according to an example embodiment.
[0035] Figure 3 A block diagram of a computing device is shown according to an exemplary embodiment. DETAILED DESCRIPTION
[0036] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that the present invention will be comprehensive and complete and fully convey the concepts of the example embodiments to those skilled in the art. The same reference numerals in the figures represent the same or similar parts, and thus their repeated description will be omitted.
[0037] In addition, the described features, structures or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present invention. However, those skilled in the art will appreciate that the technical solution of the present invention can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, known methods, devices, implementations or operations are not shown or described in detail to avoid blurring various aspects of the present invention.
[0038] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0039] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.
[0040] It should be understood that although the terms first, second, third, etc. may be used herein to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another component. Therefore, the first component discussed below can be referred to as the second component without departing from the teachings of the present inventive concept. As used herein, the term "and / or" includes any one of the associated listed items and all combinations of one or more.
[0041] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the present invention are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0042] Those skilled in the art will appreciate that the drawings are merely schematic diagrams of example embodiments, and the modules or processes in the drawings are not necessarily necessary for implementing the present invention, and therefore cannot be used to limit the protection scope of the present invention.
[0043] There are many types of AI chips on the market. Each chip manufacturer will launch its own runtime for model reasoning. At the same time, different chips have different characteristics. For AI algorithm application development engineers, if they want to develop an AI algorithm with good enough performance, they must be familiar with the chip characteristics and how to use different runtimes. The requirements for algorithm engineers are very high. At the same time, if they need to connect to different AI chips, a lot of manpower investment is required.
[0044] An algorithm application generally consists of two parts: model reasoning and business processing. Model reasoning is further divided into data preparation, reasoning, and data processing. Data preparation is to process the input of the algorithm accordingly, such as normalization and standardization, and data processing filters out useless outputs of the model. An algorithm application requires multiple models to complete. Different algorithms may call the same model at the bottom layer. The relationship between AI algorithms and AI models is many-to-many. Visual processing algorithms such as pedestrian intrusion and work clothes recognition use the same head detection model for reasoning, but there will be differences in business processing. Therefore, the same model resources can be reused during the operation of multiple algorithms.
[0045] To this end, the present invention proposes a method for artificial intelligence computing that can shield the differences between the underlying software and hardware, reduce repeated development work, achieve resource reuse, standardize algorithm code writing, and ensure the ultimate performance of AI reasoning.
[0046] Exemplary embodiments of the present invention are described below with reference to the accompanying drawings.
[0047] Figure 1 A flow chart of a method for artificial intelligence computing according to an example embodiment is shown.
[0048] See also Figure 1 In S101, an algorithm task set for artificial intelligence application is obtained, wherein the algorithm task set includes multiple algorithm tasks, each algorithm task includes multiple model calculations and multiple business calculations related to the multiple model calculations.
[0049] According to some embodiments, the multiple business calculations are implemented using a data structure provided by a predetermined framework, the data structure including an input data structure and an output data structure. The input data structure includes a result threshold, a picture focus area and / or an input source, and the output data structure includes an alarm identifier, an analysis category and / or a detection box.
[0050] The framework provides a set of business-related data structures at the algorithm level. The input data includes result threshold, image ROI, input source, etc., which are used to control the algorithm detection effect and the detection area for the input. The threshold can be understood as the accuracy of the algorithm. The smaller the threshold, the better the algorithm effect. Image ROI refers to the detection area. For image algorithms, such as detecting whether a helmet is worn at the factory gate, for the area captured by the entire camera, it is only necessary to identify the gate. At this time, it is necessary to draw an ROI box to circle the area that needs to be detected. The input source is the input data, such as a picture.
[0051] The output data includes information such as whether to alarm, analysis category, detection frame, etc., which is used for the upper-level platform to further process the algorithm results. The upper-level software platform is used to configure the combination of cameras and algorithms, and to count and summarize the result information output by the algorithm and display it on the interface.
[0052] In S103, the algorithm task set is reconstructed into a fusion task, in which at least one model calculation is shared by a first algorithm task group and at least one business calculation is shared by a second algorithm task group, the first algorithm task group and the second algorithm task group each include at least two algorithm tasks, and the first algorithm task group and the second algorithm task group are the same or different.
[0053] According to some embodiments, the multiple model calculations are constructed as nodes, the multiple business calculations are constructed as edges, and the nodes and edges are combined into a directed acyclic graph. The same model calculations in the algorithm task set correspond to the same node in the directed acyclic graph, and the same business calculations in the algorithm task set correspond to the same edge in the directed acyclic graph.
[0054] According to some embodiments, the directed acyclic graph is optimized. The graph optimization is to cache the first input of the same input passing through the same node, so that only one calculation is required and subsequent inputs can be directly reused.
[0055] According to some embodiments, a data structure provided by an underlying platform is called to perform model calculations, and the data structure provided by the underlying platform is transparent to the developer.
[0056] At the model level, for different docking chips and runtimes, the data structures provided by different chip manufacturers and different runtimes are used for model reasoning. The data structures are used within the framework and are invisible to the developers of intelligent (AI) algorithms. They cannot see the data structures provided by the chip manufacturers, but only the data structures provided by the reasoning framework.
[0057] For AI algorithm developers, they need to write two parts of code, namely model data preparation and model data processing, and algorithm business data preparation and algorithm business data processing. In order to save more resources, each model only corresponds to one data preparation and data processing, and the algorithm business is only related to the business. Relatively speaking, the code logic is relatively simple and unique to each algorithm. In this way, AI algorithm developers can be divided into two parts, algorithm business developers and model developers, making the process of developing algorithm business more efficient.
[0058] In S105, computing resources are scheduled to execute computing tasks according to the fusion task, and the computing resources include a central processing unit, memory, and artificial intelligence computing power.
[0059] According to some embodiments, a unified thread pool is used to schedule CPU resources based on the predetermined framework to execute the multiple business calculations. According to the calculation priority configuration, the computing resources are preferentially allocated to the model calculations and business calculations with high priority. For example, each node has a waiting queue, and the computing requests with high priority can jump the queue to realize the priority allocation of resources.
[0060] The framework is responsible for coordinating hardware resources such as CPU, memory, and AI computing power, and converts model reasoning into nodes and algorithm business processing into edges to form a directed acyclic graph. The bottom layer uses the same thread pool to perform business calculations. The bottom layer refers to the inside of the framework. For many algorithms running on the framework, a unified thread pool is used to schedule CPU resources.
[0061] The same model only exists in one copy in the device, which takes up less memory. Different algorithms may use the same model. For example, the yolov model is used for algorithms related to human heads. If the framework is not used, each algorithm will occupy a share of computing resources. The framework can reuse resources at the bottom layer. After the same input source is processed by different businesses, the reasoning of the model with the same input is reduced to save computing resources. The framework can also perform load balancing, so that when computing resources are insufficient, more resources can be allocated to algorithms with high priority.
[0062] For example, smoking and making phone calls are not allowed in gas stations. A camera is used for monitoring, and two different algorithms are used in the background to identify smoking and making phone calls. These two algorithms use the same human body detection model. For each input from the camera, these two algorithms need to be run. The input of the smoking algorithm is first inferred by the human body detection model. At this time, the output can be cached and returned directly after the input of the phone call algorithm arrives.
[0063] The framework described in the present invention mainly solves the problems of high requirements for AI algorithm engineers, high development manpower investment costs, and reuse of the same model resources caused by multiple algorithms. The framework of the present invention supports a variety of different AI chips, shielding the differences in underlying software and hardware, greatly reducing repetitive development work. Resource reuse, only one copy of the underlying model resources is started. Standardize the code writing methods for algorithm business processing, model data preparation, model data processing, etc., so that algorithm developers only need to focus on algorithm business processing, and the framework will perform resource scheduling at the same time to ensure the ultimate performance of AI reasoning.
[0064] Through the present invention, AI algorithm development engineers do not need to spend time researching different AI chips. The framework reasonably allocates hardware resources, coordinates the overall CPU, memory, and computing power resources, and ensures that the AI chip is never idle by using a unified thread pool internally. The framework supports allocating resources to high-priority algorithms when hardware resources are insufficient, thereby maximizing the performance of the AI chip.
[0065] Figure 2 A schematic diagram showing a directed acyclic graph consisting of algorithmic business processing and model calculation according to an example embodiment.
[0066] See also Figure 2 , Figure 2 The directed acyclic graph consisting of algorithm business processing and model calculation is shown. The framework uses algorithm business processing as edges and model calculation as nodes to build a flowchart during operation. After graph analysis and graph optimization, scheduling and reasoning are performed in a way that saves more computing resources.
[0067] At the same time, we load balance the computing resources to ensure that the resources of the algorithms with high priority are allocated when the computing resources of the device are insufficient. The computing resources on a device are fixed. A device may be configured with more than ten or twenty algorithms. The computing resources required by the algorithms in different scenarios are different. For example, the computing resources required for the smoking and phone calls algorithms in the gas station scene are much greater than the computing resources required for one person when there are ten people in the picture. There may be times when the computing resources are insufficient. At this time, the framework will allocate resources to the algorithms with high priority according to the priority of the algorithm configuration.
[0068] Graph analysis and graph optimization are to find the same model nodes and edges under different paths, and cache the output of model nodes for a short period of time to save resources. It is mainly to cache the inference results of model nodes for a short period of time. It can be used to directly output the same input.
[0069] According to some embodiments, by analyzing the edges in the graph, i.e., the dependencies between algorithmic business processes, it is possible to clearly understand which tasks must be completed before other tasks. This helps determine the correct execution order and avoid errors or repeated calculations caused by unsatisfied dependencies. Further analysis can identify whether there are redundant calculations or data transfers in the graph, such as repeated data read operations, which can be optimized or deleted to reduce resource consumption.
[0070] According to some embodiments, once the dependencies between tasks are clear, it is possible to identify which tasks can be executed simultaneously. Assigning tasks to different computing resources, such as CPU cores, GPUs, etc., for parallel processing can significantly reduce the total execution time. Based on real-time system load information, dynamically adjust the allocation and execution order of tasks to ensure that computing resources are used most effectively, such as giving priority to tasks that have a greater impact on system performance. Avoid situations where some computing nodes are overloaded while other nodes are idle, achieve load balancing through reasonable task allocation, and improve overall computing efficiency.
[0071] According to some embodiments, graph analysis can better understand data access patterns, use spatial locality and temporal locality principles to optimize data storage and loading methods, reduce unnecessary memory access, and improve cache hit rates. Frequently used intermediate results can be retained in memory for a period of time to reduce repeated calculations and data transmission. Figure 3 A block diagram of a computing device is shown according to an exemplary embodiment.
[0072] like Figure 3 As shown, computing device 30 includes processor 12 and memory 14. Computing device 30 may also include bus 22, network interface 16, and I / O interface 18. Processor 12, memory 14, network interface 16, and I / O interface 18 may communicate with each other via bus 22.
[0073] The processor 12 may include one or more general-purpose CPUs (Central Processing Units, processors), microprocessors, or application-specific integrated circuits, etc., for executing relevant program instructions. According to some embodiments, the computing device 30 may also include a high-performance graphics card (GPU) 20 for accelerating the processor 12.
[0074] The memory 14 may include a machine system readable medium in the form of a volatile memory, such as a random access memory (RAM), a read-only memory (ROM) and / or a cache memory. The memory 14 is used to store one or more programs including instructions and data. The processor 12 can read the instructions stored in the memory 14 to execute the above-mentioned method according to the embodiment of the present invention.
[0075] The computing device 30 may also communicate with one or more networks via the network interface 16. The network interface 16 may be a wireless network interface.
[0076] The bus 22 may include an address bus, a data bus, a control bus, etc. The bus 22 provides a path for exchanging information between components.
[0077] It should be noted that, in the specific implementation process, the computing device 30 may also include other components necessary for normal operation. In addition, those skilled in the art may understand that the above device may only include components necessary for implementing the embodiments of this specification, and need not include all components shown in the figure.
[0078] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above method when executed by a processor. The computer-readable storage medium may include, but is not limited to, any type of disk, including a floppy disk, an optical disk, a DVD, a CD-ROM, a microdrive, and a magneto-optical disk, a ROM, a RAM, an EPROM, an EEPROM, a DRAM, a VRAM, a flash memory device, a magnetic card or an optical card, a nanosystem (including a molecular memory IC), a network storage device, a cloud storage device, or any type of medium or device suitable for storing instructions and / or data.
[0079] An embodiment of the present invention further provides a computer program product, which includes a computer program. The computer program is operable to enable a computer to execute part or all of the steps of any one of the methods described in the above method embodiments.
[0080] Those skilled in the art can clearly understand that the technical solution of the present invention can be implemented with the help of software and / or hardware. "Unit" and "module" in this specification refer to software and / or hardware that can independently complete or cooperate with other components to complete specific functions, where the hardware can be, for example, a field programmable gate array, an integrated circuit, etc.
[0081] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.
[0082] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0083] In the several embodiments provided by the present invention, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are only schematic, such as the division of units, which is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.
[0084] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0085] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0086] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the whole or part of the technical solution can be embodied in the form of a software product, which is stored in a memory and includes several instructions for a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the methods of various embodiments of the present invention.
[0087] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0088] The exemplary embodiments of the present invention are specifically shown and described above. It should be understood that the present invention is not limited to the detailed structures, configurations or implementations described herein; on the contrary, the present invention is intended to cover various modifications and equivalent configurations included in the spirit and scope of the appended clauses.
Claims
1. A method for artificial intelligence computing, characterized in that: include: Acquire an algorithm task set for an artificial intelligence application, the algorithm task set comprising a plurality of algorithm tasks, each algorithm task comprising a plurality of model calculations and a plurality of business calculations related to the plurality of model calculations; Reconstructing the algorithm task set into a fusion task, in which at least one model calculation is shared by a first algorithm task group and at least one business calculation is shared by a second algorithm task group, the first algorithm task group and the second algorithm task group each include at least two algorithm tasks, and the first algorithm task group and the second algorithm task group are the same or different; The computing resources are scheduled to execute computing tasks according to the fusion task, and the computing resources include a central processing unit, memory, and artificial intelligence computing power.
2. The method according to claim 1, characterized in that The multiple business calculations are implemented using a data structure provided by a predetermined framework, where the data structure includes an input data structure and an output data structure.
3. The method according to claim 2, characterized in that The input data structure includes a result threshold, a picture focus area and / or an input source, and the output data structure includes an alarm mark, an analysis category and / or a detection box.
4. The method according to claim 1, characterized in that: Reconstruct the algorithm task set into a fusion task, including: The multiple model calculations are constructed as nodes, the multiple business calculations are constructed as edges, and the nodes and edges form a directed acyclic graph. The same model calculations in the algorithm task set correspond to the same node in the directed acyclic graph, and the same business calculations in the algorithm task set correspond to the same edge in the directed acyclic graph.
5. The method according to claim 4, characterized in that Reconstructing the algorithm task set into a fusion task also includes: Graph optimization is performed on the directed acyclic graph.
6. The method according to claim 4, characterized in that Scheduling computing resources to execute computing tasks according to the fusion task includes: The data structure provided by the underlying platform is called to perform model calculations, and the data structure provided by the underlying platform is transparent to the developer.
7. The method according to claim 2, characterized in that Scheduling computing resources to execute computing tasks according to the fusion task includes: Based on the predetermined framework, a unified thread pool is used to schedule CPU resources to execute the multiple business calculations.
8. The method according to claim 6 or 7, characterized in that: Scheduling computing resources to execute computing tasks according to the fusion task also includes: Based on the computing priority configuration, computing resources are allocated preferentially to model computing and business computing with high priority.
9. A computer program product, characterized in that The method comprises a computer program, which implements the method according to any one of claims 1 to 8 when being executed by a processor.
10. A computing device, characterized in that: include: processor; as well as A memory storing a computer program, wherein when the computer program is executed by the processor, the method according to any one of claims 1 to 8 is implemented.