Methods and systems for scheduling heterogeneous computing resources in intelligent mobile devices under multi-task concurrency
By parsing the concurrent learning task set of the learning machine terminal and constructing a heterogeneous computing resource visualization, combined with consortium scheduling analysis, the problems of high response latency and uneven resource utilization of critical tasks in online education were solved, achieving efficient resource utilization and improved response speed.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HAODUO SUJIAO (ZHEJIANG) NETWORK TECH CO LTD
- Filing Date
- 2026-06-29
- Publication Date
- 2026-07-31
AI Technical Summary
In online education, existing technologies for intelligent learning terminals often result in high latency in critical task response and uneven resource utilization due to a lack of comprehensive perception of task computation characteristics, real-time requirements, and node load status.
A resource demand mechanism is introduced to analyze the concurrent learning task set of the learning machine terminal, construct a heterogeneous computing resource visualization, and optimize task distribution through alliance scheduling analysis to achieve the optimal distribution and execution of collaborative computing groups.
By constructing a heterogeneous resource visualization and an alliance-based collaborative scheduling mechanism, efficient resource utilization under multi-task concurrency of the learning machine terminal is achieved, improving response speed and optimizing terminal energy consumption.
Smart Images

Figure CN122489294A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of resource scheduling technology, specifically to a method and system for scheduling heterogeneous computing resources in intelligent mobile terminals under multi-task concurrency. Background Technology
[0002] In online education, existing smart learning machines, tablets and mobile terminals need to simultaneously handle multiple learning tasks such as voice interaction, video playback, AI recognition and homework correction. To improve the computing efficiency of the devices, existing technologies usually adopt scheduling methods such as task arrival order or simple priority queues to allocate tasks to local CPUs or external cloud resources for execution in sequence.
[0003] However, in actual scheduling, due to the lack of comprehensive perception of task computation characteristics, real-time requirements and node load status, the existing scheduling method is difficult to achieve collaborative optimization of multi-dimensional resources. It is easy to cause mixed scheduling of high real-time tasks and ordinary tasks, which leads to defects such as increased response delay of key learning tasks, overload of local computing resources and unbalanced overall resource utilization, affecting the real-time performance of online learning interaction and the overall operating efficiency of the system.
[0004] In summary, existing technologies suffer from the technical problems of scheduling heterogeneous resources by learning machine terminals according to the order of task arrival, resulting in high response latency for critical tasks and uneven resource utilization. Summary of the Invention
[0005] The purpose of this application is to provide a method and system for scheduling heterogeneous computing resources in intelligent mobile terminals under multi-task concurrency, in order to solve the technical problems of existing technologies, such as learning machine terminals scheduling heterogeneous resources according to the order of task arrival, resulting in high response latency of critical tasks and uneven resource utilization.
[0006] To achieve the above objectives, this application provides a method and system for scheduling heterogeneous computing resources in intelligent mobile terminals under multi-task concurrency.
[0007] Firstly, this application provides a method for scheduling heterogeneous computing resources in a smart mobile terminal under multi-task concurrency. The method includes: introducing a resource demand mechanism to parse the concurrent learning task set of a learning machine terminal to obtain task parsing results; constructing a heterogeneous computing resource visualization based on the heterogeneous computing resource information of the learning machine terminal obtained through dynamic sensing; performing coalition scheduling analysis on the target parallel subtasks of the target stage in the task parsing results using the heterogeneous computing resource visualization as constraints to obtain the target optimal collaborative computing group; and distributing and executing the target parallel subtasks according to the target optimal collaborative computing group.
[0008] Optionally, a first learning task is extracted from the concurrent learning task set; a first subtask set corresponding to the first learning task is obtained, and a concurrent subtask set is constructed; the first subtask in the concurrent subtask set is subjected to feature analysis based on the predetermined task characteristics of memory in the resource requirement mechanism to obtain a first requirement feature parameter; the first subtask is ordered for execution based on the first requirement feature parameter to obtain a parallel execution sequence; the parallel execution sequence is used as the task parsing result; wherein, the predetermined task characteristics include computational load, response deadline, and task type.
[0009] Optionally, a local heterogeneous computing unit group is constructed for the learning machine terminal, wherein the local heterogeneous computing unit group includes a central processing unit and a neural network processor; an online heterogeneous computing unit group is constructed for the learning machine terminal, wherein the online heterogeneous computing unit group includes an edge processor and a cloud processor; predetermined resource features are retrieved to perform feature analysis on the central processing unit, the neural network processor, the edge processor, and the cloud processor to form the heterogeneous computing resource information; wherein the predetermined resource features include computing power, real-time load, and communication latency with the learning machine terminal.
[0010] Optionally, the target serial subtask sequence with the closest response deadline among the plurality of serial subtask sequences is compared; the parallel execution sequence is segmented into stages based on the target response deadline of the target serial subtask sequence to obtain the target stage.
[0011] Optionally, multiple target serial sequences corresponding to the target stage are matched among the multiple serial subtask sequences to form the target parallel subtask.
[0012] Optionally, using the target parallel subtask as the alliance target, the heterogeneous computing resource visualization is used to filter computing nodes to obtain a first collaborative computing group; it is determined whether the first predicted time for the first collaborative computing group to complete the target parallel subtask is no later than the target response deadline; if the first predicted time is no later than the target response deadline, the first collaborative computing group is added to the candidate collaborative computing group; the candidate collaborative computing group is optimized to obtain the target optimal collaborative computing group.
[0013] Optionally, subtasks of the interactive teaching type among the target parallel subtasks are selected and denoted as core subtasks; using the core subtasks as core alliance targets, the computing nodes of the neural network processors in the heterogeneous computing resource visualization are selected to obtain a core collaborative computing group; the core subtasks among the target parallel subtasks are removed to obtain front and back subtasks; using the front and back subtasks as front and back alliance targets, the computing nodes of the central processing unit, the edge processor, and the cloud processor in the heterogeneous computing resource visualization are selected to obtain a front and back collaborative computing group; the core collaborative computing group and the front and back collaborative computing group constitute the first collaborative computing group.
[0014] Optionally, based on the heterogeneous computing resource visibility diagram, the first maximum real-time load and the first average real-time load of the first collaborative computing group are obtained sequentially; based on the heterogeneous computing resource visibility diagram, the first predicted energy consumption of the first collaborative computing group in completing the target parallel subtask is obtained; the first maximum real-time load, the first average real-time load, and the first predicted energy consumption are weighted and calculated to obtain the first fitness; with the first fitness being maximized as the objective, the candidate collaborative computing groups are optimized and screened to obtain the target optimal collaborative computing group.
[0015] Optionally, the actual execution record of the target parallel subtask is obtained; the actual energy consumption in the actual execution record is extracted, and the heterogeneous computing resource visualization is dynamically corrected based on the actual energy consumption.
[0016] Secondly, this application also provides a heterogeneous computing resource scheduling system for intelligent mobile terminals under multi-task concurrency. The system includes: a task parsing module, used to introduce a resource demand mechanism to parse the concurrent learning task set of the learning machine terminal and obtain task parsing results; a view construction module, used to construct a heterogeneous computing resource view based on the heterogeneous computing resource information of the learning machine terminal obtained by dynamic perception; a scheduling analysis module, used to perform coalition scheduling analysis on the target parallel subtasks of the target stage in the task parsing results with the heterogeneous computing resource view as a constraint, and obtain the target optimal collaborative computing group; and a task distribution module, used by the learning machine terminal to distribute and execute the target parallel subtasks according to the target optimal collaborative computing group.
[0017] One or more technical solutions provided in this application have at least the following technical effects or advantages: By introducing a resource demand mechanism to parse the concurrent learning task set of the learning machine terminal, task parsing results are obtained. Based on the heterogeneous computing resource information of the learning machine terminal obtained through dynamic sensing, a heterogeneous computing resource view is constructed. Using the heterogeneous computing resource view as a constraint, a coalition scheduling analysis is performed on the target parallel subtasks in the target stage of the task parsing results to obtain the target optimal collaborative computing group. The learning machine terminal distributes and executes the target parallel subtasks according to the target optimal collaborative computing group. Ultimately, by constructing a heterogeneous resource view and a coalition-based collaborative scheduling mechanism, the technical effect of achieving efficient resource utilization under multi-task concurrency of the learning machine terminal, improving response speed, and optimizing terminal energy consumption is achieved.
[0018] The above description is merely an overview of the technical solution of this application. To enable a clearer understanding of the technical means of this application and to facilitate its implementation according to the description, and to make the above and other objects, features, and advantages of this application more apparent, specific embodiments of this application are described below. It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent through the following description. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0020] Figure 1 This is a flowchart illustrating the heterogeneous computing resource scheduling method for intelligent mobile terminals under multi-task concurrency in this application.
[0021] Figure 2 This is a schematic diagram of the heterogeneous computing resource scheduling system for intelligent mobile terminals under multi-task concurrency in this application.
[0022] Figure labeling: Task parsing module 11, Visualization module 12, Scheduling analysis module 13, Task distribution module 14. Detailed Implementation
[0023] This application provides a method and system for scheduling heterogeneous computing resources in intelligent mobile terminals under multi-task concurrency, solving the technical problems of existing technologies where learning machine terminals schedule heterogeneous resources according to the order of task arrival, resulting in high response latency for critical tasks and uneven resource utilization. It achieves the technical effect of efficiently utilizing resources under multi-task concurrency in learning machine terminals by constructing a heterogeneous resource visualization and a federation-based collaborative scheduling mechanism, thereby improving response speed and optimizing terminal energy consumption.
[0024] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. It should be understood that this application is not limited to the exemplary embodiments described herein. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application. It should also be noted that, for ease of description, only the parts related to this application are shown in the accompanying drawings, not all of them.
[0025] Example 1, please refer to the appendix. Figure 1 This application provides a method for scheduling heterogeneous computing resources in a smart mobile device under multi-task concurrency, wherein the method specifically includes: A resource requirement mechanism is introduced to parse the concurrent learning task set of the learning machine terminal and obtain the task parsing results.
[0026] Furthermore, a resource requirement mechanism is introduced to parse the concurrent learning task set of the learning machine terminal to obtain the task parsing result, including: extracting the first learning task from the concurrent learning task set; obtaining the first subtask set corresponding to the first learning task and constructing a concurrent subtask set; performing feature analysis on the first subtask in the concurrent subtask set according to the predetermined task characteristics of memory in the resource requirement mechanism to obtain the first requirement feature parameter; sorting the first subtask according to the first requirement feature parameter to obtain a parallel execution sequence; and using the parallel execution sequence as the task parsing result; wherein, the predetermined task characteristics include computational load, response deadline, and task type.
[0027] Specifically, the learning terminal acquires the set of concurrent learning tasks in the current operating environment in real time. This set refers to multiple independent learning tasks initiated or awaiting execution by the learning terminal simultaneously within the same time window, such as interactive quiz tasks, voice reading tasks, video course playback tasks, or homework correction tasks. A first learning task is extracted from this set, where "first" refers to the currently being processed task and has no sequential meaning. Then, based on a pre-defined task description file, task runtime metadata, or task structure information, the first learning task is broken down to obtain a corresponding first sub-task set. This first sub-task set consists of multiple functional execution units constituting the learning task, such as voice acquisition, voice noise reduction, and cache writing. These functional execution units in the first sub-task set are then combined with sub-tasks of other learning tasks to form a concurrent sub-task set. For example, the interactive question-answering task can be broken down into multiple subtasks, including user interface rendering, touch event acquisition, interactive animation playback, and real-time question accuracy inference. These subtasks can be combined with multiple subtasks for tasks such as voice reading, video course playback, or homework correction to form a concurrent subtask set.
[0028] Then, based on the predetermined task characteristics of memory in the resource requirement mechanism, the first subtask in the concurrent subtask set is subjected to feature analysis. The resource requirement mechanism is a set of rules pre-defined according to actual needs, used to assess the degree of resource requirements of the task during execution, such as computing, storage, and communication. The predetermined task characteristics include computational load, response deadline, and task type. Computational load represents the scale of computation required during task execution, quantified by CPU cycles or floating-point operations. The response deadline represents the latest time node that the task must complete from the current moment, i.e., the task latency constraint, determined based on user interaction experience metrics; for example, an interactive quiz task requires a response within 100ms. The task type represents the business category to which the task belongs, such as interactive teaching, video decoding, or data caching. Different task types correspond to different hardware resource adaptation relationships, automatically marked by the API feature task tags used during task invocation.
[0029] Based on the predetermined task characteristics, combined with the task execution logs, historical execution records, and task tags, the first subtask in the concurrent subtask set is subjected to feature analysis to obtain the corresponding computational load, response deadline, and task type, forming the first requirement feature parameters. After obtaining the requirement feature parameters of all subtasks, multiple subtasks are sorted and analyzed according to the requirement feature parameters. The response deadline is taken as the primary objective to determine the urgency of each subtask; that is, the shorter the response deadline, the more urgent the corresponding subtask. When the deadlines are the same, the computational load is used as the sorting objective, and the subtasks are sorted in descending order, prioritizing the scheduling of subtasks with higher computational load to avoid blocking subsequent tasks. After sorting, subtasks that have no data dependencies and can run simultaneously on different computing units are grouped into the same parallel group. Subtasks within each group can be executed concurrently, and subtasks with data dependencies (i.e., the output of one subtask is the input of another subtask) are serialized to form a serial subtask sequence. Multiple parallel groups are integrated according to their dependencies to form a parallel execution sequence, which includes multiple serial subtask sequences. The parallel execution sequence is then used as the task parsing result for subsequent analysis and processing.
[0030] For example, the learning terminal simultaneously receives three learning tasks: an AI oral interaction task, an online video playback task, and an assignment image recognition task. The first learning task, the AI oral interaction task, is extracted and broken down into a voice acquisition subtask, a noise reduction subtask, a semantic recognition subtask, and a result feedback subtask. Then, based on the pre-defined memory requirements in the resource demand mechanism, the semantic recognition subtask is found to have a computational cost of 18 GFLOPs and a response deadline of 120ms, classifying it as an interactive teaching task. The noise reduction subtask has a computational cost of 2 GFLOPs and a deadline of 200ms. The result feedback subtask only involves interface refresh and has a low computational cost. Based on the above analysis results, the required feature parameters for each subtask are generated and sorted. According to the sorting rules, the semantic recognition subtask is a task with high real-time and high AI computing power requirements, and the noise reduction subtask is a task with medium real-time and medium computing power requirements. Therefore, a parallel execution sequence is formed: the first serial subtask sequence is speech acquisition, noise reduction processing, semantic recognition and result feedback, and the second serial subtask sequence is video cache update and video decoding. There is no strong dependency between the multiple serial subtask sequences, so they can be executed concurrently. Thus, the first serial subtask sequence and the second serial subtask sequence are combined to form an overall parallel execution sequence.
[0031] By analyzing the set of concurrent learning tasks in the learning terminal, it is transformed into a fine-grained sequence of subtasks with clear computational load, response deadlines, and task types. The parallel and serial relationships between multiple subtasks are also identified, providing a unified and standardized task input foundation for subsequent heterogeneous resource collaborative scheduling. Compared to the traditional method of scheduling based solely on task arrival order, this approach effectively reduces the response latency of critical learning tasks and improves resource utilization and system stability in multi-task concurrent scenarios.
[0032] Based on the heterogeneous computing resource information of the learning machine terminal obtained by dynamic sensing, a heterogeneous computing resource visualization is constructed.
[0033] Furthermore, based on the heterogeneous computing resource information of the learning machine terminal obtained through dynamic sensing, a heterogeneous computing resource visualization is constructed, which includes: assembling a local heterogeneous computing unit group for the learning machine terminal, wherein the local heterogeneous computing unit group includes a central processing unit (CPU) and a neural network processor (NNPM); assembling an online heterogeneous computing unit group for the learning machine terminal, wherein the online heterogeneous computing unit group includes an edge processor and a cloud processor; retrieving predetermined resource features to perform feature analysis on the CPU, the NNPM, the edge processor, and the cloud processor to form the heterogeneous computing resource information; wherein the predetermined resource features include computing power, real-time load, and communication latency with the learning machine terminal.
[0034] Specifically, heterogeneous computing resources refer to various computing units with different computing architectures, instruction sets, computing capabilities, or deployment locations, such as local CPUs on the terminal, local NPUs, edge servers, and cloud servers. Before constructing a heterogeneous computing resource visualization, the local heterogeneous computing unit group of the learning machine terminal is first established. The local heterogeneous computing unit group refers to a collection of different types of computing cores deployed on the local hardware platform of the learning machine terminal, including at least a central processing unit (CPU) and a neural network processor (NPU). Among them, the CPU is mainly used for general logic control, data scheduling, lightweight computing, and system background tasks, while the neural network processor is an AI acceleration unit integrated into the SoC chip, such as a dedicated AI computing module that supports INT8 / FP16 tensor operations, mainly used for highly parallel computing tasks such as deep learning inference, matrix parallel operations, and AI interactive recognition. For the central processing unit and neural network processor, terminal hardware information can be read through the underlying hardware abstraction layer (HAL), such as the number of CPU cores, clock speed, cache capacity, NPU tensor throughput, and current power consumption. The CPU computing power is represented by the number of integer operations per unit time, floating-point operation power, or SPEC benchmark performance value, while the NPU computing power is quantified by TOPS.
[0035] Simultaneously, an online heterogeneous computing unit group is established for the learning machine terminal. This group refers to a collection of remote computing resources located outside the learning machine terminal that can participate in collaborative computing via network connection, including edge processors and cloud processors. Edge processors are deployed on edge server nodes close to the learning machine terminal, such as edge nodes of a campus LAN or edge MEC nodes of a carrier, offering lower communication latency and faster response times. Cloud processors are deployed in remote cloud computing centers, providing higher computing power and larger storage resources. The learning machine terminal can register and connect to edge and cloud nodes through network resource discovery protocols. For example, it can establish a resource synchronization channel with the edge scheduling server via MQTT, WebSocket, or HTTP / 2 long-connection mechanisms and periodically obtain the operational status information of edge and cloud nodes.
[0036] Retrieve predetermined resource characteristics, which are a unified indicator system used to reflect the current operating status and available capabilities of each computing node. These characteristics include at least computing power, real-time load, and communication latency of the learning machine terminal. The computing power represents the scale of computation that a node can complete per unit time, the real-time load reflects the current resource occupancy status of the node, and the communication latency represents the data transmission delay between the learning machine terminal and the target node.
[0037] Based on predetermined resource characteristics, feature analysis is performed on the central processing unit (CPU), neural network processor (NNPM), edge processor, and cloud processor. For example, the CPU is represented using GFLOPS (billions of floating-point operations per second), the NNPM using TOPS (trillions of operations per second), and the edge and cloud servers using the number of GPU CUDA cores. Then, for the CPU and NNPM, real-time CPU utilization, NPU tensor queue length, edge node task queue length, and cloud server GPU utilization are calculated to obtain the corresponding real-time load, reflecting the current remaining schedulable capacity of the nodes. For the edge and cloud processors, the learning machine terminal periodically sends load query requests to them to obtain the current CPU load. Furthermore, by sending an empty data packet, such as a 64-byte heartbeat packet, from the learning machine terminal, the communication latency of the CPU, NNPM, edge processor, cloud processor, and learning machine terminal is obtained. The communication latency of the CPU and NNPM can be approximated as on-chip bus access latency, while the latency of the edge and cloud processors is quantified using network round-trip time (RTT). Normalization calculation methods, such as min-max normalization or z-score normalization, are used to convert computing power, load rate and communication latency of different dimensions into a unified numerical range.
[0038] Based on the heterogeneous computing resource information of the learning machine terminal obtained by the above dynamic perception, a graph data structure is used for modeling: the central processing unit, neural network processor, edge server and cloud server are abstracted as graph nodes, and the data transmission relationship between nodes is abstracted as graph edges. Each node stores the corresponding computing power and real-time load, and each edge stores the corresponding communication delay, forming a heterogeneous computing resource visualization. The heterogeneous computing resource visualization can be updated in real time according to the system operating status. For example, when the load of the neural network processor increases, its node weight changes automatically, and when the edge node network is congested, its communication edge delay parameter increases dynamically.
[0039] By uniformly perceiving the heterogeneous computing resources inside and outside the learning machine terminal, and dynamically modeling the central processing unit, neural network processor, edge real-time, and large-scale cloud computing power, task matching and scheduling can be performed based on real-time resource status. This improves the utilization rate of heterogeneous resources, reduces the response time of critical tasks, and reduces scheduling errors caused by resource status distortion.
[0040] Using the heterogeneous computing resource visualization as a constraint, the target parallel subtasks in the target stage of the task parsing results are subjected to coalition scheduling analysis to obtain the target optimal cooperative computing group.
[0041] Furthermore, the parallel execution sequence includes multiple serial subtask sequences. Using the heterogeneous computing resource visibility as a constraint, a coalition scheduling analysis is performed on the target parallel subtasks of the target stage in the task parsing results to obtain the target optimal collaborative computing group. This includes: comparing the target serial subtask sequence with the closest response deadline among the multiple serial subtask sequences; and segmenting the parallel execution sequence into stages based on the target response deadline of the target serial subtask sequence to obtain the target stage.
[0042] Furthermore, multiple target serial sequences corresponding to the target stage are matched among the multiple serial subtask sequences to form the target parallel subtask.
[0043] Specifically, multiple serial subtask sequences are extracted from the parallel execution sequence. Then, the response deadlines of these sequences are compared, and the sequences are sorted in ascending order. The sequence with the smallest response deadline is selected as the target serial subtask sequence. In other words, the sequence closest to the current system time and with the highest real-time requirement is chosen as the target serial subtask sequence. For example, at the current time, if the interactive speech recognition task has a remaining deadline of 80ms and the video caching task has a remaining deadline of 500ms, then the sequence corresponding to the speech recognition task is selected as the target serial subtask sequence.
[0044] The estimated execution time is obtained by dividing the computational complexity feature value by the computational capability corresponding to the subtask. Then, the target response deadline of the target serial subtask sequence is used as a benchmark, combined with the estimated execution time of each subtask in the parallel execution sequence, to divide the parallel execution sequence into stages. For example, a dynamic time window partitioning mechanism is used, with the target response deadline as the boundary of the current stage. All task sequences are mapped onto a timeline, and the set of tasks that need to be completed within the time window is selected. Subtasks that need to be started or completed within the target response deadline are divided into the current target stages.
[0045] The system matches multiple target serial sequences corresponding to the target stage within multiple serial subtask sequences to form target parallel subtasks. A target serial sequence refers to a sequence of serial subtasks whose execution time range falls within the current target stage. Target parallel subtasks are a set of subtasks extracted from the target serial sequences that can be executed synchronously within the current stage. It should be noted that target parallel subtasks are not equivalent to the entire task set, but rather a local real-time task set after stage filtering, used to reduce the scheduling search space and improve the response efficiency of critical tasks. For example, the time interval of each subtask is calculated based on its earliest start time, estimated execution duration, and deadline, and then it is determined whether this time interval falls within the current target stage. For subtasks that meet the conditions, their dependencies are further analyzed, and multiple subtasks without mutual exclusion dependencies are combined to form target parallel subtasks. For example, there is no execution dependency between the AI semantic recognition task and the video cache update task, so they are combined into a set of target parallel subtasks.
[0046] For example, by analyzing the existing three serial subtask sequences—the first, AI interactive question-and-answer sequence, which includes voice acquisition, semantic recognition, and answer generation, with a response deadline of 150ms; the second, video playback sequence, which includes video caching, video decoding, and image rendering, with a response deadline of 600ms; and the third, learning data synchronization sequence, which includes log compression and data upload, with a corresponding response deadline of 2000ms—comparing the response deadlines of the three serial subtask sequences, it is found that the first serial subtask sequence has the shortest deadline, and it is determined as the target serial subtask sequence, with 150ms as the time boundary of the current target stage. Analyzing the time intervals of subtasks in other task sequences, it is found that the video caching subtask needs to be completed within 100ms, the video decoding task is expected to be executed after 200ms, and the data upload task does not need to start within 150ms. Therefore, voice acquisition, semantic recognition, answer generation, and video caching are included in the current target stage and form a target parallel subtask set.
[0047] By breaking down concurrent tasks into stages in real time, the learning tasks with the highest real-time requirements can be prioritized, avoiding scheduling delays and resource contention caused by traditional global unified scheduling methods. Furthermore, by performing subsequent coalition scheduling analysis only on the parallel subtasks of the current target stage, the search complexity in the heterogeneous resource matching process can be reduced, improving scheduling efficiency and system response speed.
[0048] Furthermore, using the heterogeneous computing resource visibility as a constraint, a coalition scheduling analysis is performed on the target parallel subtasks in the target stage of the task parsing results to obtain the target optimal collaborative computing group. This includes: using the target parallel subtasks as coalition targets, filtering computing nodes in the heterogeneous computing resource visibility to obtain a first collaborative computing group; determining whether the first predicted time for the first collaborative computing group to complete the target parallel subtasks is no later than the target response deadline; if the first predicted time is no later than the target response deadline, then adding the first collaborative computing group to the candidate collaborative computing group; and optimizing the candidate collaborative computing group to obtain the target optimal collaborative computing group.
[0049] Furthermore, using the target parallel subtask as the alliance target, the heterogeneous computing resource view is used to filter computing nodes to obtain a first collaborative computing group, including: filtering the target parallel subtasks whose task type is interactive teaching, and denoting them as core subtasks; using the core subtasks as the core alliance target, the neural network processors in the heterogeneous computing resource view are used to filter computing nodes to obtain a core collaborative computing group; removing the core subtasks from the target parallel subtasks to obtain front and back subtasks; using the front and back subtasks as front and back alliance targets, the central processing unit, the edge processor, and the cloud processor in the heterogeneous computing resource view are used to filter computing nodes to obtain a front and back collaborative computing group; the core collaborative computing group and the front and back collaborative computing group constitute the first collaborative computing group.
[0050] Specifically, each subtask in the target parallel subtask is traversed, and the task type of each subtask is obtained according to the corresponding requirement feature parameters. Then, subtasks with the task type of interactive teaching are extracted. The interactive teaching type refers to the task type with high requirements for real-time interaction capability and AI reasoning capability, such as speech semantic recognition, real-time oral assessment, interactive question answering reasoning or action recognition. Interactive teaching type tasks usually have characteristics such as needing to perform deep learning reasoning, being highly sensitive to response latency, having a high data interaction frequency, and needing strong parallel tensor computing capability. Subtasks selected as interactive teaching type are used as core subtasks.
[0051] The core sub-tasks are designated as the core alliance objectives. These objectives refer to dynamically combining multiple heterogeneous computing nodes into a collaborative execution alliance based on the computational characteristics, task type, and real-time requirements of different sub-tasks, enabling multiple nodes to participate in the concurrent processing of tasks at the same stage. Based on the core alliance objectives, computing nodes for neural network processors are screened in the heterogeneous computing resource visualization. For example, using a constraint filtering method, neural network processor nodes with real-time loads exceeding a preset threshold (e.g., 80%) are first eliminated. Then, candidate neural network processor nodes with prediction inference time and communication latency less than or equal to the current target response deadline are selected. All candidate neural network processor nodes meeting the screening criteria are combined to form a core collaborative computing group.
[0052] The core subtasks in the target parallel subtasks are removed, and the remaining subtasks are designated as preceding and following subtasks. These preceding and following subtasks refer to auxiliary tasks located before or after the AI inference task, such as data preprocessing, cache updating, result rendering, logging, data synchronization, and network upload tasks. These preceding and following subtasks are used as preceding and following alliance targets. Computational nodes are screened from the central processing unit (CPU), edge processor, and cloud processor in the heterogeneous computing resource visualization. For the CPU: the real-time load and computing power of the CPU are read, and subtasks with lower computational load and sensitivity to communication latency are prioritized for allocation to the CPU. A greedy algorithm can be used for the matching strategy: all preceding and following subtasks are sorted in descending order of computational load × communication latency, and then sequentially allocated to the CPU node with the lowest current load. If the remaining available computing power of the CPU meets the subtask's requirements, a mapping relationship is established and the CPU load is updated; otherwise, the subtask is marked as awaiting allocation to an edge processor or cloud processor.
[0053] For edge processors: Obtain the computing power, real-time load, and communication latency of edge processor nodes in the heterogeneous computing resource visualization. For subtasks not allocated to the central processing unit (CPU) node, assign them to the edge processor. The criterion is: divide the computational load of the subtask by the remaining available computing power of the edge processor and add twice the communication latency. Determine if the result is less than a preset threshold, such as 30% of the target response deadline. If satisfied, assign the subtask to the edge processor and add it to the front-end / back-end collaborative computing group. For cloud processors: Obtain the computing power, real-time load, and communication latency of cloud processor nodes. Assign the subtask with the highest computational load among the front-end / back-end subtasks that remains unassigned after allocation by the CPU and edge processors to the cloud processor node. Integrate the successfully assigned groups to form front-end / back-end collaborative computing groups.
[0054] The core collaborative computing group and the front-end and back-end collaborative computing groups are combined to form the first collaborative computing group. The first collaborative computing group is a heterogeneous computing alliance oriented towards the current stage task, which specifies the specific computing nodes distributed to each subtask in the target parallel subtask.
[0055] The first predicted time for the first collaborative computing group to complete the target parallel subtask is compared with the target response deadline. The first predicted time is the total execution time required to complete the entire target parallel subtask based on the current resource status. Critical path analysis can be used to predict the time of parallel tasks. The first predicted time = maximum parallel sub-path execution time + inter-node communication synchronization time. The maximum parallel sub-path execution time in the first predicted time refers to the task execution path with the longest expected completion time among multiple parallel execution paths under the current task allocation result of the first collaborative computing group. First, a directed acyclic graph (DAG) is constructed based on the target parallel subtask. Each subtask node records the corresponding task computation volume, input data volume, and target allocation node. The edges between nodes represent task dependencies and data transmission relationships. Then, based on the real-time resource status in the heterogeneous computing resource visualization, the expected execution time of each subtask on the corresponding allocation node is calculated. For example, subtask execution time = subtask computation volume ÷ target node computing capacity + node queuing waiting time. The node queuing waiting time is estimated using the current node task queue length, resource utilization rate, and historical scheduling statistics. The DAG (Directed Acyclic Graph) is used to traverse all parallel execution paths, and the execution time of subtasks on each path and the data transmission time within the path are accumulated. The path with the largest accumulated time is selected as the maximum parallel sub-path execution time. For example, when the first path (speech acquisition, semantic recognition, and answer generation) and the second path (video caching and video decoding) are executed simultaneously, if the first path accumulates a time of 65ms and the second path accumulates a time of 42ms, then 65ms is determined as the maximum parallel sub-path execution time.
[0056] The inter-node communication synchronization time refers to the additional time overhead incurred by task migration, data transmission, and result synchronization during the collaborative execution of the target parallel subtask on multiple heterogeneous nodes, reflecting the data interaction latency between different computing nodes. The data dependencies between nodes are determined based on the task allocation results, and the communication latency of the corresponding links is obtained from the network status parameters in the heterogeneous computing resource visualization. Multiple communication latencyes are accumulated to form the inter-node communication synchronization time, which is then added to the execution time of the maximum parallel sub-path to form the first predicted time, used to determine whether the current first collaborative computing group meets the target response deadline requirement.
[0057] If the first predicted time is less than or equal to the target response deadline, it indicates that the collaborative computing group has the execution capability to meet real-time requirements, and this first collaborative computing group is added to the candidate collaborative computing groups. If the first predicted time is greater than the target response deadline, the collaborative computing group is deemed infeasible and discarded. Other collaborative computing groups are then evaluated, such as by changing the allocation strategy or moving some subtasks from the cloud to the edge, until a feasible collaborative computing group is found. Then, other available collaborative computing combinations are traversed to generate multiple candidate collaborative computing groups that meet the time limit requirements. These candidate collaborative computing groups are then optimized to finally obtain the target optimal collaborative computing group.
[0058] By separating and filtering core subtasks from their preceding and following subtasks, precise utilization of the neural network processor is achieved. Auxiliary tasks are dynamically allocated to central processing units, edge processors, or cloud processor nodes for execution, thereby avoiding performance bottlenecks and resource waste. Simultaneously, through deadline constraints and a coalition optimization mechanism, the response latency of critical tasks can be effectively reduced, the overall utilization rate of heterogeneous resources can be improved, and resource contention and system energy consumption in multi-task concurrent scenarios can be reduced.
[0059] Furthermore, optimizing the candidate collaborative computing groups to obtain the target optimal collaborative computing group includes: based on the heterogeneous computing resource visibility diagram, sequentially obtaining the first maximum real-time load and the first average real-time load of the first collaborative computing group; based on the heterogeneous computing resource visibility diagram, obtaining the first predicted energy consumption of the first collaborative computing group to complete the target parallel subtask; performing a weighted calculation on the first maximum real-time load, the first average real-time load, and the first predicted energy consumption to obtain a first fitness; and optimizing and screening the candidate collaborative computing groups with the first fitness as the target to obtain the target optimal collaborative computing group.
[0060] Specifically, based on the heterogeneous computing resource visualization, the real-time load of each computing node in the first collaborative computing group is obtained sequentially, resulting in multiple real-time loads. These multiple real-time loads are then compared, and the maximum value among them is taken as the first maximum real-time load. The first maximum real-time load reflects the computing node under the greatest pressure in the first collaborative computing group. The average value of the obtained multiple real-time loads is then calculated to obtain the first average real-time load, reflecting the overall resource utilization efficiency. For example, if the CPU load in the first collaborative computing group is 45%, the NPU load is 68%, the edge node load is 81%, and the cloud node load is 35%, then 81% is determined as the first maximum real-time load, and the average value is calculated to obtain the first average real-time load = 57.25%.
[0061] Then, based on the heterogeneous computing resource visibility map, the first predicted energy consumption for the first collaborative computing group to complete the target parallel sub-task is calculated. Predicted energy consumption refers to the system's calculation of the total energy consumption expected during task execution based on the current node operating status, task execution time, and communication overhead. First, the unit-time power consumption parameters of each computing node in the first collaborative computing group under the current load state are obtained. For example, for the central processing unit and neural network processor, this is obtained through a mapping model. This mapping model is obtained through offline calibration, i.e., using a power measuring instrument to record power consumption curves under different loads and computational loads, fitting them as linear or quadratic functions, such as the unit-time power consumption parameter E = Pi × Ti + k × C, where Pi is the static power consumption of node i, i.e., the power when idle, Ti is the execution time of the sub-task on that node, k is the dynamic power consumption coefficient, and C is the computational load of the sub-task. Edge processor nodes and cloud processor nodes are obtained through the server power monitoring interface. Then, the task execution energy consumption of each node is calculated based on the expected task execution time: Node execution energy consumption = Average node power consumption × Task execution time. The first predicted energy consumption is obtained by summing the unit time power consumption parameters of multiple computing nodes under the current load state with the node execution energy consumption.
[0062] The first collaborative computing group's three evaluation metrics—the first maximum real-time load, the first average real-time load, and the first predicted energy consumption—are weighted to obtain the first fitness, which is calculated as: First Fitness = w1 × Load Balancing Metric + w2 × Average Load Optimization Metric + w3 × Energy Consumption Optimization Metric. Here, w1, w2, and w3 are preset weight coefficients, satisfying w1 + w2 + w3 = 1. Specific values can be set according to the learning machine's scheduling strategy preferences. For example, for learning machines emphasizing battery life, such as children's tablets, w3 can be set higher, such as 0.5, and w1 and w2 can be 0.25 respectively. For learning machines emphasizing performance, w1 can be set higher. The load balancing metric is composed of the inverse value of the first maximum real-time load, the average load optimization metric is composed of the inverse value of the first average real-time load, and the energy consumption optimization metric is composed of the inverse value of the first predicted energy consumption. Furthermore, to avoid the influence of different parameter dimensions on the results, a minimum-maximum normalization method is used to uniformly map each metric to the [0,1] interval before weighted calculation.
[0063] Then, taking the maximum fitness as the objective, we optimize and screen all candidate collaborative computing groups, and select the candidate collaborative computing group with the maximum fitness as the target optimal collaborative computing group.
[0064] By performing multi-objective comprehensive optimization of the heterogeneous collaborative resource alliance, the optimal collaborative scheduling of the current stage task is achieved. This enables the learning machine terminal to not only meet the real-time response requirements of the target task, but also avoid long-term high-load operation of local nodes and effectively reduce the overall energy consumption in the process of mobile terminal and edge collaboration. This improves the resource utilization efficiency and long-term operating efficiency of the learning machine terminal in multi-task concurrent scenarios.
[0065] The learning machine terminal distributes and executes the target parallel subtasks according to the target optimal collaborative computing group.
[0066] Specifically, after obtaining the optimal collaborative computing group for the target, the learning machine terminal distributes and executes the target parallel subtasks according to the node allocation relationship in the optimal collaborative computing group. Through the local task scheduler, edge task agent module, and cloud task interface, multiple subtasks in the target parallel subtasks are sent to the corresponding computing nodes for execution, so that different types of subtasks can be matched with corresponding computing resources. This reduces the response latency of key interactive teaching tasks, improves the utilization rate of heterogeneous computing resources, optimizes terminal energy consumption, and reduces the problem of task blocking and increased system energy consumption caused by excessive load on a single node in multi-task concurrent scenarios.
[0067] Furthermore, the learning machine terminal distributes and executes the target parallel subtasks according to the target optimal collaborative computing group, and then further includes: obtaining the actual execution record of the target parallel subtasks; extracting the actual energy consumption from the actual execution record, and dynamically correcting the heterogeneous computing resource visualization based on the actual energy consumption.
[0068] Specifically, during the process of distributing and executing target parallel subtasks according to the target optimal collaborative computing group, the running status of each node is recorded synchronously, and the actual execution record corresponding to the target parallel subtask is generated. The actual execution record refers to the set of running logs formed by collecting running data of the learning machine terminal during the actual operation of the task through the operating system performance monitoring interface, AI driver interface or chip power management module, which reflects the actual execution status of the task in the actual resource environment. The actual execution record includes at least the following: actual task start time, actual task completion time, node real-time occupancy rate, actual task queuing waiting time, actual communication latency and actual energy consumption data.
[0069] The actual energy consumption is obtained from the actual execution records of the target parallel subtasks, i.e., the power resources actually consumed during the actual execution of the task. Then, the heterogeneous computing resource visualization is dynamically corrected based on the actual energy consumption. Specifically, the actual energy consumption is compared with the first predicted energy consumption. By calculating the difference between the actual energy consumption and the first predicted energy consumption, an energy consumption deviation value is obtained. The energy consumption deviation value is compared with a preset energy consumption threshold. If the energy consumption deviation value is greater than the preset energy consumption threshold, it indicates that there is a deviation between the current resource model and the actual running state, and the source of the deviation is further analyzed. For example, if the actual energy consumption of the NPU is higher than the predicted value, it means that the current chip temperature has increased, resulting in increased power consumption. If the energy consumption of edge node communication increases abnormally, it means that the current network congestion has led to an increase in the number of retransmissions. If the CPU execution energy consumption is higher than the predicted value, it means that background task competition has led to an increase in CPU frequency.
[0070] A sliding window mean update mechanism is used to dynamically correct the heterogeneous computing resource visibility map. For example, the corrected parameter = λ × historical parameter + (1-λ) × current actual parameter, where λ is the historical weight coefficient, used to control the weight ratio of historical resource status and current actual operating status in the heterogeneous computing resource visibility map during the dynamic correction process, thereby avoiding drastic fluctuations in the heterogeneous computing resource visibility map due to instantaneous load fluctuations or occasional network anomalies. For example, historical operating data of each computing node is continuously collected within a preset time window, such as 6 hours or 24 hours, including parameters such as CPU / NPU load rate, task execution time, communication latency, and actual energy consumption. Historical parameters are obtained by moving average, and the state fluctuation coefficient of the corresponding parameter in multiple consecutive sampling periods is calculated by standard deviation. λ is obtained by 1 / (1+state fluctuation coefficient). Specifically, when node load changes are small, communication latency is stable, and historical energy consumption fluctuations are low, it indicates that the current resource environment is relatively stable. In this case, the λ value can be increased, for example, to 0.7~0.8, making the correction process more reliant on historical parameters. Conversely, when edge network congestion and significant CPU frequency fluctuations are detected, it indicates that the current resource state has changed significantly. In this case, the λ value can be decreased, for example, to 0.3~0.5, increasing the proportion of current actual parameters in the correction, enabling the heterogeneous computing resource visualization to respond more quickly to changes in the real operating environment. By continuously learning the real resource state based on task execution, the predicted execution time, predicted energy consumption, and node fitness evaluation in subsequent resource scheduling become closer to the real operating environment.
[0071] By acquiring the actual execution records of the target parallel subtasks, a closed-loop operation feedback mechanism for heterogeneous resource scheduling is achieved. This not only enables task scheduling but also allows for the continuous visualization of heterogeneous computing resources based on the actual execution results. This improves the accuracy and stability of subsequent collaborative computing group generation, reduces resource prediction errors, and further enhances the long-term scheduling reliability and energy consumption control capabilities in multi-task concurrent scenarios.
[0072] Example 2: Based on the same inventive concept as the heterogeneous computing resource scheduling method for intelligent mobile devices under multi-task concurrency in Example 1, this application also provides a heterogeneous computing resource scheduling system for intelligent mobile devices under multi-task concurrency. Please refer to the appendix. Figure 2 The heterogeneous computing resource scheduling system for intelligent mobile terminals under multi-task concurrency includes: The task parsing module 11 is used to introduce a resource requirement mechanism to parse the concurrent learning task set of the learning machine terminal and obtain the task parsing result; the visualization construction module 12 is used to construct a heterogeneous computing resource visualization based on the heterogeneous computing resource information of the learning machine terminal obtained by dynamic perception; the scheduling analysis module 13 is used to perform coalition scheduling analysis on the target parallel subtasks of the target stage in the task parsing result with the heterogeneous computing resource visualization as a constraint to obtain the target optimal collaborative computing group; the task distribution module 14 is used for the learning machine terminal to distribute and execute the target parallel subtasks according to the target optimal collaborative computing group.
[0073] Furthermore, the task parsing module 11 is also used to: extract a first learning task from the concurrent learning task set; obtain a first subtask set corresponding to the first learning task and construct a concurrent subtask set; perform feature analysis on the first subtask in the concurrent subtask set according to the predetermined task characteristics of memory in the resource requirement mechanism to obtain a first requirement feature parameter; sort the first subtask according to the first requirement feature parameter to obtain a parallel execution sequence; and use the parallel execution sequence as the task parsing result; wherein the predetermined task characteristics include computational load, response deadline, and task type.
[0074] Furthermore, the visualization construction module 12 is also used to: construct a local heterogeneous computing unit group for the learning machine terminal, wherein the local heterogeneous computing unit group includes a central processing unit and a neural network processor; construct an online heterogeneous computing unit group for the learning machine terminal, wherein the online heterogeneous computing unit group includes an edge processor and a cloud processor; retrieve predetermined resource features to perform feature analysis on the central processing unit, the neural network processor, the edge processor, and the cloud processor to form the heterogeneous computing resource information; wherein the predetermined resource features include computing power, real-time load, and communication latency with the learning machine terminal.
[0075] Furthermore, the scheduling analysis module 13 is also used to: compare the target serial subtask sequence with the closest response deadline among the multiple serial subtask sequences; and perform stage segmentation on the parallel execution sequence based on the target response deadline of the target serial subtask sequence to obtain the target stage.
[0076] Furthermore, the scheduling analysis module 13 is also used to: match multiple target serial sequences corresponding to the target stage in the multiple serial subtask sequences to form the target parallel subtask.
[0077] Furthermore, the scheduling analysis module 13 is also used to: use the target parallel subtask as the alliance target, filter computing nodes in the heterogeneous computing resource visualization to obtain a first collaborative computing group; determine whether the first predicted time for the first collaborative computing group to complete the target parallel subtask is not later than the target response deadline; if the first predicted time is not later than the target response deadline, add the first collaborative computing group to the candidate collaborative computing group; and optimize the candidate collaborative computing group to obtain the target optimal collaborative computing group.
[0078] Furthermore, the scheduling analysis module 13 is also used to: filter the interactive teaching type subtasks among the target parallel subtasks, and denot them as core subtasks; use the core subtasks as core alliance targets to filter the computing nodes of the neural network processors in the heterogeneous computing resource view to obtain a core collaborative computing group; remove the core subtasks from the target parallel subtasks to obtain front and back subtasks; use the front and back subtasks as front and back alliance targets to filter the computing nodes of the central processing unit, the edge processor, and the cloud processor in the heterogeneous computing resource view to obtain a front and back collaborative computing group; the core collaborative computing group and the front and back collaborative computing group form the first collaborative computing group.
[0079] Furthermore, the scheduling analysis module 13 is also used to: based on the heterogeneous computing resource visibility diagram, sequentially obtain the first maximum real-time load and the first average real-time load of the first collaborative computing group; based on the heterogeneous computing resource visibility diagram, obtain the first predicted energy consumption of the first collaborative computing group to complete the target parallel subtask; perform weighted calculation on the first maximum real-time load, the first average real-time load, and the first predicted energy consumption to obtain a first fitness; and optimize and screen the candidate collaborative computing groups with the first fitness as the target to obtain the target optimal collaborative computing group.
[0080] Furthermore, the task distribution module 14 is also used to: obtain the actual execution record of the target parallel subtask; extract the actual energy consumption in the actual execution record, and dynamically correct the heterogeneous computing resource visibility based on the actual energy consumption.
[0081] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0082] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application also intends to include such modifications and variations.
Claims
1. A method for scheduling heterogeneous computing resources of an intelligent mobile terminal under multitask concurrency, characterized in that, include: A resource demand mechanism is introduced to parse the concurrent learning task set of the learning machine terminal and obtain the task parsing results; Based on the heterogeneous computing resource information of the learning machine terminal obtained by dynamic sensing, a heterogeneous computing resource visualization is constructed. Using the heterogeneous computing resource visualization as a constraint, the target parallel subtasks in the target stage of the task parsing results are subjected to coalition scheduling analysis to obtain the target optimal cooperative computing group; The learning machine terminal distributes and executes the target parallel subtasks according to the target optimal collaborative computing group.
2. The method of claim 1, wherein the method further comprises: A resource requirement mechanism is introduced to parse the concurrent learning task set of the learning machine terminal, and the task parsing results are obtained, including: Extract the first learning task from the concurrent learning task set; Obtain the first subtask set corresponding to the first learning task, and assemble a concurrent subtask set; Based on the predetermined task characteristics of memory in the resource demand mechanism, the first subtask in the concurrent subtask set is subjected to feature analysis to obtain the first demand feature parameters. The first subtasks are ordered for execution based on the first requirement feature parameters to obtain a parallel execution sequence; The parallel execution sequence is used as the task parsing result; The predetermined task characteristics include computational load, response deadline, and task type.
3. The method of claim 2, wherein the method further comprises: Based on the heterogeneous computing resource information of the learning machine terminal obtained through dynamic sensing, a heterogeneous computing resource visualization is constructed, including the following: A local heterogeneous computing unit group is constructed for the learning machine terminal, wherein the local heterogeneous computing unit group includes a central processing unit and a neural network processor; An online heterogeneous computing unit group is constructed for the learning machine terminal, wherein the online heterogeneous computing unit group includes an edge processor and a cloud processor; The predetermined resource characteristics are retrieved to perform feature analysis on the central processing unit, the neural network processor, the edge processor, and the cloud processor to form the heterogeneous computing resource information; The predetermined resource characteristics include computing power, real-time load, and communication latency with the learning machine terminal.
4. The method of claim 3, wherein the method further comprises: The parallel execution sequence includes multiple serial subtask sequences. Using the heterogeneous computing resource visibility as a constraint, a coalition scheduling analysis is performed on the target parallel subtasks of the target stage in the task parsing results to obtain the target optimal cooperative computing group. This includes: Compare the target serial subtask sequence with the closest response deadline among the multiple serial subtask sequences; The parallel execution sequence is divided into stages based on the target response deadline of the target serial subtask sequence to obtain the target stage.
5. The method of claim 4, wherein the method further comprises: The target parallel subtask is formed by matching multiple target serial sequences corresponding to the target stage in the multiple serial subtask sequences.
6. The method of claim 5, wherein the method further comprises: Using the heterogeneous computing resource visibility as a constraint, a coalition scheduling analysis is performed on the target parallel subtasks in the target stage of the task parsing results to obtain the target optimal cooperative computing group, including: Using the target parallel subtask as the alliance target, the computing nodes of the heterogeneous computing resource visualization are filtered to obtain the first collaborative computing group; Determine whether the first predicted time for the first collaborative computing group to complete the target parallel subtask is no later than the target response deadline; If the first prediction time is not later than the target response deadline, then the first collaborative computing group is added to the candidate collaborative computing group; The target optimal collaborative computing group is obtained by optimizing the candidate collaborative computing groups.
7. The heterogeneous computing resource scheduling method for intelligent mobile terminals under multi-task concurrency as described in claim 6, characterized in that, Using the target parallel subtask as the alliance target, the heterogeneous computing resource visualization is used to filter computing nodes to obtain the first collaborative computing group, including: Select the interactive teaching type subtasks from the target parallel subtasks and denote them as core subtasks; Using the core sub-task as the core alliance target, the neural network processors in the heterogeneous computing resource visualization are screened for computing nodes to obtain a core collaborative computing group. Remove the core subtask from the target parallel subtask to obtain the preceding and following subtasks; Using the aforementioned front and back subtasks as front and back alliance targets, the computing nodes of the central processing unit, the edge processor, and the cloud processor in the heterogeneous computing resource visualization are filtered to obtain a front and back collaborative computing group; The core collaborative computing group and the front-end and back-end collaborative computing groups together form the first collaborative computing group.
8. The heterogeneous computing resource scheduling method for intelligent mobile terminals under multi-task concurrency as described in claim 6, characterized in that, The optimization of the candidate collaborative computing groups to obtain the target optimal collaborative computing group includes: Based on the heterogeneous computing resource visibility diagram, the first maximum real-time load and the first average real-time load of the first collaborative computing group are obtained sequentially. Based on the heterogeneous computing resource visibility diagram, the first predicted energy consumption of the first collaborative computing group to complete the target parallel subtask is obtained; The first fitness is obtained by weighting the first maximum real-time load, the first average real-time load, and the first predicted energy consumption. With the goal of maximizing the first fitness, the candidate collaborative computing groups are optimized and screened to obtain the target optimal collaborative computing group.
9. The heterogeneous computing resource scheduling method for intelligent mobile terminals under multi-task concurrency as described in claim 1, characterized in that, The learning machine terminal distributes and executes the target parallel subtasks according to the target optimal collaborative computing group, and then further includes: Obtain the actual execution record of the target parallel subtask; Extract the actual energy consumption from the actual execution record, and dynamically correct the heterogeneous computing resource visualization based on the actual energy consumption.
10. A heterogeneous computing resource scheduling system for intelligent mobile terminals under multi-task concurrency, characterized in that, The heterogeneous computing resource scheduling method for a smart mobile terminal under multi-task concurrency as described in any one of claims 1 to 9, wherein the heterogeneous computing resource scheduling system for the smart mobile terminal under multi-task concurrency comprises: The task parsing module is used to introduce a resource requirement mechanism to parse the concurrent learning task set of the learning machine terminal and obtain the task parsing results; The visualization construction module is used to construct a heterogeneous computing resource visualization based on the heterogeneous computing resource information of the learning machine terminal obtained by dynamic perception. The scheduling analysis module is used to perform coalition scheduling analysis on the target parallel subtasks in the target stage of the task parsing results, using the heterogeneous computing resource visibility as a constraint, to obtain the target optimal cooperative computing group; The task distribution module is used by the learning machine terminal to distribute and execute the target parallel subtasks according to the target optimal collaborative computing group.