Server scheduling method and device, storage medium and program product

By decomposing the computing tasks into subtasks, and using precompiled acceleration template libraries and reconfigurable logic devices to dynamically match and configure hardware acceleration templates and reconstruction areas, the problem of inability to adjust hardware resource configuration in real time in the existing technology is solved, and low-latency and high-performance computing performance is achieved.

CN120066800AActive Publication Date: 2025-05-30INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510529949.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-30
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

In distributed computing scenarios of large-scale artificial intelligence models, the existing technology cannot adjust hardware resource configuration in real time and accurately, resulting in increased computing delay and excessive energy consumption.

Method used

Dynamically match and configure hardware acceleration templates and refactored areas by breaking down operational tasks into multiple subtasks and leveraging precompiled acceleration template libraries and refactorable logic devices for fast and precise resource configuration.

Benefits of technology

It significantly reduces system delay and energy consumption, improves computing performance and resource utilization, and realizes real-time adjustment and rapid update of hardware resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066800A_ABST
    Figure CN120066800A_ABST
Patent Text Reader

Abstract

The invention discloses a server scheduling method and device, a storage medium and a program product, and relates to the technical field of server scheduling, and the method comprises the steps: extracting a hardware acceleration template corresponding to a sub-task in a plurality of sub-tasks from a pre-compiling acceleration template library; a hardware acceleration template stored in the pre-compiling acceleration template library refers to hardware circuit configuration pre-compiled for tasks of different operation types; determining a target reconstruction region corresponding to the sub-task in the plurality of sub-tasks from the plurality of reconstruction regions; according to the hardware acceleration template corresponding to the sub-task in the plurality of sub-tasks, performing reconstruction operation on the target reconstruction area corresponding to the sub-task in the plurality of sub-tasks to obtain a new target reconstruction area corresponding to the sub-task in the plurality of sub-tasks, the problems of increased operation delay and high energy consumption caused by the fact that hardware resource configuration cannot be accurately adjusted in real time in the face of a dynamically changing calculation task in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of server scheduling, and in particular to a server scheduling method, device, storage medium, and program product. Background Art

[0002] In the distributed computing scenario of large-scale artificial intelligence models, the efficient utilization and dynamic adjustment of computing resources have become increasingly important. The reconfigurable technology, as a method that can flexibly adjust the hardware logic to adapt to different computing requirements, realizes the optimization of resource allocation under uninterrupted operation conditions.

[0003] In related technologies, most server systems adopt fixed hardware configurations and relatively simple scheduling algorithms. However, when facing dynamic large-scale model computing tasks, especially when the computing requirements, task nature, or resource status change, related technologies cannot adjust the hardware resource configuration in real time and accurately when dealing with dynamic and variable computing tasks, resulting in deficiencies in load change response and energy consumption control. Summary of the Invention

[0004] The present application provides a server scheduling method, device, storage medium, and program product to at least solve the problems of increased operation delay and high energy consumption caused by the inability to adjust the hardware resource configuration in real time and accurately when related technologies face dynamic computing tasks.

[0005] The present application provides a server scheduling method, including: decomposing an operation task sent by a target object to obtain multiple subtasks; extracting a hardware acceleration template corresponding to a subtask among the multiple subtasks from a pre-compiled acceleration template library; the hardware acceleration templates stored in the pre-compiled acceleration template library refer to the hardware circuit configurations pre-compiled for tasks of different operation types; determining a target reconfiguration area corresponding to a subtask among the multiple subtasks from multiple reconfiguration areas; the multiple reconfiguration areas refer to the logical blocks divided on a reconfigurable logic device that allow independent reconfiguration of the logic circuit; performing a reconfiguration operation on the target reconfiguration area corresponding to a subtask among the multiple subtasks according to the hardware acceleration template corresponding to the subtask among the multiple subtasks to obtain a new target reconfiguration area corresponding to the subtask among the multiple subtasks, and executing the corresponding subtask through the new target reconfiguration area corresponding to the subtask among the multiple subtasks.

[0006] The present application also provides a server scheduling device, including: a task decomposition module, configured to decompose an operation task sent by a target object to obtain a plurality of subtasks; an acceleration template matching module, configured to extract a hardware acceleration template corresponding to a subtask among the plurality of subtasks from a pre-compiled acceleration template library; the hardware acceleration templates stored in the pre-compiled acceleration template library refer to hardware circuit configurations pre-compiled for tasks of different operation types; a reconstruction area determination module, configured to determine a target reconstruction area corresponding to a subtask among the plurality of subtasks from a plurality of reconstruction areas; the plurality of reconstruction areas refer to logical blocks divided on a reconfigurable logic device that allow independent reconfiguration of logic circuits; a task execution module, configured to perform a reconstruction operation on the target reconstruction area corresponding to a subtask among the plurality of subtasks according to the hardware acceleration template corresponding to the subtask among the plurality of subtasks, to obtain a new target reconstruction area corresponding to the subtask among the plurality of subtasks, and execute the corresponding subtask through the new target reconstruction area corresponding to the subtask among the plurality of subtasks.

[0007] The present application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any one of the above server scheduling methods when executing the computer program.

[0008] The present application also provides a computer-readable storage medium, in which a computer program is stored, and wherein the computer program implements the steps of any one of the above server scheduling methods when executed by a processor.

[0009] The present application also provides a computer program product, including a computer program, and the computer program implements the steps of any one of the above server scheduling methods when executed by a processor.

[0010] With this application, the computing tasks are divided into multiple subtasks, and an intelligent matching algorithm is used to dynamically select the most suitable hardware acceleration template and the corresponding target reconstruction area for each subtask. In the case of improving the computing efficiency and reducing the latency, the problem that the static configuration cannot adapt to the dynamically changing computing requirements is solved; the pre-compiled acceleration template library is used to pre-store the hardware acceleration templates optimized in advance for different computing types, so that they can be quickly called when the tasks arrive, avoiding the latency of real-time compilation, improving the matching speed and accuracy of the hardware resources, reducing the latency and energy consumption, and accurately matching each subtask with the corresponding hardware acceleration template. The hardware acceleration template corresponding to each subtask is transmitted to the corresponding reconstruction area through a high-speed configuration link to perform a reconstruction operation on the reconstruction area. Then, the corresponding subtasks are executed based on the new reconstruction area after reconstruction, significantly reducing the system latency and energy consumption, improving the computing performance and resource utilization rate, and allowing partial updates to the selected reconstruction area during operation without restarting, reducing the reconstruction latency, realizing the real-time adjustment and rapid update of the hardware resources, and solving the problem of increased computing latency and high energy consumption caused by the inability to adjust the hardware resource configuration in real time and accurately in the related technologies when facing dynamically changing computing tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0012] Figure 1 FIG. is a schematic diagram of an application scenario of a server scheduling method provided by an embodiment of the present application.

[0013] Figure 2 FIG. is a schematic flowchart of an optional server scheduling method provided by an embodiment of the application.

[0014] Figure 3 FIG. is a structure of a server provided by an embodiment of the application.

[0015] Figure 4 FIG. is a schematic flowchart of a task decomposition provided by an embodiment of the application.

[0016] Figure 5 FIG. is a schematic flowchart of a hardware acceleration template matching provided by an embodiment of the application.

[0017] Figure 6 FIG. is a schematic flowchart of a partial reconstruction and hardware loading provided by an embodiment of the present application.

[0018] Figure 7It is a schematic flowchart of status monitoring and feedback regulation provided by an embodiment of the present application.

[0019] Figure 8 It is a schematic flowchart of a storage module storing task results provided by an embodiment of the present application.

[0020] Figure 9 It is a structural block diagram of an optional server scheduling device provided by an embodiment of the application. Detailed implementation manners

[0021] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0022] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0023] To enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0024] According to one aspect of the embodiments of the present application, a server scheduling method is provided. Optionally, in this embodiment, the above server scheduling method may but is not limited to be applied to a hardware environment including a terminal device 102 and a server 104 as Figure 1 shown. The server 104 can be connected to the terminal device 102 through a network and can be used to provide services (such as application services, etc.) for the terminal device 102 or a client installed on the terminal device 102. A database can be set on the server 104 or independently of the server 104 to provide data storage services for the server 104.

[0025] The above network may include, but is not limited to, at least one of the following: a wired network, a wireless network. The above wired network may include, but is not limited to, at least one of the following: a wide area network, a metropolitan area network, a local area network. The above wireless network may include, but is not limited to, at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The terminal device 102 may be, but is not limited to, a PC (Personal Computer), a mobile phone, a tablet computer, etc. The server 104 may be, but is not limited to, a cloud server, a server cluster, or other server types.

[0026] The server scheduling method of the embodiments of the present application is executed by the server 104. Embodiments of the present application provide a server scheduling method. Figure 2 It is a schematic flowchart of an optional server scheduling method according to an embodiment of the present application, as Figure 2 shown. The process of this method may include steps S202 to S208.

[0027] Step S202: Decompose the computing task sent by the target object to obtain multiple subtasks.

[0028] Among them, the computing task refers to large-scale artificial intelligence model (such as deep neural network, large language model) calculation instructions or data processing requests submitted by the target object or upper-layer application and that need to be executed on the server system. Computing tasks usually have high computational complexity and data throughput requirements, and have strict requirements for the real-time performance and energy consumption control of task execution. Decomposing the computing task can obtain multiple subtasks. Among them, the subtask refers to an independently executable computing unit obtained by decomposing the original computing task. Each subtask is attached with a clear operation type (such as matrix multiplication, convolution operation, activation function, etc.), data block size, real-time requirement, and energy consumption budget for subsequent template matching and resource scheduling.

[0029] Optionally, Figure 3 is a structure of a server provided by an embodiment of the present application, as Figure 3 shown. The internal components of the server include a control and scheduling module. The control and scheduling module includes a dedicated high-performance processor and custom control logic, and integrates a task decomposition engine, a status monitoring unit, a reconstruction decision module, and a template matching module, realizing functions of task decomposition, system status monitoring, reconstruction decision, template matching, and task scheduling. The control and scheduling module dynamically adjusts the resource configuration of each computing unit through real-time monitoring and intelligent scheduling algorithms, optimizes the task decomposition and scheduling strategies, and improves the overall system operation efficiency and stability. Figure 4 It is a schematic flowchart of a task decomposition provided by an embodiment of the present application, as Figure 4As shown in the figure, after the server receives the computing task of the target object, the control and scheduling module parses the running task description. Through the task decomposition engine, a large-scale or complex computing task is decomposed into multiple smaller and more manageable subtasks according to its computing requirements, data dependency relationships, and real-time requirements, and the real-time load and energy consumption data of the current server system are collected.

[0030] Step S204, extract the hardware acceleration template corresponding to the subtask in the multiple subtasks from the pre-compiled acceleration template library; the hardware acceleration templates stored in the pre-compiled acceleration template library refer to the hardware circuit configurations pre-compiled for tasks of different computing types.

[0031] Among them, the pre-compiled acceleration template library refers to a preset hardware circuit configuration database, which stores hardware acceleration templates pre-compiled and optimized for different computing types (such as matrix multiplication, convolution operation, activation function calculation, etc.). The establishment of the pre-compiled acceleration template library is based on a large number of previous experiments and data analysis to ensure that each hardware acceleration template can efficiently execute specific types of computing tasks while meeting the system design goals of low power consumption and low latency. During the operation of the server system, the pre-compiled acceleration template library provides the basic resources required for dynamic matching for the control and scheduling module to achieve fast task-to-hardware mapping.

[0032] The hardware acceleration template refers to the hardware circuit configuration pre-compiled to accelerate specific computing types. After compilation, the hardware acceleration template has been solidified into the configuration data that can be directly loaded into the available reconfiguration area of the reconfigurable logic device, and realizes the hardware acceleration of different computing tasks by quickly retrieving and loading at runtime. In the embodiments of the present application, the hardware acceleration template includes detailed hardware resource allocation (such as the number and layout of DSP (Digital Signal Processor), BRAM (Block RAM), and logic units), circuit design (such as pipeline structure, parallel processing unit), and control logic (such as data flow management, state machine control) information, and can be directly loaded into the reconfiguration area to quickly achieve hardware-level customized acceleration. Each hardware acceleration template is optimized for specific computing requirements, can significantly reduce the computing latency, improve the resource utilization rate, and control the energy consumption at the same time, and is the key component to realize the adaptive optimization of the system.

[0033] Optionally, as Figure 3 shown, the template matching module in the control and scheduling module is used to query the pre-compiled acceleration template library according to the task attributes (such as computing type, computing requirement data, etc.) of each subtask to find the hardware acceleration template that matches each subtask.

[0034] Step S206: Determine the target reconstruction region corresponding to the subtask among multiple subtasks from multiple reconstruction regions; the multiple reconstruction regions refer to the logical blocks divided on the reconfigurable logic device that allow independent reconfiguration of the logic circuit.

[0035] Among them, a reconfigurable logic device is an integrated circuit, which is characterized in that its internal logic circuit can still be reconfigured by software after leaving the factory to adapt to different application requirements. For example, the reconfigurable logic device can be an FPGA (Field-Programmable Gate Array), and as the core computing acceleration module in the server system, the FPGA realizes the acceleration and optimization of large-scale artificial intelligence model computing tasks through dynamic partial reconfiguration technology.

[0036] Inside the reconfigurable logic device, in order to support the dynamic partial reconfiguration technology, the entire reconfigurable logic device is divided into several independent logic regions. Each region contains a certain number of logic resources (such as programmable logic units, digital signal processors, BRAMs, etc.) and has independent configuration capabilities. This region is the reconstruction region. Each reconstruction region has clear boundaries and resource configurations (such as DSP, BRAM, logic units) in hardware. This design allows the system to dynamically adjust the hardware logic within the specified reconstruction region according to task requirements without interrupting the overall operation, so as to achieve the purpose of fast response and resource optimization.

[0037] In the task decomposition and template matching stage, the control and scheduling module determines the reconstruction region that is most suitable for executing a certain subtask from multiple reconstruction regions according to the specific computing requirements of the subtask and the real-time resource status of the reconstruction regions. This region is called the "target reconstruction region", and it will be loaded with a specific hardware acceleration template to execute the corresponding subtask operation, realizing the efficient utilization of computing resources and the acceleration of task execution.

[0038] Optionally, Figure 5 is a schematic flow diagram of hardware acceleration template matching provided by an embodiment of the present application. As Figure 5 shown, the control and scheduling module extracts the hardware acceleration templates matched by each subtask from the pre-compiled acceleration template library, and obtains the detailed parameters and configuration instructions corresponding to each subtask. The reconstruction decision module in the control and scheduling module continuously monitors the current resource status of all reconstruction regions, including resource occupancy, computing power, and reconstruction delay estimation, etc., and performs optimal matching calculations according to the current resource status of each reconstruction region to determine the target reconstruction region corresponding to each subtask.

[0039] Step S208: According to the hardware acceleration templates corresponding to the subtasks among the multiple subtasks, perform a reconstruction operation on the target reconstruction regions corresponding to the subtasks among the multiple subtasks to obtain new target reconstruction regions corresponding to the subtasks among the multiple subtasks, and execute the corresponding subtasks through the new target reconstruction regions corresponding to the subtasks among the multiple subtasks.

[0040] Among them, the reconstruction operation refers to the process of updating or reconfiguring the logic circuit based on a specific hardware acceleration template for the selected reconstruction region inside the reconfigurable logic device. This operation allows for the dynamic adjustment of hardware resources during system operation to adapt to the changing computational task requirements. Through the high-speed configuration link, the control and scheduling module can quickly load the pre-compiled acceleration template into the target reconstruction region, and a partial reconstruction control logic inside the reconfigurable logic device is responsible for completing the update of hardware resources, such as the reallocation of DSP, BRAM, and logic units, within milliseconds, thereby achieving hardware acceleration for computational tasks.

[0041] The new target reconstruction region refers to the reconstruction region in the reconfigurable logic device used to execute a specific subtask after the reconstruction operation. The internal logic circuit of this reconstruction region has been updated according to the corresponding hardware acceleration template to meet the computational requirements of the subtask, such as accelerating computationally intensive operations, optimizing memory bandwidth, or controlling reconstruction latency. Through the intelligent matching algorithm and dynamic partial reconstruction technology, the new target reconstruction region is given the optimal hardware configuration required to execute the subtask, thereby improving the efficiency of task execution and the overall performance of the system.

[0042] As Figure 3 shown, the internal components of the server also include a reconfigurable computing acceleration module, an auxiliary computing module, and a storage module. The reconfigurable computing acceleration module is internally divided into multiple independent reconstruction regions, and each reconstruction region includes fixed interface logic, computing units, and local storage units. The reconfigurable computing acceleration module has a built-in pre-compiled acceleration template library, and each hardware acceleration template is designed for specific neural network operations (matrix multiplication, convolution, activation function calculation, etc.) to ensure meeting the high-efficiency operation requirements. The reconfigurable computing acceleration module uses a high-speed configuration link to support partial reconstruction within the reconstruction region, ensuring hardware logic update without terminating the overall operation of the system. The control and scheduling module and the reconfigurable computing acceleration module use the PCIe high-speed bus for signaling interaction to ensure low-latency transmission of reconstruction instructions and status feedback, and each reconstruction region inside the reconfigurable computing acceleration module is interconnected through an internal bus. The reconfigurable computing acceleration module uses dynamic partial reconstruction technology and the pre-compiled template library to provide dedicated hardware acceleration for large model computing tasks, quickly respond to different workload requirements, and effectively reduce power consumption and latency.

[0043] The auxiliary computing module includes a high-performance processor and a dedicated coprocessor, which are designed for data preprocessing, postprocessing, and auxiliary computing tasks. It provides parallel computing support for the reconfigurable computing acceleration module to ensure the coordinated processing of data streams and computing tasks in large model calculations. The auxiliary computing module is physically connected to the control and scheduling module, the reconfigurable computing acceleration module, and the storage module through high-speed network protocols and independent network interfaces to ensure data consistency and task coordination.

[0044] The storage module integrates NVMe (Non-Volatile Memory Express) solid-state drives and a large-capacity high-speed memory, and is responsible for storing model parameters, intermediate data, configuration files of the reconfigurable computing acceleration module, and log data. High-speed data transmission between the storage module and each module is achieved through independent dedicated data buses, and all physical connections are implemented using welding, slots, and dedicated high-speed cables to achieve data transmission and persistent storage, ensuring that the data read and write rates meet the requirements of real-time computing. The auxiliary computing module and the storage module implement data preprocessing, postprocessing, and high-speed data transmission to ensure the coordinated working ability of the system in high-concurrency and large-data-volume computing scenarios.

[0045] Special communication protocols and DMA (Direct Memory Access) data transmission mechanisms are used to transmit control signals and data between modules. All data interactions use strict timing synchronization mechanisms and verification algorithms to ensure that the system can accurately transmit any control signaling or computing data during continuous operation.

[0046] Optionally, as Figure 3 shown, after the reconfiguration decision module in the control and scheduling module calculates the optimal reconfiguration plan in real time according to a predetermined algorithm, it issues a reconfiguration instruction (a command used to direct the reconfigurable logic device to update the local circuit, including the hardware acceleration template corresponding to the subtask and configuration parameters (such as timing, data width, on-chip storage allocation)) to the reconfigurable computing acceleration module through a high-speed configuration link to instruct the reconfigurable computing acceleration module to perform a reconfiguration operation on the reconfiguration area corresponding to each subtask. After receiving the reconfiguration instruction, the reconfigurable computing acceleration module completes the update of the local logic circuit within milliseconds to obtain a new target reconfiguration area corresponding to each subtask. Among them, a partial reconfiguration control logic built into the reconfigurable computing acceleration module is responsible for applying the new bitstream and updating the corresponding logic, wiring, and storage resources.

[0047] Through the embodiments of the present application, the computing task is divided into multiple subtasks, and an intelligent matching algorithm is used to dynamically select the most suitable hardware acceleration template and the corresponding reconstruction area for each subtask. While improving the computing efficiency and reducing the latency, the problem that the static configuration cannot adapt to the dynamically changing computing requirements is solved; the pre-compiled acceleration template library is used to pre-store the hardware acceleration templates optimized in advance for different computing types, so that they can be quickly called when the task arrives, avoiding the latency of real-time compilation, improving the matching speed and accuracy of hardware resources, reducing the latency and energy consumption, and accurately matching each subtask with the corresponding hardware acceleration template. The hardware acceleration template corresponding to each subtask is transmitted to the corresponding reconstruction area through a high-speed configuration link to perform a reconstruction operation on the reconstruction area. Then, the corresponding subtask is executed based on the new reconstructed area after reconstruction, significantly reducing the system latency and energy consumption, improving the computing performance and resource utilization rate, and allowing partial updates to the selected reconstruction area during operation without restarting, reducing the reconstruction latency, realizing the real-time adjustment and rapid update of hardware resources, and solving the problem of increased computing latency and high energy consumption caused by the inability to adjust the hardware resource configuration in real time and accurately in the related technology when facing dynamically changing computing tasks.

[0048] In an exemplary embodiment, the hardware acceleration templates stored in the pre-compiled acceleration template library include the configuration descriptions of tasks of different computing types.

[0049] The configuration description of the hardware acceleration template is the instruction set for the circuit design of the reconfigurable logic device, used to quickly configure the reconfigurable logic device to execute tasks of a specific computing type. The configuration description details how to achieve efficient hardware acceleration in the reconstruction area of the reconfigurable logic device, including the allocation and usage methods of resources such as logic units, DSPs, and BRAMs, as well as the wiring details of the circuit. In this way, when the system receives a task of the corresponding computing type, it can directly call the corresponding hardware acceleration template from the pre-compiled acceleration template library for local reconstruction without time-consuming real-time compilation, thus significantly improving the system response speed and resource utilization efficiency.

[0050] In some embodiments, extracting the hardware acceleration template corresponding to the subtask among the multiple subtasks from the pre-compiled acceleration template library includes:

[0051] Determining the hardware acceleration template in the pre-compiled acceleration template library that matches the computing type of the subtask among the multiple subtasks as the hardware acceleration template corresponding to the subtask among the multiple subtasks.

[0052] Among them, the operation type refers to the specific mathematical operations or processing flows in the large model computing tasks, such as matrix multiplication, convolution operation, activation function calculation, etc. The operation type defines the computing characteristics of the subtasks and is the basis for selecting the hardware acceleration templates in the pre-compiled acceleration template library. By identifying the operation type of each subtask, the most suitable hardware circuit configuration can be accurately matched to achieve the acceleration and optimization of specific operations, thereby improving the overall computing efficiency and reducing energy consumption and latency.

[0053] Optionally, the template matching module in the control and scheduling module extracts the hardware acceleration template that matches the operation type of the subtask from the pre-compiled acceleration template library for each subtask and determines it as the hardware acceleration template corresponding to the subtask.

[0054] Through this embodiment, matching the operation type of the subtask with the hardware acceleration template can ensure that the selected hardware acceleration template best meets the actual computing requirements of the subtask, avoiding waste of resources and unnecessary time delay.

[0055] In an exemplary embodiment, determining the target reconstruction region corresponding to the subtask among multiple subtasks from multiple reconstruction regions includes:

[0056] Performing matching calculations on the subtask among multiple subtasks according to the calculation requirement data of the subtask among multiple subtasks and the current resource status of multiple reconstruction regions to determine the target reconstruction region corresponding to the subtask among multiple subtasks.

[0057] Among them, the calculation requirement data refers to the specific hardware computing resources and characteristics required by each subtask during execution, including but not limited to computing power (such as floating-point operations per second FPoS), memory bandwidth or capacity, task priority, expected running time, energy consumption budget, etc. The calculation requirement data is parsed and generated by the control and scheduling module during the task decomposition stage and is used to guide the selection of hardware acceleration templates and the matching of reconstruction regions to ensure that the task can efficiently utilize hardware resources while meeting performance requirements and reducing energy consumption.

[0058] The current resource status of the reconstruction region refers to the actual available resource situation and running state of the reconstruction region at the current moment, which is fed back to the control and scheduling module by the internal monitoring unit and status register of the reconfigurable computing acceleration module, etc., and is used for dynamic decision-making of subtask matching and resource scheduling to ensure smooth system operation, reasonable resource allocation, efficient task execution and controllable energy consumption. The current resource status of the reconstruction region usually includes whether the current reconstruction region is idle or executing a task, the remaining available resources (such as logic units, DSPs, BRAMs, etc.), whether the region allows local reconstruction, the current power consumption or temperature, etc.

[0059] Optionally, Figure 5It is a schematic flowchart of a hardware-accelerated template matching provided by an embodiment of the present application. As Figure 5 shown, the control and scheduling module extracts the hardware-accelerated templates matching each subtask from the pre-compiled acceleration template library, and obtains the detailed parameters and configuration descriptions corresponding to each subtask. The reconstruction decision module in the control and scheduling module continuously monitors the current resource status of all reconstruction areas, including resource occupancy, computing power, and reconstruction delay estimation, etc. The reconstruction decision module analyzes the computing requirement data of each subtask, including the computing requirements, memory requirements, and latency sensitivity of each subtask, performs an optimal matching calculation on the computing requirement data of each subtask and the current resource status of each reconstruction area, selects the reconstruction area with the highest matching degree with the subtask as the target reconstruction area corresponding to the subtask, and issues the matching result and reconstruction instruction to the reconfigurable computing acceleration module.

[0060] Through this embodiment, the computing requirement data of the subtasks is matched with the current resource status of the reconstruction areas, ensuring that each subtask can be allocated the most suitable hardware resources for its needs, avoiding overabundance or shortage of resources, thereby improving resource utilization rate, reducing operation latency and energy consumption; by considering the specific requirements of the subtasks for computing resources and the current status of the reconstruction areas, the energy consumption of each reconstruction operation can be predicted and controlled more accurately, while ensuring that the task performance is not affected. This data-driven matching strategy enables the system to automatically optimize resource allocation when processing high-load and variable computing tasks, achieving the best balance between energy consumption and performance.

[0061] In an exemplary embodiment, according to the computing requirement data of the subtasks in multiple subtasks and the current resource status of multiple reconstruction areas, a matching calculation is performed on the subtasks in multiple subtasks to determine the target reconstruction area corresponding to the subtasks in multiple subtasks, including:

[0062] Based on the computing requirement data of the subtasks in multiple subtasks, a task requirement vector of the subtasks in multiple subtasks is generated; based on the current resource status of multiple reconstruction areas, multiple regional resource vectors are generated; the multiple regional resource vectors correspond to the multiple reconstruction areas one by one; calculate the distances between the task requirement vectors of the subtasks in multiple subtasks and the multiple regional resource vectors, and determine the reconstruction area corresponding to the minimum distance as the target reconstruction area corresponding to the subtasks in multiple subtasks.

[0063] Among them, the task requirement vector corresponding to each subtask is a multi-dimensional vector composed of specific attributes of the corresponding subtask. For example, the task requirement vector of each subtask includes the type of calculation (such as matrix multiplication, convolution operation, etc.), the required computing power (such as floating-point operations per second), the memory requirement (bandwidth or capacity), the latency sensitivity, and the energy consumption budget, etc. The task requirement vector of each subtask comprehensively reflects the hardware resource requirements of the subtask and is used to guide the intelligent matching algorithm to select the most suitable reconstruction area.

[0064] The area resource vector corresponding to each reconstruction area represents a multi-dimensional vector of the current available resource status of each independent reconstruction area within the reconfigurable computing acceleration module. For example, the area resource vector corresponding to each reconstruction area covers the utilization rate of computing units (such as DSP, BRAM), the remaining logic resources, the estimated current reconstruction latency, power consumption, and temperature, etc. It provides real-time resource information for the control and scheduling module to calculate the matching degree with the task requirement vector, so as to determine the reconstruction strategy.

[0065] The distance between the task requirement vector corresponding to each subtask and the area resource vector corresponding to each reconstruction area refers to the difference metric between the task requirement vector and the area resource vector. The distance reflects the matching degree between the task requirements and the current resource status of the reconstruction area. The smaller the distance, the more the resource status can meet the task requirements, thus determining which reconstruction area the subtask should be assigned to for execution to achieve the best matching of resources and tasks.

[0066] Optionally, the control and scheduling module parses the calculation requirement data of each subtask and converts it into a multi-dimensional task requirement vector. The state detection unit of the reconfigurable computing acceleration module real-time collects the current resource status of each reconstruction area to form an area resource vector corresponding to each reconstruction area one by one. The control and scheduling module calculates the distance between the task requirement vector of each subtask and the area resource vectors of all reconstruction areas. The control and scheduling module identifies the reconstruction area with the minimum distance and takes it as the target reconstruction area of the subtask.

[0067] Through this embodiment, the concepts of task requirement vector and reconstruction area resource vector are introduced, covering multiple aspects of attributes, forming a comprehensive mapping relationship between resource requirements and supply, laying a foundation for the efficient allocation of dynamic tasks; calculating the distance between the task requirement vector of a subtask in multiple subtasks and the area resource vectors of multiple areas, and determining the reconstruction area corresponding to the minimum distance as the target reconstruction area corresponding to the subtask in multiple subtasks. This distance calculation method takes into account multiple dimensions, ensuring the accuracy and efficiency of resource allocation.

[0068] In an exemplary embodiment, the calculation requirement data of a subtask in multiple subtasks includes the expected computing power, the expected memory capacity, and the upper limit of the expected reconstruction latency of the subtask in multiple subtasks.

[0069] Among them, the expected computing power of a subtask reflects the requirement of the subtask for the operation speed, such as the number of floating-point operations per second. The expected memory capacity refers to the size of the storage space required for the subtask to process data (unit: GB / s), which is used to ensure the fast access and caching of data. The upper limit of the expected reconstruction delay defines the maximum allowable time for the subtask to wait for the reconstruction to complete before starting to execute, so as to meet the real-time requirements.

[0070] In some embodiments, based on the computing requirement data of the subtasks among multiple subtasks, a task requirement vector of the subtasks among multiple subtasks is generated, including:

[0071] Generating a first task element based on the expected computing power of the subtasks among multiple subtasks; generating a second task element based on the expected memory capacity of the subtasks among multiple subtasks; generating a third task element based on the upper limit of the expected reconstruction delay of the subtasks among multiple subtasks; and generating a task requirement vector of the subtasks among multiple subtasks according to the first task element, the second task element, and the third task element.

[0072] Among them, the first task element refers to the task attribute value generated based on the expected computing power of the subtask (such as the number of floating-point operations per second), which reflects the specific requirements of the subtask for the processing speed and computing resources and is used for the construction of the task requirement vector.

[0073] The second task element refers to the task attribute value generated according to the expected memory capacity of the subtask (such as the number of GBs required to store data), which reflects the demand of the subtask for memory resources during the execution process and is an important part of the task requirement vector, used to match the reconstruction area with sufficient memory resources to ensure the efficiency of data processing.

[0074] The third task element refers to the task attribute value generated by the upper limit of the expected reconstruction delay of the subtask (such as the maximum allowable waiting time), which ensures that the reconfigurable logic device performs necessary local reconstruction operations while meeting the real-time requirements of the task and is one of the key parameters for constructing the task requirement vector, used to avoid additional delays caused by waiting for reconstruction.

[0075] Optionally, the control and scheduling module evaluates the computing requirements of the subtask, extracts its expected computing power index, quantifies it into the first task element; analyzes the memory occupation situation of the subtask, extracts its expected memory capacity, and quantifies it into the second task element; pre-computes the maximum allowable waiting time (i.e., the upper limit of the expected reconstruction delay) of the subtask, and quantifies it into the third task element. The control and scheduling module combines the first task element, the second task element, and the third task element to generate a complete task requirement vector, providing an accurate parameter basis for subsequent resource matching.

[0076] For example, define a subtask The corresponding task requirement vector is the following formula (1):

[0077] (1)

[0078] Wherein, represents the task corresponding task requirement vector; represents the first task element, i.e., the expected computing power (floating-point operations per second) of task ; represents the second task element, i.e., the expected memory capacity (GB / s) of task ; represents the third task element, i.e., the upper limit of the expected reconstruction delay (in milliseconds) of task .

[0079] Through this embodiment, by quantifying the expected computing power of subtasks, the specific requirements of each subtask for computing resources can be accurately grasped, ensuring that sufficient hardware support can be provided for compute-intensive jobs at the initial stage of task scheduling, avoiding task delay phenomena caused by insufficient computing resources, and thus improving the overall computing efficiency and response speed of the system; by quantifying the expected memory capacity of subtasks, an accurate assessment of the memory resources required by subtasks can be made. When dealing with large-scale models, it ensures the fast access and caching of data, reduces unnecessary data transfer delays, and further improves the task processing speed and system throughput: by introducing the upper limit of the expected reconstruction delay of subtasks as part of the task requirement vector, by estimating the time for tasks to wait for the partial reconstruction of reconfigurable logic devices to complete, those computing tasks that may not be able to respond in time due to long reconstruction cycles can be avoided in advance, ensuring the real-time nature of tasks and the stability of the system, and reducing the additional system burden brought by frequent reconstruction; by combining the above three key task elements to generate a task requirement vector and expressing the comprehensive requirements of each subtask in vector form, the compatibility between subtasks and the reconfiguration area can be accurately evaluated, thus realizing the effective allocation of resources. It not only considers computing power and memory resources, but also pays special attention to reconstruction delay, ensuring the low-latency characteristics of task execution, while optimizing resource utilization, reducing energy consumption, and overcoming the performance bottleneck and high energy consumption problems during large-scale model deployment.

[0080] In an exemplary embodiment, the current resource status of the reconfiguration areas in the multiple reconfiguration areas includes the current computing power, current memory capacity, and current reconstruction delay estimate value of the reconfiguration areas in the multiple reconfiguration areas.

[0081] Among them, the current computing power refers to the computing and processing capabilities that the reconstruction area can provide at a certain moment, including resources such as available DSP modules, logic units, and BRAM units, which are used to evaluate the support degree of this area for computationally intensive subtasks, ensuring that real-time computing resource availability is considered during task allocation.

[0082] The current memory capacity represents the amount of storage resources currently available in the reconstruction area, including the remaining space in on-chip memories (such as BRAMs) and caches, which is used to measure data processing and storage capabilities, ensuring that subtasks can quickly access the required memory resources during execution, avoiding data latency and task interruption caused by insufficient memory.

[0083] The current reconstruction delay estimate reflects the time estimate required for the reconstruction area to perform partial logic updates in the current state, which is used to evaluate the impact of the reconstruction operation on the execution delay of subtasks, ensuring that the timeliness of reconstruction is fully considered during task scheduling, and avoiding task waiting and system efficiency degradation caused by too long reconstruction delay.

[0084] In some embodiments, based on the current resource states of multiple reconstruction areas, multiple area resource vectors are generated, including:

[0085] Taking the reconstruction area in the multiple reconstruction areas as the current reconstruction area, performing the following vector generation operations to obtain multiple area resource vectors:

[0086] Generating a first area element based on the current computing power of the current reconstruction area; generating a second area element based on the current memory capacity of the current reconstruction area; generating a third area element based on the current reconstruction delay estimate of the current reconstruction area; generating the area resource vector corresponding to the current reconstruction area according to the first area element, the second area element, and the third area element.

[0087] Among them, the first area element represents the computing power state of the current reconstruction area. By quantifying the current available computing resources such as DSP units, BRAM modules, and logic units in this area, the first area element is generated. The first area element reflects the immediate computing and processing capabilities of the current reconstruction area, which is used to match the computing requirements of subtasks, ensuring the efficiency and timeliness of task execution.

[0088] The second area element is generated based on the memory capacity state of the current reconstruction area. By quantifying the remaining BRAM capacity, on-chip storage resources, and external memory bandwidth, etc., the second area element is used to evaluate whether the reconstruction area can meet the data storage and transmission requirements of subtasks, ensuring the coherence and reliability of data processing.

[0089] The third region element represents the estimated reconstruction delay value of the current reconstruction region, which is generated by predicting the time required for local logic updates. The third region element ensures that the maximum waiting time allowed for subtasks is not exceeded during the dynamic resource adjustment process, thus avoiding task delays and system performance degradation.

[0090] The region resource vector is composed of the first region element (computing power), the second region element (memory capacity), and the third region element (estimated reconstruction delay value), comprehensively describing the resource status and characteristics of the current reconstruction region. The region resource vector is used for resource matching calculations within the system to ensure that each subtask can be assigned to the reconstruction region that best matches its requirements, thereby achieving efficient resource utilization and fast task response, and reducing the overall energy consumption and delay of the system.

[0091] Optionally, the control and scheduling module selects a reconstruction region from multiple reconstruction regions as the current reconstruction region, queries the DSP, BRAM, and logic resource status of each reconstruction region, quantifies it as the current computing power metric, and generates the first region element; through the internal monitoring mechanism, obtains the current memory usage of the reconstruction region, converts it into the second region element; uses the historical data of the reconstruction region and the current task queue status to predict the time required to execute partial reconstruction, that is, the current estimated reconstruction delay value, to form the third region element; fuses the first, second, and third region elements, and uses the existing vector construction algorithm to generate a complete region resource vector representing the resource status of the current reconstruction region.

[0092] For example, for each reconstruction region , its corresponding region resource vector is defined by the following formula (2):

[0093] (2)

[0094] Where represents the region resource vector corresponding to the reconstruction region ; represents the first region element, that is, the currently available computing power within the reconstruction region ; represents the second region element, that is, the current memory resources within the reconstruction region ; represents the third region element, that is, the current estimated reconstruction delay value of the reconstruction region .

[0095] Through this embodiment, by monitoring and quantifying the current computing power of the current reconstruction area in real time to generate the first area element, which can immediately reflect the operation and processing potential of this area, enabling the system to perform intelligent matching based on the most real-time computing resource data in the current task scheduling instead of fixed preset values, ensuring the maximum support for compute-intensive tasks, reducing the task queuing waiting time, and improving the overall computing efficiency and system response speed; converting the current memory capacity of the current reconstruction area into the second area element, enabling the system to accurately grasp the actual requirements of data storage and transmission, solving the problem of unreasonable memory resource allocation, ensuring the fluidity and integrity of data during processing, avoiding task processing delays caused by memory bottlenecks, and simultaneously increasing the data processing speed and the overall throughput of the system; generating the third area element based on the current reconstruction delay estimate value, which can avoid allocating delay-sensitive tasks to areas that are in reconstruction or about to be reconstructed, thus significantly reducing the additional waiting time of tasks and ensuring the real-time performance of task execution and the quality of user experience; synthesizing the above three area elements into an area resource vector to achieve a comprehensive consideration and dynamic adjustment of the resource status. This real-time resource monitoring and dynamic matching down to each reconstruction area greatly improves the utilization rate of hardware resources, while ensuring the intelligence and efficiency of task scheduling, effectively coping with the variability of system load, reducing energy consumption, and achieving precise allocation and optimal scheduling of computing resources.

[0096] In one exemplary embodiment, calculating the distance between the task demand vector of a subtask among multiple subtasks and multiple area resource vectors includes:

[0097] Using the Euclidean distance calculation formula to calculate the distance between the task demand vector of a subtask among multiple subtasks and multiple area resource vectors; the Euclidean distance calculation formula characterizes the difference between corresponding elements in two vectors, as well as the correlation relationship between the difference, the normalization factor, and the index weight. The selection of the normalization factor is based on previous experimental data and system configuration to ensure that the values are calculated within the same magnitude.

[0098] Among them, the normalization factor is a set of values used to eliminate the dimensional differences of each element in the task demand vector and the area resource vector, ensuring that all indicators are compared on the same scale when calculating the Euclidean distance.

[0099] The index weight refers to assigning different weights to each element in the task demand vector and the area resource vector, reflecting the relative importance of each indicator to the matching result. The weight setting takes into account the key roles of computing power, memory capacity, and reconstruction delay in large model calculations, making the distance calculation more in line with actual requirements and optimizing resource allocation.

[0100] If the two point vectors are respectively and , the traditional Euclidean distance calculation is as shown in the following formula (3):

[0101] (3)

[0102] The traditional Euclidean distance calculation is applicable to the comparison of data with the same dimension. It simplifies the calculation process, eliminates the need for additional normalization steps, and does not consider the importance differences of different dimensions. Moreover, the Euclidean distance is a natural distance metric based on geometric intuition and linear algebra principles, which assumes that each dimension in space contributes equally to the distance. Therefore, the traditional Euclidean distance calculation formula does not contain a normalization factor and index weights. If there are significant dimension differences in different dimensions of a vector (for example, one dimension is in centimeters and the other is in kilometers), then the larger dimension will overly dominate the distance calculation, and the subtle changes in the smaller dimension may be ignored, resulting in unfair or inaccurate matching results. Additionally, in actual application scenarios, different indicators do not contribute equally to the overall similarity or matching degree. For example, in resource matching, computing power may be more important than memory capacity. Without weighting, the Euclidean distance calculation formula cannot reflect this priority and may make suboptimal matching decisions.

[0103] Therefore, to solve the above problems, in this embodiment, the traditional Euclidean distance calculation formula is optimized by adding a normalization factor and index weights. The normalization factor ensures that the values of all dimensions are compared within the same scale, eliminates the influence brought by dimension differences, and makes the distance calculation more fair and accurate. By setting weights for each dimension, the importance of different indicators can be distinguished according to the actual situation, ensuring that resource matching and task scheduling are more reasonable and reflecting the priority of actual business requirements.

[0104] Through this embodiment, calculating the normalized weighted distance between the task demand vector and multiple regional resource vectors can intelligently identify the most matching reconstruction area, thereby realizing the precise scheduling of computing resources. This method not only considers the linear distance between task requirements and resource status, but also enhances the accuracy and practicality of distance calculation through weights and normalization factors, effectively solving the problem of low matching degree between resource scheduling strategies and actual computing requirements.

[0105] In an exemplary embodiment, the task demand vector of the subtasks in multiple subtasks includes a first task element, a second task element, and a third task element; the regional resource vector of the reconstruction areas in multiple reconstruction areas includes a first regional element, a second regional element, and a third regional element; the Euclidean distance calculation formula includes a first normalization factor, a second normalization factor, a third normalization factor, a first index weight, a second index weight, and a third index weight.

[0106] Among them, the first normalization factor refers to a constant used to normalize the computational power requirement of a subtask, ensuring the dimensional consistency of the difference between the first task element (computational power requirement) of the task requirement vector and the first region element (computational power provision) of the region resource vector, thereby making the consideration of computational power in task-resource matching more fair and accurate.

[0107] The second normalization factor refers to a constant used to normalize the memory bandwidth or capacity requirement of a subtask, ensuring the standardization of the difference between the second task element (memory bandwidth or capacity requirement) and the second region element (memory bandwidth or capacity provision), making the representation of memory resources in distance calculation of the same order of magnitude as computational power and reconstruction latency, and enhancing the comprehensiveness of resource matching.

[0108] The third normalization factor refers to a constant used to normalize the upper limit requirement of the reconstruction latency of a subtask, ensuring the standardization of the difference between the third task element (reconstruction latency requirement) and the third region element (estimated reconstruction latency), so that the impact of reconstruction latency in task-resource matching can be reasonably reflected.

[0109] The first metric weight is used to calculate the relative importance of computational power in the matching decision between the task requirement vector and the region resource vector. Its value reflects the priority of computational power compared to other resources (such as memory and reconstruction latency) in a specific scenario. Through weighted processing, the first metric weight ensures the appropriate consideration of computational power in resource matching.

[0110] The second metric weight refers to the weighting coefficient of memory bandwidth or capacity in the task-resource matching decision. Its value reflects the contribution degree of memory resources to task execution efficiency and stability.

[0111] The third metric weight refers to the importance weight of reconstruction latency in task requirement-resource matching. By quantifying the correlation between reconstruction latency and task execution efficiency, the third metric weight ensures the reasonable consideration of reconstruction latency requirements in the resource matching algorithm. Especially in scenarios that require quick response, the weight of reconstruction latency will be increased to reduce the waiting time of task execution.

[0112] In some embodiments, the task requirement vector of a subtask among multiple subtasks is used as the current task requirement vector, and the region resource vector among multiple region resource vectors is used as the current region resource vector. A distance calculation operation is performed to obtain the distance between the task requirement vector of a subtask among multiple subtasks and the multiple region resource vectors:

[0113] 1. Calculate the first difference between the first task element of the current task requirement vector and the first region element of the current region resource vector, the second difference between the second task element of the current task requirement vector and the second region element of the current region resource vector, and the third difference between the third task element of the current task requirement vector and the third region element of the current region resource vector.

[0114] 2. Determine the first squared difference according to the first ratio between the first difference and the first normalization factor; determine the second squared difference according to the second ratio between the second difference and the second normalization factor; determine the third squared difference according to the third ratio between the third difference and the third normalization factor.

[0115] 3. Weight the first squared difference with the first index weight to obtain the first weighted value; weight the second squared difference with the second index weight to obtain the second weighted value; weight the third squared difference with the third index weight to obtain the third weighted value.

[0116] 4. Take the square root of the sum of the first weighted value, the second weighted value, and the third weighted value to determine the distance between the current task requirement vector and the current region resource vector.

[0117] Among them, as can be seen from the above embodiments, represents the first task element, represents the second task element, represents the third task element, represents the first region element, represents the second region element, represents the third region element. In this embodiment, the first normalization factor is defined as the second normalization factor is , the third normalization factor is , the first index weight is defined as , the second index weight is , the third index weight is , the current task requirement vector and the distance between the current region resource vector is shown in the following formula (4):

[0118] (4)

[0119] Where: .

[0120] After calculating the subtask After calculating the distances from all regional resource vectors, among all available reconstruction regions, the control scheduling module selects the reconstruction region with the minimum distance as the best matching region (i.e., the target reconstruction region) corresponding to each subtask. The corresponding mathematical expression is shown in formula (5) below:

[0121] (5)

[0122] where, represents the number of the target reconstruction region corresponding to the subtask ; This decision-making process ensures that the selected region best matches the task requirements in terms of computing power, memory resources, and reconstruction latency, achieving optimal resource allocation and overall system performance optimization.

[0123] In this embodiment, the differences between the corresponding elements in the two vectors are normalized using the first normalization factor, the second normalization factor, and the third normalization factor, eliminating the dimensional differences between different metrics (such as computing power, memory resources, and reconstruction latency), ensuring that each metric can be fairly compared when calculating the distance, and solving the matching deviation problem caused by inconsistent dimensions in different dimensions in traditional Euclidean distance calculation; Using the first metric weight, the second metric weight, and the third metric weight allows adjusting the relative importance of each metric in distance calculation according to specific scenarios and task requirements. This weighting mechanism enables the algorithm to more intelligently evaluate the resource matching degree, prioritize the metrics that have a greater impact on task execution efficiency and system performance, thereby optimizing resource allocation, improving computing efficiency, and response speed; After processing multiple resource attributes (computing power, memory resources, reconstruction latency) through difference, normalization, and weighting, a comprehensive evaluation is carried out, which can comprehensively reflect the matching degree between task requirements and the regional resources of the reconfigurable logic device, avoiding a single resource attribute dominating the task scheduling decision, and realizing multi-dimensional optimization of resource allocation.

[0124] In an exemplary embodiment, the normalization factor and the metric weight can be dynamically adjusted according to the system state and task requirements, enabling the distance calculation method to flexibly adapt to the demand changes in different scenarios, and thus achieving more accurate task scheduling and resource matching.

[0125] For example, input the task requirement vectors corresponding to each subtask and the multiple regional resource vectors corresponding to all reconstruction regions into a pre-trained neural network to obtain the predicted normalization factor and predicted metric weight corresponding to each subtask. Among them, the task requirement vector corresponding to each subtask includes a first normalization factor, a second normalization factor, and a third normalization factor; the predicted metric weight corresponding to each subtask includes a first metric weight, a second metric weight, and a third metric weight; the predicted normalization factor and predicted metric weight corresponding to each subtask are used to calculate the normalized weighted Euclidean distance between the task requirement vector and the regional resource vector corresponding to each subtask.

[0126] Among them, the neural network is based on an adaptive learning mechanism and automatically adjusts the normalization factor and weight parameters according to historical operation data and patterns. The neural network can learn the resource consumption laws of different types of tasks, and can predict the predicted normalization factor and predicted metric weight applicable to the current scenario according to the current system state and task characteristics, improving the matching accuracy and system response speed.

[0127] In this embodiment, the neural network uses machine learning techniques, such as reinforcement learning, to learn historical data and train a neural network that can predict the resource requirements of different types of tasks. Specifically, obtain multiple training samples. Among them, each training sample includes the operation state data (computing load, memory usage, energy consumption, reconstruction delay, etc.) of each reconstruction region when performing different tasks and label data. The label data of each training sample is used to label the historical normalization factor and historical metric weight corresponding to the operation state data in each training sample. Use multiple training samples to train an untrained neural network to obtain a trained neural network. Among them, during the training process, the neural network outputs the predicted result corresponding to each training sample. The predicted result corresponding to each training sample includes the predicted normalization factor and predicted metric weight corresponding to each training sample, and updates the model parameters of the neural network according to the difference between the predicted result of each training sample and the label data until the iteration stop condition is met, and a trained neural network is obtained.

[0128] Through this embodiment, take the task requirement vector of each subtask and each regional resource vector as inputs and send them into a pre-trained neural network model to predict the normalization factor and metric weight of each subtask. Utilize the powerful pattern recognition and learning ability of the neural network to automatically adjust the algorithm parameters according to historical operation data and the real-time system state, improving the accuracy and flexibility of task matching; the predicted normalization factor and predicted metric weight can adaptively reflect the resource requirements and system state in the current scenario, ensuring that the distance calculation method can accurately evaluate the matching degree of task requirements and hardware resources, and effectively solving the problem in the related technology that static parameter settings cannot adapt to dynamic resource requirement changes.

[0129] In an exemplary embodiment, before performing a reconstruction operation on a target reconstruction area corresponding to a subtask in the multiple subtasks according to a hardware acceleration template corresponding to the subtask in the multiple subtasks, the server scheduling method further includes:

[0130] Preliminary inspection is performed on target reconstruction areas corresponding to subtasks among multiple subtasks; when there is no real-time data transmission task in the target reconstruction areas corresponding to subtasks among multiple subtasks, the steps of performing reconstruction operations on the target reconstruction areas corresponding to subtasks among multiple subtasks according to the hardware acceleration templates corresponding to the subtasks among the multiple subtasks.

[0131] In an exemplary embodiment, when a real-time data transmission task exists in a target reconstruction area corresponding to a subtask among multiple subtasks, the real-time data transmission task is migrated to another reconstruction area, or the real-time data transmission task is waited for to be completed.

[0132] Among them, after the reconfigurable computing acceleration module receives the reconstruction instruction sent by the control scheduling module, before the reconstruction process starts, the state detection unit solidified in the reconfigurable computing acceleration module pre-checks the area to be updated. Pre-check refers to a series of state verification operations performed by the internal state detection unit of the reconfigurable computing acceleration module on the designated reconstruction area before the hardware update of the reconstruction area. Its main purpose is to confirm that there is no real-time data transmission or computing task in progress in the target reconstruction area, to ensure that the reconstruction operation will not interrupt the ongoing data processing flow or computing task, and avoid data loss and calculation errors. The pre-check mechanism includes but is not limited to checking the current task status, data transmission status and internal resource usage of the reconstruction area. Once it is confirmed that the area is idle, the reconstruction operation can be safely performed and the new hardware acceleration template can be loaded, thereby realizing dynamic resource adjustment and efficient utilization of reconfigurable logic devices.

[0133] Optionally, Figure 6 is a schematic diagram of a partial reconstruction and hardware loading process provided by an embodiment of the present application, such as Figure 6 As shown, after the reconfigurable computing acceleration module receives the reconstruction instruction sent by the control scheduling module, the state detection unit solidified in the reconfigurable computing acceleration module pre-checks the target reconstruction area indicated by the reconstruction instruction. When the target reconstruction area indicated by the reconstruction instruction has no real-time data transmission task, the hardware acceleration template corresponding to the target reconstruction area is transmitted to the target reconstruction area through the high-speed configuration link; the reconfigurable computing acceleration module completes the local logic circuit update within milliseconds to obtain a new target reconstruction area corresponding to each subtask. When the target reconstruction area indicated by the reconstruction instruction has a real-time data transmission task, the real-time data transmission task is migrated to other reconstruction areas, or the real-time data transmission task is waited for to be completed, and then the reconstruction is performed.

[0134] As Figure 6 shown, after the update is completed, the status feedback unit immediately transmits the reconstruction success signal to the control and scheduling module to ensure that subsequent task scheduling depends on the latest hardware status. After the reconstruction operation is completed, the reconfigurable computing acceleration module executes the corresponding subtasks according to the new target reconstruction area corresponding to each subtask. At the same time, the auxiliary computing module starts the data preprocessing process, formats and caches the original input data (referring to the initial data required by the large model or subtasks) according to the task requirements. High-speed data transmission between the reconfigurable computing acceleration module and the auxiliary computing module is achieved through the DMA technology. Specifically, the auxiliary computing module preprocesses, formats, and caches the original data according to the task requirements; transmits these processed data to the target reconstruction area for calculation at high speed through the DMA; after the calculation is completed, the reconfigurable computing acceleration module can send the intermediate results or final results back to the auxiliary computing module or the storage module for caching through the DMA or a high-speed interface.

[0135] Through this embodiment, before the hardware update of the reconstruction area, the target reconstruction area is pre-checked first to confirm that there is no real-time data transmission or computing task in this area, solving the problems of data loss and operation interruption that may be caused during the dynamic reconstruction process, and ensuring the continuity of tasks and the integrity of data; when the pre-check finds that the target reconstruction area is executing a real-time data transmission task, a task migration or waiting strategy is adopted, which can intelligently migrate the current data transmission task to other idle reconstruction areas, or wait until the current task is completed before performing the reconstruction, solving the conflict between dynamic reconstruction and task execution in complex and changeable computing scenarios, and ensuring the timely response of tasks and the efficient utilization of hardware resources.

[0136] In an exemplary embodiment, the above server scheduling method further includes:

[0137] Real-time monitoring of the operating status data of the reconfigurable logic device; in the case of the first reconstruction area where the operating status data indicates an abnormality, using a spare hardware acceleration template to perform a reconstruction operation on the first reconstruction area; the spare hardware acceleration template is compatible with the hardware characteristics of multiple reconstruction areas.

[0138] Among them, the operating status data refers to various index data reflecting the current working status and performance of the reconstruction area on the reconfigurable logic device, including but not limited to computing load, energy consumption data, signal transmission status, and reconstruction effect, etc.

[0139] The first reconstruction area refers to the reconstruction area where the monitored status deviates from the normal range, which may show abnormal conditions such as overheating, decreased computing performance, data transmission errors, faults, or failure to meet real-time requirements, etc.

[0140] A spare hardware acceleration template refers to a set of pre-compiled hardware circuit configuration options. The spare hardware acceleration template is compatible with the hardware characteristics of multiple reconfiguration regions, that is, when designing the spare hardware acceleration template, the hardware architecture characteristics of different reconfiguration regions are fully considered, such as the number and layout of logic units, storage resources (BRAM, URAM), and digital signal processing units (DSP). This compatibility ensures that the spare hardware acceleration template can operate effectively in the reconfiguration regions with various hardware configurations, improving the robustness and scalability in dealing with hardware diversity. The spare hardware acceleration template covers multiple operation types and is enabled when the original template fails or an exception occurs in the reconfiguration region. The spare hardware acceleration template fully considers the generality and flexibility of hardware resources, ensuring that even when problems occur in some reconfiguration regions, it can still respond quickly and find a suitable alternative solution to ensure that the task execution is not affected.

[0141] Optionally, Figure 7 is a schematic flowchart of a state monitoring and feedback regulation provided by an embodiment of the present application. As Figure 7 shown, the control and scheduling module maintains real-time communication with each module throughout the task execution, monitors the computing load, energy consumption data, signal transmission status, and reconfiguration effect of each module. In the case of a first reconfiguration region where the running state data indicates an anomaly, the control and scheduling module retrieves the spare hardware acceleration template from the pre-compiled acceleration template library and uses the spare hardware acceleration template to perform a reconfiguration operation on the first reconfiguration region. After the reconfiguration is completed, the state monitoring unit built into the control and scheduling module immediately records the abnormal situation of the reconfiguration region and reports the reconfiguration status to the control and scheduling module, confirming that the first reconfiguration region has returned to the normal operating state and has the ability to execute new tasks.

[0142] Through this embodiment, when a first reconfiguration region with an anomaly is detected, the spare hardware acceleration template can be automatically selected and loaded to perform a reconfiguration operation on the first reconfiguration region. Among them, when designing the spare hardware acceleration template, the hardware characteristic differences of all reconfiguration regions are fully considered, ensuring the generality and compatibility of the template, which means that a spare template can be applicable to the reconfiguration regions with multiple different hardware configurations. By quickly switching to the spare hardware acceleration template, the normal function of the first reconfiguration region can be restored without affecting the overall operation, enhancing the fault tolerance of the system and the continuity of computing tasks, effectively overcoming the scheduling difficulties brought by the heterogeneity of hardware resources, and realizing the efficient utilization of computing resources.

[0143] In an exemplary embodiment, the above server scheduling method further includes:

[0144] In the case of a second reconfiguration region where the running state data indicates that the load exceeds the preset threshold, migrate the subtasks on the second reconfiguration region to other reconfiguration regions.

[0145] Among them, the second reconstruction area refers to those reconstruction areas with excessively high current load levels.

[0146] Optionally, as Figure 7 shown, the control scheduling module maintains real-time communication with each module throughout the task execution, monitors the computing load, energy consumption data, signal transmission status, and reconstruction effect of each module. In the case where the running state data indicates a second reconstruction area with a load exceeding the preset threshold, the reconfigurable computing acceleration module migrates the subtasks on the second reconstruction area to other available reconstruction areas for execution.

[0147] Figure 8 is a schematic flowchart of a storage module storing task results provided by an embodiment of the present application. As Figure 8 shown, after the arithmetic task is completed, the reconfigurable computing acceleration module and the auxiliary computing module transmit the final calculation result to the storage module through a high-speed interface. After the storage module performs data verification on the result, it stores the result in a non-volatile storage device and generates a task execution log. The control scheduling module sorts out the task execution records, energy consumption indicators, and reconstruction records, and transmits them to the system monitoring platform for long-term system performance analysis and template library optimization.

[0148] Through this embodiment, when it is detected that the load of the second reconstruction area is too high, the subtasks on this area are migrated to other reconstruction areas with lower loads for execution, solving the problem of excessive resource concentration caused by static task allocation in the related art. By dynamically migrating subtasks, the balanced allocation of computing resources is achieved, preventing system performance degradation and increased energy consumption caused by local overload.

[0149] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0150] An embodiment of the present application also provides a server scheduling device. As Figure 9 shown, the server scheduling device includes:

[0151] A task decomposition module 902, configured to decompose an arithmetic task sent by a target object to obtain a plurality of subtasks.

[0152] An acceleration template matching module 904, configured to extract a hardware acceleration template corresponding to a subtask among the plurality of subtasks from a pre-compiled acceleration template library; the hardware acceleration templates stored in the pre-compiled acceleration template library refer to hardware circuit configurations pre-compiled for tasks of different arithmetic types.

[0153] The reconstruction area determination module 906 is configured to determine, from multiple reconstruction areas, a target reconstruction area corresponding to a subtask in multiple subtasks; the multiple reconstruction areas refer to logical blocks divided on a reconfigurable logic device that allow independent reconfiguration of logic circuits.

[0154] The task execution module 908 is configured to perform a reconstruction operation on the target reconstruction area corresponding to a subtask in multiple subtasks according to the hardware acceleration template corresponding to the subtask in multiple subtasks, obtain a new target reconstruction area corresponding to the subtask in multiple subtasks, and execute the corresponding subtask through the new target reconstruction area corresponding to the subtask in multiple subtasks.

[0155] It should be noted that the task decomposition module 902 in this embodiment can be used to execute the above step S202, the acceleration template matching module 904 in this embodiment can be used to execute the above step S204, the reconstruction area determination module 906 in this embodiment can be used to execute the above step S206, and the task execution module 908 in this embodiment can be used to execute the above step S208.

[0156] In an exemplary embodiment, the hardware acceleration templates stored in the pre-compiled acceleration template library include configuration descriptions of tasks of different operation types. The acceleration template matching module 904 is further configured to determine, as the hardware acceleration template corresponding to the subtask in multiple subtasks, the hardware acceleration template in the pre-compiled acceleration template library that matches the operation type of the subtask in multiple subtasks.

[0157] In an exemplary embodiment, the reconstruction area determination module 906 is further configured to perform matching calculations on the subtasks in multiple subtasks according to the calculation requirement data of the subtasks in multiple subtasks and the current resource status of multiple reconstruction areas, and determine the target reconstruction area corresponding to the subtasks in multiple subtasks.

[0158] In an exemplary embodiment, the reconstruction area determination module 906 is further configured to generate a task requirement vector of the subtasks in multiple subtasks based on the calculation requirement data of the subtasks in multiple subtasks; generate a plurality of regional resource vectors based on the current resource status of multiple reconstruction areas; the plurality of regional resource vectors correspond one-to-one to the multiple reconstruction areas; calculate the distance between the task requirement vector of the subtasks in multiple subtasks and the plurality of regional resource vectors, and determine the reconstruction area corresponding to the minimum distance as the target reconstruction area corresponding to the subtasks in multiple subtasks.

[0159] In an exemplary embodiment, the calculation requirement data of a subtask among multiple subtasks includes the expected computing power, expected memory capacity, and expected upper limit of reconstruction delay of the subtask among the multiple subtasks. The reconstruction area determination module 906 is further configured to generate a first task element based on the expected computing power of the subtask among the multiple subtasks; generate a second task element based on the expected memory capacity of the subtask among the multiple subtasks; generate a third task element based on the expected upper limit of reconstruction delay of the subtask among the multiple subtasks; and generate a task requirement vector of the subtask among the multiple subtasks according to the first task element, the second task element, and the third task element.

[0160] In an exemplary embodiment, the current resource state of a reconstruction area among multiple reconstruction areas includes the current computing power, current memory capacity, and current reconstruction delay estimated value of the reconstruction area among the multiple reconstruction areas. The reconstruction area determination module 906 is further configured to use the reconstruction area among the multiple reconstruction areas as the current reconstruction area and perform the following vector generation operations to obtain multiple area resource vectors: generate a first area element based on the current computing power of the current reconstruction area; generate a second area element based on the current memory capacity of the current reconstruction area; generate a third area element based on the current reconstruction delay estimated value of the current reconstruction area; and generate an area resource vector corresponding to the current reconstruction area according to the first area element, the second area element, and the third area element.

[0161] In an exemplary embodiment, the reconstruction area determination module 906 is further configured to use the Euclidean distance calculation formula to calculate the distance between the task requirement vector of the subtask among the multiple subtasks and the multiple area resource vectors; the Euclidean distance calculation formula characterizes the difference between the corresponding elements in the two vectors, and the association relationship between the difference, the normalization factor, and the index weight.

[0162] In an exemplary embodiment, the task requirement vector of a subtask among multiple subtasks includes a first task element, a second task element, and a third task element; the regional resource vector of a reconstruction region among multiple reconstruction regions includes a first region element, a second region element, and a third region element; the Euclidean distance calculation formula includes a first normalization factor, a second normalization factor, a third normalization factor, a first index weight, a second index weight, and a third index weight. The reconstruction region determination module 906 is further configured to use the task requirement vector of a subtask among multiple subtasks as the current task requirement vector, use the regional resource vector among multiple regional resource vectors as the current regional resource vector, perform a distance calculation operation to obtain the distance between the task requirement vector of a subtask among multiple subtasks and the regional resource vector: calculate a first difference between the first task element of the current task requirement vector and the first region element of the current regional resource vector, a second difference between the second task element of the current task requirement vector and the second region element of the current regional resource vector, and a third difference between the third task element of the current task requirement vector and the third region element of the current regional resource vector; determine a first squared difference according to a first ratio between the first difference and the first normalization factor; determine a second squared difference according to a second ratio between the second difference and the second normalization factor; determine a third squared difference according to a third ratio between the third difference and the third normalization factor; weight the first squared difference with the first index weight to obtain a first weighted value; weight the second squared difference with the second index weight to obtain a second weighted value; weight the third squared difference with the third index weight to obtain a third weighted value; and determine the square root of the sum of the first weighted value, the second weighted value, and the third weighted value as the distance between the current task requirement vector and the current regional resource vector.

[0163] In an exemplary embodiment, before performing a reconstruction operation on a target reconstruction region corresponding to a subtask among multiple subtasks according to the hardware acceleration template corresponding to the subtask among multiple subtasks, the task execution module 908 is further configured to perform a pre-check on the target reconstruction region corresponding to the subtask among multiple subtasks; and in the case that there is no real-time data transmission task in the target reconstruction region corresponding to the subtask among multiple subtasks, perform the step of performing a reconstruction operation on the target reconstruction region corresponding to the subtask among multiple subtasks according to the hardware acceleration template corresponding to the subtask among multiple subtasks.

[0164] In an exemplary embodiment, the task execution module 908 is further configured to, in the case that there is a real-time data transmission task in the target reconstruction region corresponding to a subtask among multiple subtasks, migrate the real-time data transmission task to another reconstruction region or wait for the real-time data transmission task to be completed.

[0165] In an exemplary embodiment, the task execution module 908 is further configured to monitor in real time the operating state data of the reconfigurable logic device; in the case where the operating state data indicates an abnormal first reconfiguration area, use a spare hardware acceleration template to perform a reconfiguration operation on the first reconfiguration area; the spare hardware acceleration template is compatible with the hardware characteristics of multiple reconfiguration areas.

[0166] In an exemplary embodiment, the task execution module 908 is further configured to, in the case where the operating state data indicates a second reconfiguration area where the load exceeds a preset threshold, migrate the subtasks on the second reconfiguration area to other reconfiguration areas.

[0167] For the description of the features in the corresponding embodiments of the server scheduling device, reference may be made to the relevant descriptions in the corresponding embodiments of the server scheduling method, which will not be elaborated here one by one.

[0168] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above embodiments of the server scheduling method.

[0169] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the above embodiments of the server scheduling method when running.

[0170] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs, and other various media that can store computer programs.

[0171] An embodiment of the present application further provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any one of the above embodiments of the server scheduling method are implemented.

[0172] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any one of the above embodiments of the server scheduling method are implemented.

[0173] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered as exceeding the scope of this application.

[0174] The above has introduced in detail a server scheduling method, device, storage medium, and program product provided by this application. Specific examples are used herein to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A server scheduling method, characterized in that: include: Decompose the computing task sent by the target object into multiple subtasks; Extracting a hardware acceleration template corresponding to a subtask among the plurality of subtasks from a precompiled acceleration template library; the hardware acceleration template stored in the precompiled acceleration template library refers to hardware circuit configurations precompiled for tasks of different operation types; Determine a target reconstruction area corresponding to a subtask in the plurality of subtasks from a plurality of reconstruction areas; the plurality of reconstruction areas refer to logic blocks divided on a reconfigurable logic device that allow independent reconfiguration of logic circuits; According to the hardware acceleration template corresponding to the subtask among the multiple subtasks, a reconstruction operation is performed on the target reconstruction area corresponding to the subtask among the multiple subtasks to obtain a new target reconstruction area corresponding to the subtask among the multiple subtasks, and the corresponding subtask is executed through the new target reconstruction area corresponding to the subtask among the multiple subtasks.

2. The server scheduling method according to claim 1, characterized in that: The hardware acceleration templates stored in the precompiled acceleration template library include configuration instructions for tasks of different computing types; The extracting a hardware acceleration template corresponding to a subtask among the plurality of subtasks from a precompiled acceleration template library comprises: A hardware acceleration template in the precompiled acceleration template library that matches the operation type of a subtask among the multiple subtasks is determined as the hardware acceleration template corresponding to the subtask among the multiple subtasks.

3. The server scheduling method according to claim 1, characterized in that: The determining, from the plurality of reconstruction regions, a target reconstruction region corresponding to a subtask among the plurality of subtasks comprises: According to the computational requirement data of the subtasks among the multiple subtasks and the current resource states of the multiple reconstruction areas, matching computations are performed on the subtasks among the multiple subtasks to determine target reconstruction areas corresponding to the subtasks among the multiple subtasks.

4. The server scheduling method according to claim 3, characterized in that: The performing matching calculations on the subtasks in the multiple subtasks according to the computational requirement data of the subtasks in the multiple subtasks and the current resource states of the multiple reconstruction areas to determine the target reconstruction areas corresponding to the subtasks in the multiple subtasks includes: generating a task requirement vector of a subtask in the plurality of subtasks based on the computational requirement data of the subtask in the plurality of subtasks; Based on the current resource states of the multiple reconstruction areas, a plurality of area resource vectors are generated; the plurality of area resource vectors correspond one-to-one to the plurality of reconstruction areas; The distance between the task requirement vector of the subtask among the multiple subtasks and the multiple regional resource vectors is calculated, and a reconstruction region corresponding to the minimum distance is determined as a target reconstruction region corresponding to the subtask among the multiple subtasks.

5. The server scheduling method according to claim 4, characterized in that: The computing requirement data of a subtask in the plurality of subtasks includes the expected computing capability, expected memory capacity, and expected reconstruction delay upper limit of the subtask in the plurality of subtasks; The step of generating a task requirement vector of a subtask among the plurality of subtasks based on the computing requirement data of the subtask among the plurality of subtasks comprises: generating a first task element based on an expected computing capability of a subtask among the plurality of subtasks; generating a second task element based on an expected memory capacity of a subtask among the plurality of subtasks; generating a third task element based on an expected reconstruction delay upper limit of a subtask among the plurality of subtasks; A task requirement vector of a subtask among the plurality of subtasks is generated according to the first task element, the second task element, and the third task element.

6. The server scheduling method according to claim 4, characterized in that: The current resource status of the reconstruction area among the multiple reconstruction areas comprises the current computing capability, the current memory capacity, and the current reconstruction delay estimation value of the reconstruction area among the multiple reconstruction areas; The generating a plurality of regional resource vectors based on the current resource states of the plurality of reconstruction regions comprises: A reconstructed region among the multiple reconstructed regions is used as a current reconstructed region, and the following vector generation operation is performed to obtain the multiple region resource vectors: generating a first region element based on a current computing capability of the current reconstruction region; generating a second region element based on a current memory capacity of the current reconstruction region; generating a third region element based on a current reconstruction delay estimate of the current reconstruction region; A regional resource vector corresponding to the current reconstruction area is generated according to the first regional element, the second regional element and the third regional element.

7. The server scheduling method according to claim 4, characterized in that: The calculating the distance between the task requirement vector of the subtask in the multiple subtasks and the multiple regional resource vectors includes: The Euclidean distance calculation formula is used to calculate the distance between the task requirement vector of the subtask in the multiple subtasks and the multiple regional resource vectors; the Euclidean distance calculation formula represents the difference between corresponding elements in two vectors, and the correlation relationship between the difference and the normalization factor and the indicator weight.

8. The server scheduling method according to claim 7, characterized in that: The task requirement vector of the subtask in the multiple subtasks includes a first task element, a second task element and a third task element; the regional resource vector of the reconstructed area in the multiple reconstructed areas includes a first regional element, a second regional element and a third regional element; the Euclidean distance calculation formula includes a first normalization factor, a second normalization factor, a third normalization factor, a first indicator weight, a second indicator weight and a third indicator weight; The using of the Euclidean distance calculation formula to calculate the distance between the task requirement vector of the subtask in the multiple subtasks and the multiple regional resource vectors includes: Taking the task requirement vector of a subtask in the multiple subtasks as the current task requirement vector, taking the regional resource vector in the multiple regional resource vectors as the current regional resource vector, performing a distance calculation operation, and obtaining the distance between the task requirement vector of the subtask in the multiple subtasks and the multiple regional resource vectors: Calculating a first difference between a first task element of the current task requirement vector and a first regional element of the current regional resource vector, a second difference between a second task element of the current task requirement vector and a second regional element of the current regional resource vector, and a third difference between a third task element of the current task requirement vector and a third regional element of the current regional resource vector; determining a first square difference according to a first ratio between the first difference and the first normalization factor; determining a second square difference according to a second ratio between the second difference and the second normalization factor; and determining a third square difference according to a third ratio between the third difference and the third normalization factor; The first square difference is weighted by the first indicator weight to obtain a first weighted value; the second square difference is weighted by the second indicator weight to obtain a second weighted value; the third square difference is weighted by the third indicator weight to obtain a third weighted value; The square root of the sum of the first weighted value, the second weighted value, and the third weighted value is determined as the distance between the current task demand vector and the current area resource vector.

9. The server scheduling method according to claim 1, characterized in that: Before performing a reconstruction operation on a target reconstruction area corresponding to a subtask among the multiple subtasks according to a hardware acceleration template corresponding to a subtask among the multiple subtasks, the method further includes: Pre-check the target reconstruction area corresponding to the subtask among the multiple subtasks; if there is no real-time data transmission task in the target reconstruction area corresponding to the subtask among the multiple subtasks, execute the step of reconstructing the target reconstruction area corresponding to the subtask among the multiple subtasks according to the hardware acceleration template corresponding to the subtask among the multiple subtasks.

10. The server scheduling method according to claim 9, characterized in that: The method further comprises: When a real-time data transmission task exists in a target reconstruction area corresponding to a subtask among the multiple subtasks, the real-time data transmission task is migrated to another reconstruction area, or the real-time data transmission task is waited for to be completed.

11. The server scheduling method according to any one of claims 1 to 10, characterized in that: The method further comprises: Monitor the operating status data of the reconfigurable logic device in real time; when the operating status data indicates that there is an abnormal first reconstruction area, use a spare hardware acceleration template to reconstruct the first reconstruction area; the spare hardware acceleration template is compatible with the hardware features of the multiple reconstruction areas.

12. The server scheduling method according to claim 11, characterized in that: The method further comprises: When the running status data indicates that there is a second reconstruction area whose load exceeds a preset threshold, the subtask on the second reconstruction area is migrated to other reconstruction areas.

13. A server scheduling device, characterized in that: include: The task decomposition module is used to decompose the computing task sent by the target object into multiple subtasks; An acceleration template matching module, used to extract a hardware acceleration template corresponding to a subtask in the plurality of subtasks from a precompiled acceleration template library; the hardware acceleration template stored in the precompiled acceleration template library refers to a hardware circuit configuration precompiled for tasks of different operation types; A reconfiguration region determination module, used to determine a target reconfiguration region corresponding to a subtask among the multiple subtasks from multiple reconfiguration regions; the multiple reconfiguration regions refer to logic blocks divided out of a reconfigurable logic device and allowing independent reconfiguration of logic circuits; A task execution module is used to perform a reconstruction operation on a target reconstruction area corresponding to a subtask among the multiple subtasks according to a hardware acceleration template corresponding to the subtask among the multiple subtasks, obtain a new target reconstruction area corresponding to the subtask among the multiple subtasks, and execute the corresponding subtask through the new target reconstruction area corresponding to the subtask among the multiple subtasks.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the server scheduling method according to any one of claims 1 to 12.

15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the server scheduling method according to any one of claims 1 to 12 are implemented.

Citation Information

Patent Citations

  • Intelligent deployment and reconstruction method for baseband resource pool

    CN110430591A

  • Fast search method for task layout of dynamic partially reconfigurable system

    CN114647504A

  • Multi-task cooperative processing method and system based on digital employees

    CN119179563A

  • Multi-task data analysis method and device and storage medium

    CN119847752A

  • Heterogeneous chip design method and system based on Internet of Things

    CN119862833A