Hardware resource scheduling methods, devices, systems, storage media, and program products

By recording and parsing the hardware resource access information of user-mode drivers, hardware resource scheduling is achieved, solving the problems of high development cost and low efficiency in the pre-silicon verification stage, optimizing the hardware resource scheduling mechanism, and improving testing flexibility and verification efficiency.

CN120610827BActive Publication Date: 2025-10-28MOORE THREADS TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511093027.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-10-28
Estimated Expiration
2045-08-05

AI Technical Summary

Technical Problem

The hardware resource scheduling method in the pre-silicon verification stage relies on kernel-mode drivers or firmware, resulting in strong coupling between user-mode and kernel-mode development, high development costs, low overall efficiency, and difficulty in rapid adjustment and optimization.

Method used

By recording hardware resource access information of user-mode drivers in real time, parsing and replaying operation instruction information, hardware resource scheduling is achieved, avoiding the involvement of kernel-mode drivers and firmware, and sending instructions in user mode using the PCIe bus.

Benefits of technology

It reduces software development costs in the pre-silicon verification stage, improves testing flexibility and verification efficiency, and decouples user-mode and kernel-mode drivers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120610827B_ABST
    Figure CN120610827B_ABST
Patent Text Reader

Abstract

This disclosure relates to the field of pre-silicon verification technology, proposing a hardware resource scheduling method, apparatus, system, storage medium, and program product. The method includes: real-time recording of hardware resource access information of the target hardware during the execution of a user-mode driver; parsing the recorded access information to obtain operation instruction information; accessing hardware resources by replaying the operation instruction information to obtain hardware execution results, wherein the process of replaying the operation instruction information is implemented by sending instructions to PCIe in user mode, without the involvement of firmware or kernel-mode drivers; and determining the test results of the target hardware based on the hardware execution results and the expected execution results. According to embodiments of this disclosure, the development of user-mode and kernel-mode drivers can be decoupled, reducing software development costs in the pre-silicon verification stage, optimizing the hardware resource scheduling mechanism, thereby improving test flexibility and verification efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of pre-silicon verification technology, and in particular to a hardware resource scheduling method, apparatus, system, storage medium, and program product. Background Technology

[0002] Pre-silicon verification (PSC) refers to the process of thoroughly verifying a chip's functionality, performance, and power consumption before tape-out, using methods such as simulation, FPGA prototyping, and formal verification. This ensures the design meets specifications and reduces the risk of discovering problems after tape-out. The PSC stage typically requires the complete development of user-mode drivers, kernel-mode drivers, and firmware code to load workloads onto the hardware, schedule hardware resources, and verify hardware execution results.

[0003] Taking graphics processing units (GPUs) as an example, current hardware resource scheduling methods in the pre-silicon verification stage are mainly divided into firmware-based scheduling and hardware-based scheduling. GPU scheduling relies on kernel-mode drivers or firmware, resulting in strong coupling between user-mode and kernel-mode development (requiring simultaneous development of user-mode and kernel-mode drivers) to enable correct interaction with the hardware. However, due to the need for continuous adjustment and optimization in the pre-silicon verification stage, this fixed scheduling method has high development costs and low overall efficiency. Therefore, optimizing the hardware resource scheduling mechanism and improving testing flexibility and verification efficiency are important challenges faced by the pre-silicon verification stage. Summary of the Invention

[0004] In view of this, this disclosure proposes a hardware resource scheduling method, apparatus, system, storage medium, and program product.

[0005] According to one aspect of this disclosure, a hardware resource scheduling method is provided. The method includes:

[0006] It records in real time the access information of hardware resources of the target hardware during the execution of the user-mode driver. The target hardware includes the graphics processing unit (GPU). The user-mode driver is pre-developed and associated with the feature to be tested in the target hardware.

[0007] The operation instruction information is obtained by parsing the recorded access information;

[0008] Hardware resources are accessed by replaying operation instruction information to obtain hardware execution results. The process of replaying operation instruction information is achieved by sending instructions to the high-speed serial bus PCIe in user space, without the need for firmware or kernel-mode drivers.

[0009] The test results for the target hardware are determined based on the hardware execution results and the expected execution results.

[0010] In one possible implementation, the access information includes access text. Based on the recorded access information, parsing is performed to obtain operation instruction information, including:

[0011] Each line of text in the accessed record is segmented into words to obtain the segmentation result of each line of text, and each line of text corresponds to an operation instruction.

[0012] Based on the word segmentation result of any text line, determine the type of operation instruction corresponding to any text line;

[0013] Based on the type of operation instruction, determine the operation instruction information indicated in the word segmentation result of any text line.

[0014] In one possible implementation, the type of operation instruction includes memory access instructions. Based on the type of operation instruction, the operation instruction information indicated in the word segmentation result of any text line is determined, including:

[0015] When the type of operation instruction is a memory access instruction, determine one or more of the following as operation instruction information: the type of memory access instruction indicated in the word segmentation result, the information associated with the memory access instruction, and the memory data file associated with the memory access instruction.

[0016] In one possible implementation, the type of operation instruction also includes a task submission instruction. Based on the type of operation instruction, the operation instruction information indicated in the word segmentation result of any text line is determined, including:

[0017] When the type of operation instruction is a task submission instruction, determine one or more of the following as operation instruction information: the task type indicated in the word segmentation result, the information associated with the task submission instruction, and the task file associated with the task submission instruction.

[0018] In one possible implementation, hardware resources are accessed by replaying operation instruction information to obtain the hardware execution result, including:

[0019] When the operation instruction information is a memory request instruction or a memory update instruction, data is written to the memory address of the target hardware through the PCIe bus based on the information associated with the memory access instruction and the memory data file associated with the memory access instruction to replay the memory operation and obtain the hardware execution result.

[0020] When the operation instruction information is any one of page table mapping instruction, memory release instruction, or page table demapping instruction, the hardware page table is configured via the PCIe bus based on the information associated with the memory access instruction to replay the operation and obtain the hardware execution result.

[0021] In one possible implementation, hardware resources are accessed by replaying operation instruction information to obtain the hardware execution result, including:

[0022] When the operation instruction information is a task submission instruction, the register configuration sequence is obtained by parsing the task file associated with the task submission instruction. The register configuration sequence is an instruction used to access hardware registers.

[0023] Based on the task type and information associated with the task submission instruction, the register configuration sequence is sent to the target hardware via the PCIe bus to replay the task submission and obtain the hardware execution result. The replay process does not require firmware to participate in the scheduling.

[0024] In one possible implementation, the test results of the target hardware are determined based on the hardware execution results and the expected execution results, including:

[0025] By comparing the hardware execution result with the expected execution result, and in the case where there are inconsistent data items between the hardware execution result and the expected execution result, the location information and content of the inconsistent data items are determined;

[0026] Add the location information and content of inconsistent data items to the test results of the target hardware.

[0027] According to another aspect of this disclosure, a hardware resource scheduling apparatus is provided. The apparatus includes:

[0028] The first determination module is used to record in real time the access information of hardware resources of the target hardware during the execution of the user-mode driver. The target hardware includes a graphics processing unit (GPU). The user-mode driver is pre-developed and associated with the feature to be tested in the target hardware.

[0029] The second determining module is used to parse the recorded access information to obtain operation instruction information;

[0030] The third determination module is used to access hardware resources by replaying operation instruction information and obtain hardware execution results. The process of replaying operation instruction information is achieved by sending instructions to the high-speed serial bus PCIe in user mode, without the need for firmware or kernel mode drivers to participate.

[0031] The fourth determination module is used to determine the test results of the target hardware based on the hardware execution results and the expected execution results.

[0032] In one possible implementation, the access information includes access text, and a second determining module is used for:

[0033] Each line of text in the accessed record is segmented into words to obtain the segmentation result of each line of text, and each line of text corresponds to an operation instruction.

[0034] Based on the word segmentation result of any text line, determine the type of operation instruction corresponding to any text line;

[0035] Based on the type of operation instruction, determine the operation instruction information indicated in the word segmentation result of any text line.

[0036] In one possible implementation, the type of operation instruction includes memory access instructions. Based on the type of operation instruction, the operation instruction information indicated in the word segmentation result of any text line is determined, including:

[0037] When the type of operation instruction is a memory access instruction, determine one or more of the following as operation instruction information: the type of memory access instruction indicated in the word segmentation result, the information associated with the memory access instruction, and the memory data file associated with the memory access instruction.

[0038] In one possible implementation, the type of operation instruction also includes a task submission instruction. Based on the type of operation instruction, the operation instruction information indicated in the word segmentation result of any text line is determined, including:

[0039] When the type of operation instruction is a task submission instruction, determine one or more of the following as operation instruction information: the task type indicated in the word segmentation result, the information associated with the task submission instruction, and the task file associated with the task submission instruction.

[0040] In one possible implementation, the third determining module is used for:

[0041] When the operation instruction information is a memory request instruction or a memory update instruction, data is written to the memory address of the target hardware through the PCIe bus based on the information associated with the memory access instruction and the memory data file associated with the memory access instruction to replay the memory operation and obtain the hardware execution result.

[0042] When the operation instruction information is any one of page table mapping instruction, memory release instruction, or page table demapping instruction, the hardware page table is configured via the PCIe bus based on the information associated with the memory access instruction to replay the operation and obtain the hardware execution result.

[0043] In one possible implementation, the third determining module is used for:

[0044] When the operation instruction information is a task submission instruction, the register configuration sequence is obtained by parsing the task file associated with the task submission instruction. The register configuration sequence is an instruction used to access hardware registers.

[0045] Based on the task type and information associated with the task submission instruction, the register configuration sequence is sent to the target hardware via the PCIe bus to obtain the hardware execution result. The playback process does not require firmware to participate in the scheduling.

[0046] In one possible implementation, the fourth determining module is used for:

[0047] By comparing the hardware execution result with the expected execution result, and in the case where there are inconsistent data items between the hardware execution result and the expected execution result, the location information and content of the inconsistent data items are determined;

[0048] Add the location information and content of inconsistent data items to the test results of the target hardware.

[0049] According to another aspect of this disclosure, a hardware resource scheduling system is provided. The system includes a parser and an executor.

[0050] The parser is used for:

[0051] It records in real time the access information of hardware resources of the target hardware during the execution of the user-mode driver. The target hardware includes the graphics processing unit (GPU). The user-mode driver is pre-developed and associated with the feature to be tested in the target hardware.

[0052] The operation instruction information is obtained by parsing the recorded access information;

[0053] The actuator is used for:

[0054] Hardware resources are accessed by replaying operation instruction information to obtain hardware execution results. The process of replaying operation instruction information is achieved by sending instructions to the high-speed serial bus PCIe in user space, without the need for firmware or kernel-mode drivers.

[0055] The test results for the target hardware are determined based on the hardware execution results and the expected execution results.

[0056] In one possible implementation, the access information includes access text. Based on the recorded access information, parsing is performed to obtain operation instruction information, including:

[0057] Each line of text in the accessed record is segmented into words to obtain the segmentation result of each line of text, and each line of text corresponds to an operation instruction.

[0058] Based on the word segmentation result of any text line, determine the type of operation instruction corresponding to any text line;

[0059] Based on the type of operation instruction, determine the operation instruction information indicated in the word segmentation result of any text line.

[0060] In one possible implementation, the type of operation instruction includes memory access instructions. Based on the type of operation instruction, the operation instruction information indicated in the word segmentation result of any text line is determined, including:

[0061] When the type of operation instruction is a memory access instruction, determine one or more of the following as operation instruction information: the type of memory access instruction indicated in the word segmentation result, the information associated with the memory access instruction, and the memory data file associated with the memory access instruction.

[0062] In one possible implementation, the type of operation instruction also includes a task submission instruction. Based on the type of operation instruction, the operation instruction information indicated in the word segmentation result of any text line is determined, including:

[0063] When the type of operation instruction is a task submission instruction, determine one or more of the following as operation instruction information: the task type indicated in the word segmentation result, the information associated with the task submission instruction, and the task file associated with the task submission instruction.

[0064] In one possible implementation, hardware resources are accessed by replaying operation instruction information to obtain the hardware execution result, including:

[0065] When the operation instruction information is a memory request instruction or a memory update instruction, data is written to the memory address of the target hardware through the PCIe bus based on the information associated with the memory access instruction and the memory data file associated with the memory access instruction to replay the memory operation and obtain the hardware execution result.

[0066] When the operation instruction information is any one of page table mapping instruction, memory release instruction, or page table demapping instruction, the hardware page table is configured via the PCIe bus based on the information associated with the memory access instruction to replay the operation and obtain the hardware execution result.

[0067] In one possible implementation, hardware resources are accessed by replaying operation instruction information to obtain the hardware execution result, including:

[0068] When the operation instruction information is a task submission instruction, the register configuration sequence is obtained by parsing the task file associated with the task submission instruction. The register configuration sequence is an instruction used to access hardware registers.

[0069] Based on the task type and information associated with the task submission instruction, the register configuration sequence is sent to the target hardware via the PCIe bus to replay the task submission and obtain the hardware execution result. The replay process does not require firmware to participate in the scheduling.

[0070] In one possible implementation, the test results of the target hardware are determined based on the hardware execution results and the expected execution results, including:

[0071] By comparing the hardware execution result with the expected execution result, and in the case where there are inconsistent data items between the hardware execution result and the expected execution result, the location information and content of the inconsistent data items are determined;

[0072] Add the location information and content of inconsistent data items to the test results of the target hardware.

[0073] According to another aspect of this disclosure, a hardware resource scheduling apparatus is provided, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above method.

[0074] According to another aspect of this disclosure, a non-volatile computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the above-described method.

[0075] According to another aspect of this disclosure, a computer program product is provided, including a computer program or a non-volatile computer-readable storage medium carrying the computer program, wherein the computer program, when executed by a processor, implements the steps of the above-described method.

[0076] According to embodiments of this disclosure, by recording in real-time the access information of hardware resources of the target hardware, including a GPU, during the execution of the user-mode driver, the access information of the target hardware, including a GPU, can be obtained. The user-mode driver is pre-developed and associated with the feature to be tested in the target hardware. This allows for obtaining relevant access information of the feature to be tested in the target hardware without developing only the user-mode driver, thus improving verification efficiency. Through the process of recording access information, the user-mode driver's operations on hardware resources can be captured, and the hardware interaction logic can be stored independently. Furthermore, by parsing the recorded access information, operation instruction information is obtained. By replaying the operation instruction information, hardware resources are accessed to obtain the hardware execution result. The process of replaying the operation instruction information can be implemented by sending instructions to the PCIe bus in user mode, without the need for kernel-mode drivers or firmware. This complete recording and playback logic breaks the existing strong dependency relationship, achieving decoupling of user-mode and kernel-mode driver development and scheduling of GPU hardware resources. Among these advantages, the absence of kernel-mode drivers reduces software development costs in the pre-silicon verification phase. Furthermore, the absence of firmware development in the pre-silicon verification phase allows for the replay of operation instructions in user mode to access hardware resources and obtain hardware execution results, thereby optimizing hardware resource scheduling mechanisms and improving testing flexibility and verification efficiency.

[0077] Other features and aspects of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description

[0078] The accompanying drawings, which are included in and form part of this specification, illustrate exemplary embodiments, features, and aspects of this disclosure together with the specification and serve to explain the principles of this disclosure.

[0079] Figure 1 A schematic diagram illustrating an application scenario according to an embodiment of this disclosure is shown.

[0080] Figure 2 A flowchart illustrating a hardware resource scheduling method according to an embodiment of the present disclosure is shown.

[0081] Figure 3 A structural diagram of a hardware resource scheduling apparatus according to an embodiment of the present disclosure is shown.

[0082] Figure 4 A schematic diagram of the structure of a hardware resource scheduling system according to an embodiment of the present disclosure is shown.

[0083] Figure 5 This is a block diagram illustrating an apparatus 1900 for scheduling hardware resources according to an exemplary embodiment. Detailed Implementation

[0084] Various exemplary embodiments, features, and aspects of this disclosure will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.

[0085] As used herein, the terms “comprising,” “including,” “having,” or variations thereof are open-ended and include one or more of the stated features, integrals, elements, steps, components, or functions, but do not exclude the presence or addition of one or more other features, integrals, elements, steps, components, functions, or groups thereof.

[0086] When an element is referred to as “connected,” “coupled,” “responding,” or a variation thereof relative to another element, it may be directly connected, coupled, or responding to another element, or there may be an intermediate element present.

[0087] Although the terms first, second, third, etc., may be used herein to describe various elements / operations, these elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another. Therefore, without departing from the teachings of the inventive concept, a first element / operation in some embodiments may be referred to as a second element / operation in other embodiments.

[0088] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.

[0089] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.

[0090] Pre-silicon verification (PSC) refers to the process of thoroughly verifying a chip's functionality, performance, and power consumption before tape-out, using methods such as simulation, FPGA prototyping, and formal verification. This ensures the design meets specifications and reduces the risk of discovering problems after tape-out. The PSC stage typically requires the complete development of user-mode drivers, kernel-mode drivers, and firmware code to load workloads onto the hardware, schedule hardware resources, and verify hardware execution results.

[0091] Taking graphics processing units (GPUs) as an example, current hardware resource scheduling methods in the pre-silicon verification stage are mainly divided into firmware-based scheduling and hardware-based scheduling. GPU scheduling relies on kernel-mode drivers or firmware, resulting in strong coupling between user-mode and kernel-mode development (requiring simultaneous development of user-mode and kernel-mode drivers) to enable correct interaction with the hardware. Specifically:

[0092] (1) Firmware-based scheduling refers to the GPU's built-in embedded processor (such as ARM (Acorn RISC Machine), MIPS (Microprocessor without Interlocked Pipeline Stages), or RISC-V (Reduced Instruction Set Computer V)). User-mode drivers and kernel-mode drivers load the workload into the firmware program, which is then responsible for scheduling GPU tasks. This approach requires the development of a complete upper-layer application (User Application) during the pre-silicon verification stage, involving a large amount of software development work, resulting in high manpower and time costs. In addition, when new features are introduced into the chip, many software modifications are involved, making it difficult to quickly locate and solve problems, especially when there are many new features, making rapid convergence even more difficult.

[0093] (2) Hardware-based scheduling refers to the GPU having a dedicated hardware scheduling unit. User-mode drivers and kernel-mode drivers send workloads directly to the hardware for automatic scheduling and task management. This method also requires the complete development of user-mode and kernel-mode drivers for testing during the pre-silicon verification stage. Therefore, this mechanism still faces high manpower and time costs, posing a challenge to the efficiency of pre-silicon verification.

[0094] Because the pre-silicon verification phase requires continuous adjustment and optimization, which involves simulating various real-world workloads, it necessitates constructing reasonable and effective test inputs (i.e., stimuli) to cover a wide range of possible hardware behaviors. The aforementioned fixed scheduling methods are costly to develop and have low overall efficiency. Therefore, optimizing hardware resource scheduling mechanisms to improve test flexibility and verification efficiency is a significant challenge facing the pre-silicon verification phase.

[0095] In view of this, embodiments of this disclosure provide a hardware resource scheduling method, apparatus, system, storage medium, and program product. The method of this disclosure records in real-time access information of hardware resources of the target hardware, including a GPU, during the execution of a user-mode driver. The user-mode driver is pre-developed and associated with the feature to be tested in the target hardware. This allows obtaining relevant access information of the feature to be tested in the target hardware without developing only the user-mode driver, improving verification efficiency. Through the process of recording access information, the operation of the user-mode driver on hardware resources can be captured, and the hardware interaction logic can be stored independently. Furthermore, by parsing the recorded access information, operation instruction information is obtained. By replaying the operation instruction information to access hardware resources, the hardware execution result is obtained. The process of replaying the operation instruction information can be implemented by sending instructions to the PCIe bus in user mode without the need for kernel-mode drivers or firmware. This complete recording and playback logic breaks the existing strong dependency relationship, achieving decoupling of user-mode and kernel-mode driver development and scheduling of GPU hardware resources. Among these advantages, the absence of kernel-mode drivers reduces software development costs in the pre-silicon verification phase. Furthermore, the absence of firmware development in the pre-silicon verification phase allows for the replay of operation instructions in user mode to access hardware resources and obtain hardware execution results, thereby optimizing hardware resource scheduling mechanisms and improving testing flexibility and verification efficiency.

[0096] Figure 1 This diagram illustrates an application scenario according to embodiments of the present disclosure. Embodiments of the present disclosure can be used in scenarios involving pre-silicon verification of target hardware (such as a GPU), such as... Figure 1As shown, in the application scenario of this disclosure embodiment, the target hardware (e.g., GPU) is not limited to a physical chip, but can be a simulation model in the pre-silicon verification stage, such as a GPU simulation model described in a hardware description language. This model can record in real time the access information of the GPU's hardware resources during the execution of the user-mode driver (UMD). After processing the access information through the parser and executor of this disclosure embodiment, the parser parses the access information, and the executor replays the parsed result. Hardware instructions recognizable by the GPU are then issued to the GPU via a high-speed serial bus (Peripheral Component Interconnect Express, PCIe) to schedule the GPU's hardware resources and obtain the GPU's execution result. Furthermore, the GPU's execution result can be compared with the expected result to perform pre-silicon verification on the GPU and obtain the GPU pre-silicon verification result.

[0097] The parser and executor can be ELF format program files. The above pre-silicon verification scenario can be implemented in a Linux environment. The GPU pre-silicon verification results can be obtained by entering the corresponding program files of the parser and executor, as well as the files required for pre-silicon verification (i.e., access information) in the command line.

[0098] The method described in this disclosure can be used in the user mode of a central processing unit (CPU) without involving the CPU's kernel mode. In other words, this disclosure does not involve the development of a kernel mode driver (KMD) in the scenario of pre-silicon verification of a GPU, thus decoupling user mode and kernel mode. The CPU described above can be deployed on terminal devices or servers, and this disclosure does not impose any limitations on this.

[0099] Figure 2 A flowchart illustrating a hardware resource scheduling method according to an embodiment of this disclosure is shown. Figure 2 As shown, the method may include:

[0100] Step S201: Record the access information of hardware resources of the target hardware during the execution of the user-mode driver in real time.

[0101] The target hardware may include a GPU, and this method can be applied to the pre-silicon verification process for GPUs.

[0102] User-mode drivers (such as) Figure 1In this context, UMD (User-Defined Device) can represent a user-facing application. This application can be pre-developed and associated with the characteristics to be tested in the target hardware (such as GPU hardware architecture). Therefore, it allows direct access information related to the GPU's characteristics to be tested without relying on the underlying kernel-mode driver or firmware. It captures user-mode driver operations on hardware resources (such as memory access and task submission) and stores hardware interaction logic independently. Furthermore, when the GPU's characteristics to be tested change, only the user-mode driver needs to be updated; there is no need to redevelop or adapt the kernel-mode driver or firmware. This improves the flexibility of the testing process and effectively reduces the development and maintenance costs of the pre-silicon verification stage.

[0103] Step S202: Parse the recorded access information to obtain operation instruction information.

[0104] Among them, the following can be used: Figure 1 The parser shown is used to execute step S202.

[0105] Step S203: Access hardware resources by replaying operation instruction information to obtain hardware execution results.

[0106] It can be used as Figure 1 The executor shown executes step S203. The executor can convert different operation instruction types into instructions that can directly access hardware resources, and the process of replaying operation instruction information can be achieved by sending instructions to the PCIe in user mode to schedule GPU hardware resources and obtain the GPU's execution result.

[0107] Step S204: Determine the test results of the target hardware based on the hardware execution results and the expected execution results.

[0108] The hardware execution result and the expected execution result can be either computational results or graphics rendering results. The test results of the target hardware can be determined by comparing the computational results or graphics rendering results corresponding to the hardware execution result and the expected execution result, respectively. The test results of the target hardware can be pre-silicon verification results for the GPU, indicating the correctness of the GPU design.

[0109] According to embodiments of this disclosure, by recording in real-time the access information of hardware resources of the target hardware, including a GPU, during the execution of the user-mode driver, the access information of the target hardware, including a GPU, can be obtained. The user-mode driver is pre-developed and associated with the feature to be tested in the target hardware. This allows for obtaining relevant access information of the feature to be tested in the target hardware without developing only the user-mode driver, thus improving verification efficiency. Through the process of recording access information, the user-mode driver's operations on hardware resources can be captured, and the hardware interaction logic can be stored independently. Furthermore, by parsing the recorded access information, operation instruction information is obtained. By replaying the operation instruction information, hardware resources are accessed to obtain the hardware execution result. The process of replaying the operation instruction information can be implemented by sending instructions to the PCIe bus in user mode, without the need for kernel-mode drivers or firmware. This complete recording and playback logic breaks the existing strong dependency relationship, achieving decoupling of user-mode and kernel-mode driver development and scheduling of GPU hardware resources. Among these advantages, the absence of kernel-mode drivers reduces software development costs in the pre-silicon verification phase. Furthermore, the absence of firmware development in the pre-silicon verification phase allows for the replay of operation instructions in user mode to access hardware resources and obtain hardware execution results, thereby optimizing hardware resource scheduling mechanisms and improving testing flexibility and verification efficiency.

[0110] Before executing the user-mode driver, the hardware control data block can be determined based on the high-level tasks passed by the upper-layer application or runtime application programming interface (API) according to the predefined GPU control flow specification, and the hardware register information of the target hardware can be configured according to the architecture document (which describes the architecture and design information of the target hardware; the architecture document may be different for different hardware). This allows the target hardware to be accessed during the execution of the user-mode driver by utilizing the hardware control data block and the hardware register information of the target hardware.

[0111] Hardware control data blocks can be binary code and may include information about the target hardware's task execution, which can be determined based on pre-silicon testing requirements. For example, this information could include graphics rendering tasks, computation tasks, etc. Task information could include hardware operating status information, an index list for indexing vertex data, vertex data (representing the number of attributes of each vertex of a graphics object), and instruction code (for instructing the hardware to execute tasks according to predetermined steps). For different chip types (e.g., different GPUs), due to inconsistent architecture implementations, the GPU control flow specifications also differ, and user-mode control programs can determine hardware control data blocks based on different architecture implementations. For example, examples of determining hardware control data blocks during graphics state settings can be found in Table 1.

[0112] Table 1

[0113]

[0114] According to Table 1, a user-mode driver can request a 64-bit data block and fill the 64-bit data with the bit arrangement shown in Table 1. Bits 0-45 (46 bits in total) are the start address field, bits 46-53 (8 bits in total) are the size field, and bits 56-63 (8 bits in total) are the command field.

[0115] During the process of configuring the hardware register information of the target hardware according to the architecture document, the user-mode driver can, for example, configure the GPU's hardware registers to the correct values ​​according to the description in the architecture document in order to correctly address the aforementioned hardware control data blocks.

[0116] Since the execution of the aforementioned UMD involves interaction with the hardware resources of the target hardware (such as the GPU), it is usually necessary to use KMD to further transform the execution result of the UMD in order to achieve actual access to the GPU hardware resources. This requires the additional development of KMD during pre-silicon testing. In order to achieve rapid verification of the correctness of GPU design in the pre-silicon testing scenario, this embodiment of the disclosure records the access information of the hardware resources of the GPU hardware involved in the execution of the UMD in a preset format, and achieves actual access to the hardware resources through the parser and executor of this embodiment of the disclosure. In this process, it is not necessary to develop KMD or additional firmware.

[0117] In one possible implementation, the aforementioned access information to the hardware resources of the target hardware may include access text, and in step S202, it may be:

[0118] Each line of text in the recorded access text is segmented into words to obtain the segmentation result of each line of text; based on the segmentation result of any line of text, the type of operation instruction corresponding to any line of text is determined; based on the type of operation instruction, the operation instruction information indicated in the segmentation result of any line of text is determined.

[0119] Access text can be used to represent information related to hardware resource interaction, such as memory allocation and release, memory read and write, page table editing, task submission, device information acquisition, and interrupt control. This information can be recorded in formats such as text documents, so that the interaction process between user-mode drivers and hardware can be represented in a structured form, which is convenient for subsequent parsing and playback.

[0120] According to the embodiments of this disclosure, the recorded access information can be parsed to extract the structured operation information from the access text and obtain the key information required to access hardware resources. This allows the operation instruction information to be converted into actual hardware access operations in the subsequent process, enabling access to hardware resources in user mode. Structured extraction can improve the efficiency and accuracy of subsequent operation instruction generation, effectively enhancing the automation and intelligence level of the verification process.

[0121] The aforementioned access information to the hardware resources of the target hardware may also include one or more of the following: memory data files, task files, and expected execution results. This access information may also include other files related to hardware resource interaction, such as those for device information acquisition and interrupt control; this disclosure does not impose any limitations on this.

[0122] Different information can be recorded in the access text for different types of operation instructions. The types of operation instructions can include one or more of the following: memory access instructions, task submission instructions, hardware information retrieval instructions, and interrupt control instructions. Memory access instructions can include one or more of the following: memory allocation instructions, memory release instructions, memory update instructions, page table mapping instructions, and page table unmapping instructions. Specifically, memory allocation instructions can be used to request memory and write data of a preset size into memory; memory release instructions can be used to release memory at a preset location; memory update instructions can be used to update memory at a preset location with data of a preset size; page table mapping instructions can be used to associate virtual memory addresses with physical memory addresses; and page table unmapping instructions can be used to unassociate virtual memory addresses with physical memory addresses.

[0123] For example, Table 2 provides examples of some of the operation instruction types and their corresponding information mentioned above. The operation instruction types and corresponding information related to hardware resource interaction during UMD execution can be recorded in the access text.

[0124] Table 2

[0125]

[0126] For memory allocation instructions, the memory segment name identifies the requested memory region, the memory size represents the required memory space, the alignment size indicates the alignment requirements for the memory start address to improve access efficiency, the buffer type indicates the purpose or type of the requested memory (e.g., data storage, instruction execution), the initial value represents the default value filled after memory allocation, and the virtual / physical address identifier indicates whether the requested memory address is a virtual or physical memory address. For memory deallocation instructions, the memory segment name identifies the memory region to be deallocated. For memory update instructions, the memory segment name identifies the memory region to be updated, the offset represents the offset from the memory segment start address, indicating the start position of the update, the size represents the length of the updated data, and the memory value represents the updated data value. For page table mapping instructions, the virtual memory segment name identifies the virtual memory region to be associated with the physical memory region, the physical memory segment name identifies the physical memory region to be associated with the virtual memory region, and the size represents the size of the memory region to be mapped. For page table unmapping instructions, the virtual memory segment name can represent the identifier of the virtual memory region to be unassociated from the physical memory region, the physical memory segment name can represent the identifier of the physical memory region to be unassociated from the virtual memory region, and the size can represent the size of the memory region to be unmapped. For task submission instructions, the task type can represent a computation task or a graphics rendering task, etc., and the task content can be a description of the task execution.

[0127] In the above information, the data sources of memory data involved in memory access instructions, such as "initial values" and "memory values," can be saved separately as memory data files, which can be saved in binary format. Register operation sequences involved in the task can also be saved separately as task files, which can be saved in .json format. The access information may also include the expected execution result, representing the correct output of the target hardware. For example, for a computation task, the expected execution result could be the computation result; for a graphics rendering task, the expected execution result could be one or more frames of images.

[0128] In one exemplary organization of access information in this disclosure, different types of access information can be categorized and stored using a multi-file separation approach. For example, a unified data file (such as a buf_data file) can be used to store various memory data files within the access information, with each memory data file storing different types of memory data involved in different memory access instructions. A dedicated access text file (such as a mr.trace file) can be used to record operation instructions related to hardware resource interaction; while task description information for different task types can be stored separately in task files (such as regs files). Furthermore, the expected execution results can also be stored separately in an independent result data file. This improves the structuring and manageability of the access information.

[0129] In modern computer architectures, GPUs typically appear as PCIe Endpoints, and the CPU can access the GPU hardware via PCIe. Therefore, all access operations to GPU memory and registers can be initiated via PCIe. Consequently, in this embodiment, to quickly verify the correctness of the GPU design without relying on KMD or firmware, a specific CPU program (such as...) has been developed. Figure 1 By using the parser and executor in the KMD (Kinetic Depository) to access PCIe, it is possible to control the GPU hardware and access its hardware resources without relying on KMD or firmware. This reduces the amount of software development required in the pre-silicon verification process.

[0130] In one possible implementation, during the process of determining the operation instruction information indicated in the word segmentation result of any text line based on the type of operation instruction, it is possible to:

[0131] When the type of operation instruction is a memory access instruction, determine one or more of the following as operation instruction information: the type of memory access instruction indicated in the word segmentation result, the information associated with the memory access instruction, and the memory data file associated with the memory access instruction.

[0132] According to the embodiments of this disclosure, accurate parsing of memory access instructions can be achieved, thereby obtaining complete operation information required to reconstruct the actual hardware access behavior, providing data support for accurate replay of the memory access process in the user-space environment, and improving the reconfigurability of the access operation.

[0133] The accessed text may include at least one text line, and each text line may correspond to an operation instruction. The text lines can be segmented by spaces, and the segmented characters can be stored in an array as array elements. The operation instruction information within a text line can then be determined based on the array elements. Specifically, the type of operation instruction can be determined based on the first element of the array, and further related operation instruction information can be determined based on the remaining array elements. This operation instruction information can be a data structure recognizable by the executor.

[0134] For a text line corresponding to an exemplary memory allocation instruction, which may include information such as memory segment name, memory size, virtual / physical address identifier, etc., the next element ('GenericBuffer0') corresponding to the type of the operation instruction can be used as the memory segment name, and at least one set of key-value pairs can be constructed in the remaining elements of the array. In this case, the fields related to the memory segment name, memory size, virtual / physical address identifier, etc., are determined in the array elements as the keys of the key-value pairs, and the next element of the element corresponding to the key is used as the value of the key-value pair.

[0135] In one possible implementation, during the process of determining the operation instruction information indicated in the word segmentation result of any text line based on the type of operation instruction, it is possible to:

[0136] When the type of operation instruction is a task submission instruction, determine one or more of the following as operation instruction information: the task type indicated in the word segmentation result, the information associated with the task submission instruction, and the task file associated with the task submission instruction.

[0137] According to the embodiments of this disclosure, the task submission instruction can be accurately parsed, thereby fully restoring the task scheduling behavior and ensuring that the task submission process can be accurately simulated in the user-space environment, thus improving the fidelity of task scheduling and the authenticity of verification.

[0138] For any given text line, the text line can be segmented into words separated by spaces to obtain an array. Based on the first array element, it can be determined that the text line corresponds to a task submission instruction. At least one set of key-value pairs can be constructed from the remaining elements in the array. In this case, fields related to the task type and task content are determined as the keys of the key-value pairs, and the next element corresponding to the key is used as the value of the key-value pair.

[0139] In one possible implementation, in step S203, the following can be done:

[0140] When the operation instruction information is a memory request instruction or a memory update instruction, data is written to the memory address of the target hardware via the PCIe bus based on the information associated with the memory access instruction and the memory data file associated with the memory access instruction to replay the memory operation and obtain the hardware execution result; when the operation instruction information is any one of a page table mapping instruction, a memory release instruction, or a page table demapping instruction, the hardware page table is configured via the PCIe bus based on the information associated with the memory access instruction to replay the operation and obtain the hardware execution result.

[0141] According to the embodiments of this disclosure, the playback of operations involving interaction with GPU hardware resources, such as memory allocation and release, memory read and write, and page table editing, can be implemented in user space. This allows for accurate reproduction of the user-space driver's access behavior to GPU resources without the need for kernel-space drivers or firmware, effectively supporting pre-silicon verification of GPUs and improving the flexibility and efficiency of verification.

[0142] For example, for a memory allocation instruction, the executor can send a hardware instruction to the GPU via PCIe based on the memory name, virtual address, memory size, and associated memory data file contained in the operation instruction information. The GPU will then complete the allocation of the corresponding memory and write the specified data to the corresponding memory address of the GPU to replay the memory operation.

[0143] For example, for memory update instructions, the executor can update the data at a specified memory address in the GPU via PCIe based on the target memory name, starting offset address, write data size and corresponding memory data file specified in the operation instruction information, thus completing the corresponding write operation.

[0144] For example, for memory release instructions, the executor can send instructions to the GPU via PCIe based on the memory name identified in the operation instruction information, release the corresponding memory resources on the GPU side, and unconfigure the related hardware page table.

[0145] In one possible implementation, in step S203, the following can be done:

[0146] When the operation instruction information is a task submission instruction, the register configuration sequence is obtained by parsing the task file associated with the task submission instruction. The register configuration sequence is an instruction used to access hardware registers. Based on the task type and the information associated with the task submission instruction, the register configuration sequence is sent to the target hardware through the PCIe bus to replay the task submission and obtain the hardware execution result. The replay process does not require firmware to participate in the scheduling.

[0147] According to the embodiments of this disclosure, the execution results of the GPU can be obtained through operations such as user-mode execution environment playback task submission and other interactions with GPU hardware resources without the need for kernel-mode drivers or firmware. This makes the pre-silicon verification of the embodiments of this disclosure applicable to various use cases such as graphics rendering and computing, and makes the pre-silicon verification process more efficient.

[0148] The executor can call the Application Programming Interface (API) provided by the GPU. The API functions send hardware instructions (including the aforementioned hardware instructions for memory access and instructions for accessing hardware registers) to the GPU via the PCIe bus to enable the interaction between the executor and the GPU hardware resources.

[0149] For example, given any task submission instruction, the executor can manipulate the content of the execution information to schedule a compute task to the GPU. Subsequent operations can only continue after the task has been executed by the GPU. When executing the task, the predefined register operations in the task file can be converted into a register configuration sequence as instructions for accessing hardware registers. This register configuration sequence is then sent to the GPU via the PCIe bus to replay the task submission and execute the compute task.

[0150] In one possible implementation, in step S204, the following can be done:

[0151] By comparing the hardware execution result with the expected execution result, and in the case where there are inconsistent data items between the hardware execution result and the expected execution result, the location information and content of the inconsistent data items are determined;

[0152] Add the location information and content of inconsistent data items to the test results of the target hardware.

[0153] According to embodiments of this disclosure, potential design problems can be quickly located and tracked by utilizing the location information and content of the identified inconsistent data items, thereby achieving efficient verification of the correctness of GPU design in pre-silicon verification scenarios and significantly improving testing efficiency and the accuracy of problem identification.

[0154] Inconsistent data items can refer to inconsistent calculated values, image pixel values, etc. For example, for computational tasks, the calculated values ​​in the hardware execution result can be compared byte-by-byte with the expected values ​​in the expected execution result; for graphics rendering tasks, the image in the hardware execution result can be compared byte-by-byte with the pixel values ​​at corresponding positions in the expected execution result. Furthermore, the test results of the target hardware can be visualized, allowing developers to intuitively identify the location and corresponding error content, thereby optimizing the GPU design.

[0155] Figure 3 A structural diagram of a hardware resource scheduling apparatus according to an embodiment of the present disclosure is shown. Figure 3 As shown, the device includes:

[0156] The first determining module 301 is used to record in real time the access information of hardware resources of the target hardware during the execution of the user-mode driver. The target hardware includes a graphics processing unit (GPU). The user-mode driver is pre-developed and associated with the feature to be tested in the target hardware.

[0157] The second determining module 302 is used to parse the recorded access information to obtain operation instruction information;

[0158] The third determining module 303 is used to access hardware resources by replaying operation instruction information and obtain hardware execution results. The process of replaying operation instruction information is implemented by sending instructions to the high-speed serial bus PCIe in user mode, without the need for firmware or kernel mode drivers to participate.

[0159] The fourth determination module 304 is used to determine the test results of the target hardware based on the hardware execution results and the expected execution results.

[0160] In one possible implementation, the access information includes access text, and the second determining module 302 is used for:

[0161] Each line of text in the accessed record is segmented into words to obtain the segmentation result of each line of text, and each line of text corresponds to an operation instruction.

[0162] Based on the word segmentation result of any text line, determine the type of operation instruction corresponding to any text line;

[0163] Based on the type of operation instruction, determine the operation instruction information indicated in the word segmentation result of any text line.

[0164] In one possible implementation, the type of operation instruction includes memory access instructions. Based on the type of operation instruction, the operation instruction information indicated in the word segmentation result of any text line is determined, including:

[0165] When the type of operation instruction is a memory access instruction, determine one or more of the following as operation instruction information: the type of memory access instruction indicated in the word segmentation result, the information associated with the memory access instruction, and the memory data file associated with the memory access instruction.

[0166] In one possible implementation, the type of operation instruction also includes a task submission instruction. Based on the type of operation instruction, the operation instruction information indicated in the word segmentation result of any text line is determined, including:

[0167] When the type of operation instruction is a task submission instruction, determine one or more of the following as operation instruction information: the task type indicated in the word segmentation result, the information associated with the task submission instruction, and the task file associated with the task submission instruction.

[0168] In one possible implementation, the third determining module 303 is used for:

[0169] When the operation instruction information is a memory request instruction or a memory update instruction, data is written to the memory address of the target hardware through the PCIe bus based on the information associated with the memory access instruction and the memory data file associated with the memory access instruction to replay the memory operation and obtain the hardware execution result.

[0170] When the operation instruction information is any one of page table mapping instruction, memory release instruction, or page table demapping instruction, the hardware page table is configured via the PCIe bus based on the information associated with the memory access instruction to replay the operation and obtain the hardware execution result.

[0171] In one possible implementation, the third determining module 303 is used for:

[0172] When the operation instruction information is a task submission instruction, the register configuration sequence is obtained by parsing the task file associated with the task submission instruction. The register configuration sequence is an instruction used to access hardware registers.

[0173] Based on the task type and information associated with the task submission instruction, the register configuration sequence is sent to the target hardware via the PCIe bus to obtain the hardware execution result. The playback process does not require firmware to participate in the scheduling.

[0174] In one possible implementation, the fourth determining module 304 is used for:

[0175] By comparing the hardware execution result with the expected execution result, and in the case where there are inconsistent data items between the hardware execution result and the expected execution result, the location information and content of the inconsistent data items are determined;

[0176] Add the location information and content of inconsistent data items to the test results of the target hardware.

[0177] According to embodiments of this disclosure, by recording in real-time the access information of hardware resources of the target hardware, including a GPU, during the execution of the user-mode driver, the access information of the target hardware, including a GPU, can be obtained. The user-mode driver is pre-developed and associated with the feature to be tested in the target hardware. This allows for obtaining relevant access information of the feature to be tested in the target hardware without developing only the user-mode driver, thus improving verification efficiency. Through the process of recording access information, the user-mode driver's operations on hardware resources can be captured, and the hardware interaction logic can be stored independently. Furthermore, by parsing the recorded access information, operation instruction information is obtained. By replaying the operation instruction information, hardware resources are accessed to obtain the hardware execution result. The process of replaying the operation instruction information can be implemented by sending instructions to the PCIe bus in user mode, without the need for kernel-mode drivers or firmware. This complete recording and playback logic breaks the existing strong dependency relationship, achieving decoupling of user-mode and kernel-mode driver development and scheduling of GPU hardware resources. Among these advantages, the absence of kernel-mode drivers reduces software development costs in the pre-silicon verification phase. Furthermore, the absence of firmware development in the pre-silicon verification phase allows for the replay of operation instructions in user mode to access hardware resources and obtain hardware execution results, thereby optimizing hardware resource scheduling mechanisms and improving testing flexibility and verification efficiency.

[0178] Figure 4 A schematic diagram of the structure of a hardware resource scheduling system according to an embodiment of the present disclosure is shown. Figure 4 As shown, the system includes a parser and an executor.

[0179] The parser is used for:

[0180] It records in real time the access information of hardware resources of the target hardware during the execution of the user-mode driver. The target hardware includes the graphics processing unit (GPU). The user-mode driver is pre-developed and associated with the feature to be tested in the target hardware.

[0181] The operation instruction information is obtained by parsing the recorded access information;

[0182] The actuator is used for:

[0183] Hardware resources are accessed by replaying operation instruction information to obtain hardware execution results. The process of replaying operation instruction information is achieved by sending instructions to the high-speed serial bus PCIe in user space, without the need for firmware or kernel-mode drivers.

[0184] The test results for the target hardware are determined based on the hardware execution results and the expected execution results.

[0185] In one possible implementation, the access information includes access text. Based on the recorded access information, parsing is performed to obtain operation instruction information, including:

[0186] Each line of text in the accessed record is segmented into words to obtain the segmentation result of each line of text, and each line of text corresponds to an operation instruction.

[0187] Based on the word segmentation result of any text line, determine the type of operation instruction corresponding to any text line;

[0188] Based on the type of operation instruction, determine the operation instruction information indicated in the word segmentation result of any text line.

[0189] In one possible implementation, the type of operation instruction includes memory access instructions. Based on the type of operation instruction, the operation instruction information indicated in the word segmentation result of any text line is determined, including:

[0190] When the type of operation instruction is a memory access instruction, determine one or more of the following as operation instruction information: the type of memory access instruction indicated in the word segmentation result, the information associated with the memory access instruction, and the memory data file associated with the memory access instruction.

[0191] In one possible implementation, the type of operation instruction also includes a task submission instruction. Based on the type of operation instruction, the operation instruction information indicated in the word segmentation result of any text line is determined, including:

[0192] When the type of operation instruction is a task submission instruction, determine one or more of the following as operation instruction information: the task type indicated in the word segmentation result, the information associated with the task submission instruction, and the task file associated with the task submission instruction.

[0193] In one possible implementation, hardware resources are accessed by replaying operation instruction information to obtain the hardware execution result, including:

[0194] When the operation instruction information is a memory request instruction or a memory update instruction, data is written to the memory address of the target hardware through the PCIe bus based on the information associated with the memory access instruction and the memory data file associated with the memory access instruction to replay the memory operation and obtain the hardware execution result.

[0195] When the operation instruction information is any one of page table mapping instruction, memory release instruction, or page table demapping instruction, the hardware page table is configured via the PCIe bus based on the information associated with the memory access instruction to replay the operation and obtain the hardware execution result.

[0196] In one possible implementation, hardware resources are accessed by replaying operation instruction information to obtain the hardware execution result, including:

[0197] When the operation instruction information is a task submission instruction, the register configuration sequence is obtained by parsing the task file associated with the task submission instruction. The register configuration sequence is an instruction used to access hardware registers.

[0198] Based on the task type and information associated with the task submission instruction, the register configuration sequence is sent to the target hardware via the PCIe bus to replay the task submission and obtain the hardware execution result. The replay process does not require firmware to participate in the scheduling.

[0199] In one possible implementation, the test results of the target hardware are determined based on the hardware execution results and the expected execution results, including:

[0200] By comparing the hardware execution result with the expected execution result, and in the case where there are inconsistent data items between the hardware execution result and the expected execution result, the location information and content of the inconsistent data items are determined;

[0201] Add the location information and content of inconsistent data items to the test results of the target hardware.

[0202] According to embodiments of this disclosure, by having a parser record in real time the access information of hardware resources of the target hardware during the execution of the user-mode driver, the target hardware including a GPU, wherein the user-mode driver is pre-developed and associated with the feature to be tested in the target hardware, it is possible to obtain the relevant access information of the feature to be tested in the target hardware without developing only the user-mode driver, thereby improving verification efficiency. By recording access information as described above, user-mode driver operations on hardware resources can be captured, and hardware interaction logic can be stored independently. Furthermore, the executor parses the recorded access information to obtain operation instruction information. By replaying the operation instruction information to access hardware resources, the hardware execution result is obtained. This process of replaying operation instruction information can be achieved by sending instructions to the PCIe bus in user mode, without the need for kernel-mode drivers or firmware. This complete recording and playback logic breaks the existing strong dependencies, decoupling the development of user-mode and kernel-mode drivers and enabling the scheduling of GPU hardware resources. Since no kernel-mode driver needs to be developed, the software development cost in the pre-silicon verification stage can be reduced. Furthermore, since no firmware program needs to be developed in the pre-silicon verification stage, the hardware resource scheduling mechanism can be optimized by replaying operation instruction information in user mode to access hardware resources and obtain hardware execution results. This improves testing flexibility and verification efficiency.

[0203] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0204] This disclosure also provides a hardware resource scheduling device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above method.

[0205] This disclosure also provides a non-volatile computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the above-described method.

[0206] This disclosure also provides a computer program product, including a computer program or a non-volatile computer-readable storage medium carrying the computer program, wherein the computer program, when executed by a processor, implements the steps of the above method.

[0207] Figure 5 This is a block diagram illustrating an apparatus 1900 for scheduling hardware resources according to an exemplary embodiment. For example, apparatus 1900 may be provided as a server or terminal device. (Refer to...) Figure 5 The apparatus 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by memory 1932 for storing instructions, such as application programs, that can be executed by the processing component 1922. The application programs stored in memory 1932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 1922 is configured to execute instructions to perform the methods described above.

[0208] Device 1900 may also include a power supply component 1926 configured to perform power management of device 1900, a wired or wireless network interface 1950 configured to connect device 1900 to a network, and an input / output interface 1958 (I / O interface). Device 1900 can operate on an operating system, such as Windows Server, stored in memory 1932. TM macOS X TM Unix TM Linux TM FreeBSD TM Or similar.

[0209] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by a processing component 1922 of the device 1900 to perform the above-described method.

[0210] Computer-readable storage media can be tangible devices capable of holding and storing programs / instructions used by instruction execution devices. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0211] The computer program (or computer-readable program instructions) described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage medium in the respective computing / processing device.

[0212] The computer program (or computer program instructions) used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing state information from the computer-readable program instructions to implement various aspects of this disclosure.

[0213] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0214] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0215] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0216] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0217] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A hardware resource scheduling method, characterized in that, The method includes: The system records in real time the access information of hardware resources of the target hardware during the execution of the user-mode driver. The target hardware includes a GPU simulation model in the pre-silicon verification stage. The user-mode driver is pre-developed and associated with the features to be tested in the target hardware. The recorded access information is parsed to obtain operation instruction information; The hardware resources are accessed by replaying the operation instruction information to obtain the hardware execution result. The process of replaying the operation instruction information is implemented by sending instructions to the high-speed serial bus PCIe in user mode, without the need for firmware or kernel mode drivers. The test results of the target hardware are determined based on the hardware execution results and the expected execution results.

2. The method according to claim 1, characterized in that, The access information includes access text, and the parsing of the recorded access information to obtain operation instruction information includes: Each line of text in the recorded access text is segmented into words to obtain the segmentation result of each line of text, and each line of text corresponds to an operation instruction. Based on the word segmentation result of any text line, determine the type of operation instruction corresponding to the text line; Based on the type of the operation instruction, determine the operation instruction information indicated in the word segmentation result of any text line.

3. The method according to claim 2, characterized in that, The type of the operation instruction includes memory access instructions. Determining the operation instruction information indicated in the word segmentation result of any text line based on the type of the operation instruction includes: If the type of the operation instruction is the memory access instruction, determine one or more of the following as the operation instruction information: the type of the memory access instruction indicated in the word segmentation result, the information associated with the memory access instruction, and the memory data file associated with the memory access instruction.

4. The method according to claim 3, characterized in that, The types of operation instructions also include task submission instructions. The step of determining the operation instruction information indicated in the word segmentation result of any text line based on the type of the operation instruction includes: If the type of the operation instruction is the task submission instruction, determine one or more of the following as the operation instruction information: the task type indicated in the word segmentation result, the information associated with the task submission instruction, and the task file associated with the task submission instruction.

5. The method according to claim 4, characterized in that, The step of accessing the hardware resources by replaying the operation instruction information to obtain the hardware execution result includes: When the operation instruction information is a memory request instruction or a memory update instruction, based on the information associated with the memory access instruction and the memory data file associated with the memory access instruction, data is written to the memory address of the target hardware via the PCIe bus to replay the memory operation and obtain the hardware execution result. When the operation instruction information is any one of page table mapping instruction, memory release instruction, or page table demapping instruction, the hardware page table is configured via the PCIe bus based on the information associated with the memory access instruction to replay the operation and obtain the hardware execution result.

6. The method according to claim 4, characterized in that, The step of accessing the hardware resources by replaying the operation instruction information to obtain the hardware execution result includes: When the operation instruction information is a task submission instruction, a register configuration sequence is obtained by parsing the task file associated with the task submission instruction. The register configuration sequence is an instruction for accessing hardware registers. Based on the task type and the information associated with the task submission instruction, the register configuration sequence is sent to the target hardware via the PCIe bus to replay the task submission and obtain the hardware execution result. The replay process does not require firmware to participate in the scheduling.

7. The method according to any one of claims 1-6, characterized in that, Determining the test result of the target hardware based on the hardware execution result and the expected execution result includes: By comparing the hardware execution result with the expected execution result, if there are inconsistent data items between the hardware execution result and the expected execution result, the location information and content of the inconsistent data items are determined. The location information and content of the inconsistent data items are added to the test results of the target hardware.

8. A hardware resource scheduling device, characterized in that, The device includes: The first determining module is used to record in real time the access information of hardware resources of the target hardware during the execution of the user-mode driver. The target hardware includes a GPU simulation model in the pre-silicon verification stage. The user-mode driver is pre-developed and associated with the test feature in the target hardware. The second determining module is used to parse the recorded access information to obtain operation instruction information; The third determining module is used to access the hardware resources by replaying the operation instruction information to obtain the hardware execution result. The process of replaying the operation instruction information is implemented by sending instructions to the high-speed serial bus PCIe in user mode, without the need for firmware or kernel mode drivers to participate. The fourth determining module is used to determine the test result of the target hardware based on the hardware execution result and the expected execution result.

9. A hardware resource scheduling system, characterized in that, The system includes a parser and an executor. The parser is used for: The system records in real time the access information of hardware resources of the target hardware during the execution of the user-mode driver, the target hardware including the graphics processing unit (GPU), and the user-mode driver is pre-developed and associated with the feature to be tested in the target hardware. The recorded access information is parsed to obtain operation instruction information; The actuator is used for: The hardware resources are accessed by replaying the operation instruction information to obtain the hardware execution result. The process of replaying the operation instruction information is implemented by sending instructions to the high-speed serial bus PCIe in user mode, without the need for firmware or kernel mode drivers. The test results of the target hardware are determined based on the hardware execution results and the expected execution results.

10. A hardware resource scheduling device, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.

11. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

12. A computer program product comprising a computer program, or a non-volatile computer-readable storage medium carrying a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • NVMe solid state disk test module and test method

    CN113821393A