System and Method for Data Interaction, Electronic Device, and Storage Medium

The data management module obtains and divides GPU data blocks in advance, and uses multi-channel parallel transmission and dynamic prefetching algorithms to solve the problem that BMC cannot monitor the GPU operating status in time, achieving efficient data interaction and monitoring.

CN120029856BActive Publication Date: 2025-07-08INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510490580.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-08
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

The substrate management controller (BMC) cannot monitor the operating status of the graphics processor (GPU) in a timely and accurate manner, especially when there are too many GPUs, the data interaction time is too long, which affects the stability and performance of the server system.

Method used

The data management module obtains operational parameter data from multiple GPUs in advance, divides it into multiple data blocks, encapsulates it into parameter data packets according to the preset data format, and transmits it to the data management module in parallel through multiple channels. After receiving the query instruction, the BMC extracts cached data based on the dynamic prefetch algorithm.

Benefits of technology

It reduces the time for BMC to interact with GPU to obtain data, improves data acquisition efficiency, ensures timely monitoring of GPU operating status, improves BMC's operating performance and data transmission speed, and avoids excessive use of memory resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029856B_ABST
    Figure CN120029856B_ABST
Patent Text Reader

Abstract

The present disclosure provides a system and method for data interaction, an electronic device, and a storage medium. Before the baseboard management controller obtains the operating status parameter data from multiple graphics processors, a data management module obtains the operating parameter data from the multiple graphics processors; the multiple graphics processors split the operating parameter data to generate multiple data blocks, encapsulate the multiple data blocks into multiple parameter data packets, and transmit the multiple parameter data packets to the data management module through multi-channel parallel transmission; after receiving a query instruction for obtaining the operating status data of the graphics processor sent by the baseboard management controller, the data management module extracts the cached operating parameter data based on a dynamic prefetching algorithm and sends the cached operating parameter data back to the baseboard management controller. Compared with the related art, the present disclosure can save the time for the baseboard management controller to directly interact with the graphics processor to obtain data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of servers, and particularly to a system and method for data interaction, an electronic device, and a storage medium. Background Art

[0002] With the development of the mobile Internet, the demand for servers has climbed. As an independent system on the server, the Baseboard Management Controller (BMC) is responsible for functions such as monitoring the hardware status, remote management, and fault alarm. The Graphics Processing Unit (GPU) is an important computing resource, and its operating status is related to the system stability and performance. The BMC needs to monitor key parameters of the GPU, such as dynamic information like temperature, usage rate, power consumption, video memory usage, fan speed, etc., and static information like serial number, bus topology information, etc.

[0003] Currently, the BMC obtains GPU data through a polling method, that is, sending commands to each GPU separately, and the GPU returns data after receiving and processing. However, it takes a long time for the GPU to receive, process, and then return data. When the BMC monitors all GPUs, it needs to interact with each GPU in turn. When the number of GPUs increases, the total interaction time increases significantly, resulting in the BMC being unable to monitor the GPU operating status in a timely and accurate manner, seriously affecting the stability and performance of the server system. Summary of the Invention

[0004] The present disclosure provides a system and method for data interaction, an electronic device, and a storage medium. Its main purpose is to solve the problem that the Baseboard Management Controller cannot monitor the GPU operating status in a timely and accurate manner.

[0005] According to a first aspect of the present disclosure, a system for data interaction is provided, including: a Baseboard Management Controller, a data management module, and multiple Graphics Processing Units;

[0006] Before the Baseboard Management Controller obtains the operating status parameter data from the multiple Graphics Processing Units, the data management module obtains the operating parameter data from the multiple Graphics Processing Units;

[0007] The multiple Graphics Processing Units divide the operating parameter data into multiple data blocks, encapsulate the multiple data blocks into multiple parameter data packets according to a preset data format, and transmit the multiple parameter data packets to the data management module through multi-channel parallel transmission;

[0008] After receiving the query instruction for obtaining the Graphics Processing Unit operating status data sent by the Baseboard Management Controller, the data management module extracts the cached operating parameter data based on a dynamic prefetching algorithm and sends the cached operating parameter data back to the Baseboard Management Controller.

[0009] Optionally, the data management module includes a plurality of data processing units, and one data processing unit is provided on each graphics processor; the data processing unit includes a data processing subunit and a storage subunit;

[0010] The data processing subunit sends a pre-query instruction to the graphics processor at a sampling interval to pre-obtain the operating parameter data at different time periods, and sends the operating parameter data to the storage subunit;

[0011] The storage subunit receives and caches the operating parameter data;

[0012] The data processing subunit responds to the query instruction sent by the baseboard management controller, obtains the operating parameter data from the storage subunit, and sends it to the baseboard management controller.

[0013] Optionally, the baseboard management controller is further configured to send query instructions to the plurality of data processing units at a sampling interval to obtain the operating parameter data of the plurality of graphics processors.

[0014] Optionally, the data management module includes at least one processor management unit provided in the baseboard management controller; the processor management unit includes a processor management unit;

[0015] The processor management unit sends a pre-query instruction to the plurality of graphics processors at a sampling interval to pre-obtain the operating parameter data at different time periods, and caches the operating parameter data of the plurality of graphics processors into the memory of the data management module based on the page-locked memory algorithm;

[0016] The processor management unit responds to the query instruction sent by the baseboard management controller, extracts the operating parameter data of the plurality of graphics processors from the memory of the data management module, and sends it to the baseboard management controller.

[0017] Optionally, the operating parameter data of the plurality of graphics processors is stored in the storage partition corresponding to each graphics processor.

[0018] Optionally, the baseboard management controller is further configured to send query instructions to at least one processor management unit at a sampling interval to obtain the operating parameter data of the plurality of graphics processors.

[0019] Optionally, the baseboard management controller is further configured to store the status data in a preset database; in response to a data display instruction, perform data display on the status data in the preset database.

[0020] Optionally, the plurality of graphics processors obtain a plurality of data blocks after splitting the operating parameter data, and perform compression processing on the plurality of data blocks.

[0021] According to a second aspect of the present disclosure, a method for data interaction is provided, including:

[0022] The control data management module sends a pre-query instruction to multiple image processors at a sampling interval;

[0023] Respectively perform segmentation processing on the operation parameter data of multiple graphics processors, and transmit the segmented operation parameter data to the data management module through multi-channel parallel transmission, and perform data caching on the operation parameter data;

[0024] Control the baseboard management controller to send a query instruction to the data management module at a sampling interval, and extract the operation parameter data corresponding to multiple image processors cached in the data management module;

[0025] Send the operation parameter data corresponding to multiple image processors to the baseboard management controller.

[0026] Optionally, respectively perform segmentation processing on the operation parameter data of multiple graphics processors, and transmit the segmented operation parameter data to the data management module through multi-channel parallel transmission, including:

[0027] Segment the operation parameter data into multiple data chunks, and perform compression processing on the data chunks;

[0028] Encapsulate the compressed multiple data chunks into multiple parameter data packets according to a preset data format;

[0029] Transmit the multiple parameter data packets to the data management module through multi-channel parallel transmission.

[0030] Optionally, extracting the operation parameter data corresponding to multiple image processors cached in the data management module includes:

[0031] Construct a probability matrix of the operation parameter data based on the historical data access sequence;

[0032] According to the target data block identifier with a transition probability value greater than a preset threshold in the probability matrix, preload the operation parameter data corresponding to the target data block identifier.

[0033] Optionally, the method further includes:

[0034] If the operation parameter data is stored in the target memory, cache the operation parameter data based on the page-locked memory algorithm.

[0035] According to the third aspect of the present disclosure, there is provided an electronic device, including:

[0036] At least one processor; and

[0037] A memory communicatively connected to the at least one processor; wherein,

[0038] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the data interaction method described in the foregoing second aspect.

[0039] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the data interaction method described in the foregoing second aspect.

[0040] According to a fifth aspect of the present disclosure, there is provided a computer program product including a computer program, and the computer program implements the data interaction method described in the foregoing second aspect when executed by a processor.

[0041] The present disclosure provides a data interaction system and method, an electronic device, and a storage medium. The data interaction system includes: a baseboard management controller, a data management module, and multiple graphics processors; before the baseboard management controller obtains the operating status parameter data from the multiple graphics processors, the data management module obtains the operating parameter data from the multiple graphics processors; the multiple graphics processors divide the operating parameter data into multiple data blocks, encapsulate the multiple data blocks into multiple parameter data packets according to a preset data format, and transmit the multiple parameter data packets to the data management module through multi-channel parallel transmission; after receiving the query instruction for obtaining the graphics processor operating status data sent by the baseboard management controller, the data management module extracts the cached operating parameter data based on the dynamic prefetching algorithm and sends the cached operating parameter data back to the baseboard management controller. Compared with the related art, in the present disclosure, before the BMC obtains the operating status parameter data from the multiple GPUs, the data management module obtains the operating parameter data from the GPUs in advance. The GPUs divide the operating parameter data into multiple data blocks, encapsulate them into parameter data packets, and then transmit them to the data management module through multi-channel parallel transmission. This method greatly saves the time for the BMC to directly interact with the GPUs to obtain data. In the traditional method, it takes a long time for the BMC to poll each GPU to obtain data, while in this system, the BMC can quickly obtain the processed data from the data management module and timely obtain the GPU operating status parameters, improving the data acquisition efficiency. The multiple graphics processors divide the data into data blocks, encapsulate them into data packets according to a preset data format, and use multi-channel parallel transmission to transmit them to the data management module. Multi-channel parallel transmission makes full use of the transmission channel resources. Compared with single-channel transmission, it can transmit more data in the same time, reduce the data transmission delay, improve the overall bandwidth and speed of data transmission, enable the data management module to obtain the operating parameter data of the GPUs faster, and provide more timely data support for the BMC. After the BMC receives the query instruction sent by the data management module, the data management module extracts the cached operating parameter data based on the dynamic prefetching algorithm and sends it back to the BMC. The dynamic prefetching algorithm analyzes the historical data access sequence, predicts the data blocks that the BMC may access, and prefetches them into the cache in advance. This enables the BMC to obtain data without waiting for the data management module to read from the original storage, reduces the waiting time of the BMC for data, improves the operating performance of the BMC, saves the memory space of the BMC at the same time, and avoids the occupation of memory resources by frequent data reading operations. During the operation of the server, there is a data acquisition blank period in the traditional BMC monitoring GPU method before the system starts. In this system, the data management module obtains the GPU operating parameter data in advance and starts collecting data before the BMC starts. When the BMC starts and sends a query instruction, the data management module can immediately provide data, reducing the blank period of GPU-related parameter collection, ensuring the continuity and integrity of the BMC's monitoring of the GPU operating status, and enabling the BMC to more comprehensively and accurately grasp the operating conditions of the GPUs.

[0042] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0044] Figure 1 is a schematic structural diagram of a data interaction system provided by an embodiment of the present disclosure;

[0045] Figure 2 is a schematic structural diagram of another data interaction system provided by an embodiment of the present disclosure;

[0046] Figure 3 is a schematic structural diagram of another data interaction system provided by an embodiment of the present disclosure;

[0047] Figure 4 is a schematic flowchart of a data interaction method provided by an embodiment of the present disclosure;

[0048] Figure 5 is a schematic flowchart of another data interaction method provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0050] The following describes a data interaction system, method, electronic device, and storage medium according to embodiments of the present disclosure with reference to the accompanying drawings.

[0051] Figure 1 is a schematic structural diagram of a data interaction system provided by an embodiment of the present disclosure. As Figure 1 shown, the system includes: a baseboard management controller 11, a data management module 12, and a plurality of graphics processors 13;

[0052] The data management module 12 obtains operation parameter data from the plurality of graphics processors 13 before the baseboard management controller 11 obtains operation status parameter data from the plurality of graphics processors 13.

[0053] In an embodiment of the present disclosure, before the baseboard management controller 11 starts the operation of obtaining the operating status parameter data from multiple graphics processors 13, the data management module 12 actively initiates the data interaction process with the multiple graphics processors 13. The purpose is to pre-obtain various parameter data during the operation of the graphics processors 13. These data cover dynamic information such as the temperature, usage rate, power consumption, video memory usage, and fan speed of the graphics processors 13, as well as static information such as serial numbers, i2c topology information, vendor ID, and device ID. They are of great significance for comprehensively and accurately monitoring the operating status of the graphics processors 13. The data management module 12 has the hardware and software conditions to implement the above functions. From a hardware perspective, it has interfaces capable of data interaction with the graphics processors 13. These interfaces can be high-speed data transmission interfaces such as PCIE interfaces according to different implementation methods, ensuring the stability and efficiency of data transmission. From a software level, the data management module 12 is installed with a management system with specific functions. This management system has a polling mechanism. Through this mechanism, the data management module 12 can send commands to obtain data to each graphics processor 13 in sequence according to the set rules. The data management module is communicatively connected to the multiple graphics processors through a first interaction interface, and the data management module is communicatively connected to the baseboard management controller through a second interaction interface. During the data acquisition process, a specific data interaction protocol is followed between the data management module 12 and the graphics processors 13.

[0054] The multiple graphics processors 13 split the operation parameter data into multiple data blocks, encapsulate the multiple data blocks into multiple parameter data packets according to a preset data format, and transmit the multiple parameter data packets to the data management module 12 through multi-channel parallel transmission.

[0055] In an embodiment of the present disclosure, the graphics processors 13 perform specific data processing and transmission operations in the data transmission link, that is, split, encapsulate the operation parameter data, and transmit it to the data management module 12 through multi-channel parallel transmission. This series of operations can achieve efficient data interaction.

[0056] From a technical principle perspective, the amount of operation parameter data generated by multiple graphics processors 13 during operation is huge and complex. To ensure that data can be transmitted to the data management module 12 efficiently and stably, the graphics processors 13 first divide the operation parameter data into multiple data blocks according to a predefined chunking algorithm. This chunking algorithm is carefully designed to divide large data into fixed-size blocks, such as specific values between 4KB and 1MB. This range is set based on a comprehensive consideration of memory management and data transmission efficiency. Transmitting data in block order effectively reduces memory fragmentation and random access overhead, ensuring the stability and continuity of data storage and transmission. After data chunking is completed, the graphics processors 13 encapsulate the multiple data blocks into multiple parameter data packets according to a preset data format. This preset data format is clearly defined during the system design phase, aiming to unify the data structure and standard, facilitating data identification, parsing, and processing between different components. Through standardized data encapsulation, whether during data transmission or when the data management module 12 receives and processes data, the data error rate can be reduced, and the accuracy and efficiency of data processing can be improved.

[0057] In terms of data transmission, multiple graphics processors 13 use a multi-channel parallel transmission method to send multiple parameter data packets to the data management module 12. Multi-channel parallel transmission utilizes the multi-tasking engine feature of the GPU itself, divides the compressed data blocks, and distributes them to multiple CUDA streams. This method breaks the limitations of traditional single-channel transmission, allowing data to be transmitted simultaneously on multiple channels, significantly improving the overall bandwidth and speed of data transmission. Through parallel processing, the time for data to be transmitted from the graphics processors 13 to the data management module 12 is greatly shortened, enabling the data management module 12 to obtain the operation parameter data of the graphics processors 13 more timely, providing strong support for subsequent data processing and feedback to the baseboard management controller 11.

[0058] From the perspective of the integrity and logic of system operation, the data segmentation, encapsulation, and multi-channel parallel transmission operations performed by multiple graphics processors 13 closely cooperate with the functions of the data management module 12 and the baseboard management controller 11. After receiving these parameter data packets, the data management module 12 stores and processes them, and then, according to the query instructions of the baseboard management controller 11, feeds back the processed data to the baseboard management controller 11, thereby achieving comprehensive monitoring of the operating status of the graphics processors 13. This series of operations together constitute a complete and efficient data interaction process, which can optimize the data interaction process and improve system performance.

[0059] After the data management module 12 receives the query instruction sent by the baseboard management controller 11 to obtain the operating status data of the graphics processor 13, it extracts the cached operating parameter data based on the dynamic prefetching algorithm and sends the cached operating parameter data back to the baseboard management controller 11.

[0060] In an embodiment of the present disclosure, when the data management module 12 receives the query instruction sent by the baseboard management controller 11 to obtain the operating status data of the graphics processor 13, it extracts the local cached data through the dynamic prefetching algorithm and transmits it back.

[0061] During the operation of the system, the data management module 12 continuously collects and stores the operating parameter data obtained from the graphics processor 13. These data are cached in the local storage area of the data management module 12 in an orderly manner. This storage area has the characteristics of fast reading and writing to enable a quick response when receiving a query instruction. When the baseboard management controller 11 issues a query instruction to obtain the operating status data of the graphics processor 13, the data management module 12 first parses and identifies the instruction to confirm the data type and range requested by the instruction. Based on the dynamic prefetching algorithm, the data management module 12 starts to perform the data extraction operation. The dynamic prefetching algorithm realizes the prefetching of target data through in-depth analysis of the historical data access sequence. For example, by counting the reading order of the baseboard management controller 11 for various sensor data in the past, such as the reading order of sensor A → B → C, a state transition probability matrix is constructed. This matrix can accurately reflect the probability relationship of different data blocks being accessed. Based on this matrix, the data management module 12 can predict the data block that the baseboard management controller 11 may access next.

[0062] Based on the prediction, the data management module 12 prefetches the data blocks with a high probability of being accessed into the local cache in advance. When receiving a query instruction, the relevant operating parameter data is preferentially extracted from the cache. This method greatly reduces the data reading latency because the cached data can be quickly obtained without having to read from a relatively slow conventional storage medium. If all the data that meets the requirements of the query instruction exists in the cache, the data management module 12 will directly encapsulate and organize these cached operating parameter data according to the established data transmission protocol.

[0063] In the data feedback stage, the data management module 12 sends the sorted operation parameter data back to the baseboard management controller 11 through the communication link connected to the baseboard management controller 11. The selection of the communication link and the formulation of the data transmission protocol have been determined in the system design stage to ensure the stability, accuracy, and efficiency of data transmission. During the transmission process, the data management module 12 performs data verification and error correction processing to ensure that the data received by the baseboard management controller 11 is complete and error-free. If the data in the cache does not fully meet the requirements of the query instruction, the data management module 12 quickly reads the remaining part from the local storage, merges it with the cache data, and then sends it back to the baseboard management controller 11.

[0064] The present disclosure provides a data interaction system. In the present disclosure, before the BMC obtains the operation status parameter data from multiple GPUs, the data management module 12 obtains the operation parameter data from the GPUs in advance. The GPU divides the operation parameter data into multiple data blocks, encapsulates them into parameter data packets, and then transmits them to the data management module 12 in parallel through multiple channels. This method greatly saves the time for the BMC to directly interact with the GPUs to obtain data. In the traditional method, it takes a long time for the BMC to poll each GPU to obtain data, while in this system, the BMC can quickly obtain the processed data from the data management module 12, timely obtain the GPU operation status parameters, and improve the data acquisition efficiency. Multiple graphics processors 13 divide the data into data blocks, encapsulate them into data packets according to a preset data format, and transmit them to the data management module 12 in parallel through multiple channels. The multi-channel parallel transmission makes full use of the transmission channel resources. Compared with single-channel transmission, it can transmit more data in the same time, reduce the data transmission delay, improve the overall bandwidth and speed of data transmission, enable the data management module 12 to obtain the operation parameter data of the GPU faster, and provide more timely data support for the BMC. After the BMC receives the query instruction sent by the data management module 12, the data management module 12 extracts the operation parameter data cached locally based on the dynamic prefetching algorithm and sends it back to the BMC. The dynamic prefetching algorithm analyzes the historical data access sequence, predicts the data blocks that the BMC may access, and prefetches them into the cache in advance. This enables the BMC to obtain data without waiting for the data management module 12 to read from the original storage, reduces the waiting time of the BMC for data, improves the operation performance of the BMC, saves the memory space of the BMC at the same time, and avoids the occupation of memory resources by frequent data reading operations. During the operation of the server, there is a data acquisition blank period in the traditional BMC monitoring GPU method before the system starts. In this system, the data management module 12 obtains the GPU operation parameter data in advance and starts collecting data before the BMC starts. When the BMC starts and sends a query instruction, the data management module 12 can immediately provide data, reducing the blank period for collecting GPU-related parameters, ensuring the continuity and integrity of the BMC's monitoring of the GPU operation status, and enabling the BMC to more comprehensively and accurately grasp the operation of the GPU.

[0065] Further, in a possible implementation manner of this embodiment, as Figure 2As shown in the figure, the data management module 12 includes a plurality of data processing units 121, and one data processing unit 121 is provided on each graphics processor 13; the data processing unit 121 includes a data processing subunit 1211 and a storage subunit 1212; the data processing subunit 1211 sends a pre-query instruction to the graphics processor 13 at a sampling interval to pre-obtain operation parameter data at different time periods, and sends the operation parameter data to the storage subunit 1212; the storage subunit 1212 receives and caches the operation parameter data; the data processing subunit 1211 responds to the query instruction sent by the baseboard management controller 11, obtains the operation parameter data from the storage subunit 1212, and sends it to the baseboard management controller 11.

[0066] Specifically, in the embodiments of the present disclosure, the data management module 12 is composed of multiple data processing units 121. This distributed architecture design enables the system to have better parallel processing capabilities and scalability when processing data from multiple graphics processors 13. A corresponding data processing unit 121 is provided on each graphics processor 13. This one-to-one correspondence ensures the pertinence and efficiency of data processing and can accurately obtain the operation parameter data of each graphics processor 13. In the data acquisition stage, the data processing subunit 1211 sends a pre-query instruction to the corresponding graphics processor 13 at a preset sampling interval. This sampling interval is comprehensively determined based on factors such as the system's requirements for data real-time performance and the characteristics of the changes in the operation parameters of the graphics processor 13. By periodically sending query instructions, the data processing subunit 1211 can pre-obtain the operation parameter data of the graphics processor 13 at different time periods. These data cover dynamic information such as temperature, usage rate, power consumption, video memory usage, and fan speed, as well as static information such as the serial number of the GPU, i2c topology information, vendor ID, and device ID. After obtaining the operation parameter data, the data processing subunit 1211 sends it to the storage subunit 1212. The storage subunit 1212 serves as a temporary storage space for data and has the function of a cache. It receives and caches the operation parameter data transmitted by the data processing subunit 1211 and stores the data in an orderly manner using a specific storage algorithm to ensure the rapid retrieval and reading of the data. When the baseboard management controller 11 sends a query instruction, the data processing subunit 1211 is responsible for responding to this instruction. It first analyzes the specific content of the query instruction to clarify requirements such as the type and range of the required data. Then, according to these requirements, it obtains the corresponding operation parameter data from the storage subunit 1212. Due to the efficient storage and management of data by the storage subunit 1212, the data processing subunit 1211 can quickly and accurately retrieve the required data. After obtaining the data, the data processing subunit 1211 encapsulates and formats the operation parameter data according to the data transmission protocol agreed upon with the baseboard management controller 11 to ensure the integrity and accuracy of the data during transmission. Then, it sends the processed data to the baseboard management controller 11 through the established communication link within the system.

[0067] Further, in a possible implementation manner of this embodiment, as Figure 2 shown, the baseboard management controller 11 is further configured to send query instructions to multiple data processing units 121 respectively at the sampling interval to obtain the operation parameter data of multiple graphics processors 13.

[0068] Specifically, in the embodiments of the present disclosure, as the core component for monitoring and managing the server hardware status, the baseboard management controller 11 needs to obtain the operation parameter data of multiple graphics processors 13 to effectively monitor their operation status. To ensure the timeliness and accuracy of the obtained data, the baseboard management controller 11 follows a preset sampling interval and sends query instructions to multiple data processing units 121 respectively. This sampling interval is not set randomly but is comprehensively considered in many aspects. From the operation characteristics of the graphics processor 13, the dynamic information such as its operation parameters like temperature and utilization rate changes at different frequencies. For example, the temperature may fluctuate in real time with the load condition, and the utilization rate also varies greatly in different task phases. At the same time, the requirement of the system for data timeliness cannot be ignored. If the sampling interval is too long, the obtained data may be lagged and unable to reflect the actual operation status of the graphics processor 13 in a timely manner; if the sampling interval is too short, it may cause excessive consumption of system resources and affect the overall performance. Therefore, through the study of the operation rules of the graphics processor 13 and the evaluation of the system performance requirements, an appropriate sampling interval is determined.

[0069] When sending the query instructions, the baseboard management controller 11 uses the communication link established between it and multiple data processing units 121 for data transmission. According to the system design, this communication link adopts a specific communication protocol, such as based on the Peripheral Component Interconnect Express (PCIE) or other adapted communication standards, to ensure the stability and efficiency of the instruction transmission. The communication protocol defines key elements such as the format of the instructions, the transmission method, and the data verification rules.

[0070] After the data processing unit 121 receives the query instructions sent by the baseboard management controller 11, the internal data processing subunit 1211 will parse the instructions. The data processing subunit 1211 obtains the corresponding operation parameter data of the graphics processor 13 from the storage subunit 1212 according to the data type and range specified in the instructions. The storage subunit 1212 has cached the operation parameter data of different time periods in advance according to the instructions of the data processing subunit 1211. Due to the storage subunit 1212 adopting an efficient storage management strategy, such as specific data indexing and caching algorithms, the data processing subunit 1211 can quickly locate and read the required data.

[0071] After obtaining the data, the data processing subunit 1211 encapsulates and processes the data according to the format and transmission method agreed upon with the baseboard management controller 11. Subsequently, the data processing subunit 1211 transmits the processed data back to the baseboard management controller 11 through the original communication link.

[0072] Furthermore, in a possible implementation manner of this embodiment, such as Figure 3As shown in the figure, the data management module 12 includes at least one processor management unit 122 provided in the baseboard management controller 11; the processor management unit 122 sends pre-query instructions to multiple graphics processors 13 at a sampling interval to pre-obtain the operation parameter data at different time periods, and caches the operation parameter data of the multiple graphics processors 13 into the memory of the data management module 12 based on the page-locked memory algorithm; the processor management unit 122 responds to the query instructions sent by the baseboard management controller 11, extracts the operation parameter data of the multiple graphics processors 13 from the memory of the data management module 12, and sends it to the baseboard management controller 11.

[0073] Specifically, in the embodiments of the present disclosure, the data management module 12 includes at least one processor management unit 122 provided in the baseboard management controller 11. This architecture design aims to make full use of the hardware resources of the baseboard management controller 11 to achieve a deep integration of the data management function and the overall function of the BMC. The processor management unit 122, as the direct executor of data interaction, undertakes the important responsibilities of data acquisition, caching, and response. Compared with the data processing unit 121 corresponding to each GPU set in the above embodiments, in this embodiment, a processor management unit 122 is directly set in the baseboard management controller 11, and the processor management unit 122 manages all the graphics processors and executes functions similar to those of the data processing unit 121.

[0074] The processor management unit 122 sends pre-query instructions to multiple graphics processors 13 at a predetermined sampling interval. After obtaining the data, the processor management unit 122 caches the operation parameter data of the multiple graphics processors 13 into the data management module memory based on the page-locked memory algorithm. The page-locked memory algorithm, that is, allocates page-locked memory (Pinned Memory) through cudaHostAlloc of CUDA, bypasses the operating system paging mechanism, and directly binds the physical address. The application of this algorithm effectively reduces the number of memory copies and improves the speed of data storage and reading. Since the amount of GPU operation parameter data is large, frequent memory copy operations will not only consume a lot of time, but also may cause data transmission delays, affecting the real-time monitoring of the GPU operation status by the system. The page-locked memory algorithm directly stores the data in the physical memory and establishes a stable address mapping, enabling the processor management unit 122 to quickly store and access this data, greatly improving the data processing efficiency.

[0075] When the baseboard management controller 11 sends a query instruction, the processor management unit 122 is responsible for the response. The processor management unit 122 first parses the query instruction to clarify the data type, range, and other relevant parameters required by the instruction. Then, based on the parsing result, it extracts the corresponding operating parameter data of multiple graphics processors 13 from its memory. Since the page-locked memory algorithm was previously used for data caching, the processor management unit 122 can quickly locate and read the required data, avoiding the time overhead caused by memory paging and data searching. After extracting the data, the processor management unit 122 encapsulates and formats the data according to the data transmission protocol previously agreed upon with the baseboard management controller 11. This may include operations such as adding data verification information and packing according to a specific data format to ensure the integrity and accuracy of the data during transmission. Finally, the processor management unit 122 sends the processed data to the baseboard management controller 11 through the internal communication link of the system, enabling the BMC to obtain the operating parameter data of the GPU in a timely manner and realizing effective monitoring of the GPU operating state.

[0076] Further, in a possible implementation manner of this embodiment, as Figure 3 shown, the operating parameter data of multiple graphics processors 13 are stored in the storage partitions corresponding to their respective graphics processors 13.

[0077] Specifically, in the embodiment of the present disclosure, during the operation of the server, multiple graphics processors 13 each generate a large amount of operating parameter data, which is crucial for monitoring their own operating states and ensuring the stable operation of the entire server system. To achieve effective management and fast access to this data, each graphics processor 13 is provided with a corresponding storage partition. This correspondence is established based on specific rules of the system design. From a hardware perspective, the storage partition can physically be a part of the internal storage module of the graphics processor 13 or the space of an external storage device closely connected thereto. Logically, each storage partition is associated with the corresponding graphics processor 13 through a unique identifier, ensuring highly accurate and targeted storage and reading of data.

[0078] Further, in a possible implementation manner of this embodiment, the baseboard management controller 11 is also used to send query instructions to at least one processor management unit 122 according to a sampling interval to obtain the operating parameter data of multiple graphics processors 13.

[0079] Specifically, in the embodiments of the present disclosure, the baseboard management controller 11 sends query instructions to at least one processor management unit 122 at a predetermined sampling interval. The sampling interval is set according to the variation characteristics of the operating parameters of the graphics processing unit 13 and the real-time requirement of system data. After receiving the instruction, the processor management unit 122 extracts the operating parameter data of the corresponding multiple graphics processing units 13 from its own memory cache. These data were previously obtained by the processor management unit 122 at the sampling interval and cached using the page-locked memory algorithm. After the processor management unit 122 obtains the data, it performs encapsulation processing according to the protocol agreed with the baseboard management controller 11, and then transmits it back to the baseboard management controller 11 through the system communication link, helping it to effectively monitor the operating status of the multiple graphics processing units 13, optimizing the data interaction process between the BMC and the GPU, and improving the system performance.

[0080] Further, in a possible implementation manner of this embodiment, the baseboard management controller 11 is further configured to store the status data in a preset database; in response to a data display instruction, display the status data in the preset database.

[0081] Specifically, in the embodiments of the present disclosure, the baseboard management controller 11 stores the operating status data of the graphics processing unit 13 obtained from the data management module 12 in a preset database according to the established data storage rules. The preset database may be a specific storage area inside the baseboard management controller 11 or a specified database space in an external storage device connected thereto. When storing, the status data is classified and indexed according to information such as data type and timestamp for subsequent quick query and call. When the baseboard management controller 11 receives a data display instruction, it will parse the instruction to clarify the requirements such as the data range and format to be displayed. Subsequently, the corresponding status data is retrieved from the preset database according to the instruction requirements. During the retrieval process, the query optimization mechanism of the database, such as index query, is used to improve the data acquisition efficiency. After obtaining the data, the baseboard management controller 11 processes the status data according to the preset data display format. If it needs to be displayed through a web page, the data will be converted into a format suitable for web display such as HTML or JSON; if it is output through the redfish interface or ipmi command, the data will be encapsulated according to the corresponding interface or command specification. Finally, the processed data is presented to the user or other relevant systems in a suitable manner, realizing an intuitive display of the operating status of the graphics processing unit 13, facilitating the user to timely understand the operating conditions of the GPU, and providing support for the management and maintenance of the server system.

[0082] Further, in a possible implementation manner of this embodiment, the multiple graphics processing units 13 split the operating parameter data to obtain multiple data blocks, and perform compression processing on the multiple data blocks.

[0083] Specifically, in the embodiments of the present disclosure, the amount of operation parameter data generated by multiple graphics processors 13 is large. To facilitate transmission and processing, a segmentation operation is performed on it, splitting the large data into data chunks of a fixed size, such as 4KB - 1MB. This chunking method can reduce memory fragmentation and random access overhead, and improve the stability of data processing. After segmentation, the graphics processor 13 uses a specific compression algorithm, such as the LZ4 algorithm, to compress multiple data chunks. The compression process aims to reduce the amount of data transmission, and reduce the bandwidth and time required for transmission. Through compression, the originally large data chunks are converted into a more compact format, improving the efficiency of data transmission while maintaining the key information of the data. The compressed data is encapsulated in a specific format (such as the GPU - BIN format) during subsequent transmission, and sent to the data management module 12 in a multi - stream parallel transmission manner through a binary protocol (such as Cap’n Proto). This processing method makes the data transmission in the server system more efficient, and helps to solve the problem of slow data interaction between the BMC and the GPU.

[0084] Figure 4 It is a schematic flowchart of a data interaction method provided by an embodiment of the present disclosure.

[0085] As Figure 4 shown, the method includes the following steps:

[0086] Step 201, control the data management module to send a pre - query instruction to multiple image processors at a sampling interval.

[0087] In the embodiment of the present disclosure, the data management module assumes the important responsibility of obtaining the operating parameter data of the graphics processor. Among them, the setting of the sampling interval is based on the comprehensive consideration of multiple factors such as the operating characteristics of the graphics processor and the system's requirements for data real-time performance. For example, the frequency of change of the operating parameters of the graphics processor in different application scenarios is different. In complex graphics rendering tasks, its temperature, usage rate and other parameters change rapidly, and a shorter sampling interval is required to ensure the timeliness of data acquisition; in relatively stable operating scenarios, the sampling interval can be appropriately extended to balance the real-time performance of data acquisition and system resource consumption. When the system is running, according to the established control logic, the data management module follows the set sampling interval and sends pre-query instructions to multiple graphics processors in an orderly manner. These instructions are transmitted through specific communication interfaces and protocols, such as PCIE interfaces combined with corresponding high-speed communication protocols to ensure the stability and accuracy of instruction transmission. The instructions clearly contain key information such as the type and range of the required data, so that the graphics processor can accurately know the data content that needs to be fed back. After receiving the instruction, the graphics processor organizes the operating parameter data according to the instruction requirements and prepares for subsequent transmission to the data management module. This process lays the foundation for efficient BMC and GPU data interaction, which is in line with the original design intention of the patented technical solution to solve the problem of slow data interaction.

[0088] Step 202 , segmenting the operating parameter data of the plurality of graphics processors respectively, transmitting the segmented operating parameter data to the data management module in parallel through multiple channels, and caching the operating parameter data.

[0089] In an embodiment of the present disclosure, after receiving the pre-query instruction, multiple graphics processors will segment the operating parameter data. The segmentation is based on a specific algorithm to divide the big data into fixed-size blocks, such as 4KB-1MB, which can reduce memory fragmentation and random access overhead and improve data processing efficiency. After the segmentation is completed, the data is sent to the data management module through multi-channel parallel transmission technology. Multi-channel parallel transmission uses the multi-tasking engine of the GPU to divide the data into blocks and distribute it to multiple CUDA streams, and transmit and calculate at the same time, which greatly improves the data transmission speed.

[0090] After receiving the data, the data management module will cache the data. Using the page-locked memory algorithm, the page-locked memory is allocated through CUDA's cudaHostAlloc, bypassing the operating system paging mechanism, directly binding the physical address, reducing the number of memory copies, achieving fast storage, and facilitating the subsequent baseboard management controller to obtain data, optimizing the data interaction process between BMC and GPU, and meeting the requirements of patented technology.

[0091] Step 203: Control the baseboard management controller to send a query instruction to the data management module according to the sampling interval, and extract the operation parameter data corresponding to multiple graphics processors cached in the data management module.

[0092] In the embodiments of the present disclosure, the baseboard management controller (BMC) sends a query instruction to the data management module according to a preset sampling interval. The setting of the sampling interval comprehensively considers the change frequency of the graphics processor operation parameters and the system's demand for data real-time performance, ensuring that the data obtained by the BMC can timely reflect the operation state of the graphics processor without excessive consumption of system resources. After receiving the query instruction, the data management module parses the instruction to clarify the specific requirements for the data required by the BMC. Then, according to the instruction requirements, it extracts the operation parameter data corresponding to multiple graphics processors from its cache. The data management module adopts an efficient cache management strategy, such as an index-based data storage method, which can quickly locate and extract the data required by the BMC, improving the data acquisition efficiency. During the data extraction process, the data management module may also adopt a dynamic prefetching algorithm, such as a prefetching model based on a Markov chain. By statistically analyzing the historical data access sequence to construct a state transition probability matrix, it predicts the data blocks that the BMC may access and prefetches them into the cache in advance, further shortening the data acquisition time. Finally, the data management module sends the extracted data to the BMC according to the established data transmission protocol, enabling the BMC to obtain the operation parameter data of the graphics processor, providing data support for the server system to monitor and manage the graphics processor, and meeting the requirements of optimizing the data interaction between the BMC and the GPU in the technical solution of this patent.

[0093] Step 204: Send the operation parameter data corresponding to multiple graphics processors to the baseboard management controller.

[0094] In an embodiment of the present disclosure, after receiving a query instruction sent by the baseboard management controller and completing data extraction, the data management module sends the operation parameter data corresponding to multiple graphics processors to the baseboard management controller according to a predefined data transmission protocol. This data transmission protocol clearly defines key elements such as the data encapsulation format, transmission method, and verification rules. In terms of the encapsulation format, the data management module packs the extracted operation parameter data in a specific format. For example, different types of data (such as temperature, utilization rate, etc.) are arranged in a specified order and structure to ensure the integrity and accuracy of the data during transmission. In terms of the transmission method, a high-performance binary data serialization protocol such as Cap’n Proto is usually adopted, combined with a multi-stream parallel transmission method. Utilizing the multi-tasking engine of the GPU, the data is divided into chunks and distributed to multiple CUDA streams for parallel transmission to improve the speed and efficiency of data transmission and reduce transmission latency. At the same time, to ensure the accuracy of the data, the data management module adds verification information to the transmitted data according to the verification rules. For example, a cyclic redundancy check (CRC) algorithm is used to generate a verification code and append it to the data. After receiving the data, the baseboard management controller verifies the data according to the same verification rules. If the verification passes, the data is received; if the verification fails, the data management module is required to re-transmit the data.

[0095] The present disclosure provides a data interaction method. In the present disclosure, before the BMC obtains the operation status parameter data from multiple GPUs, the data management module obtains the operation parameter data from the GPUs in advance. The GPU divides the operation parameter data into multiple data blocks, encapsulates them into parameter data packets, and then transmits them to the data management module through multi-channel parallel transmission. This method greatly saves the time for the BMC to directly interact with the GPUs to obtain data. In the traditional method, it takes a long time for the BMC to poll each GPU to obtain data, while this system enables the BMC to quickly obtain the processed data from the data management module, timely obtain the GPU operation status parameters, and improve the data acquisition efficiency. Multiple graphics processors divide the data into data blocks, encapsulate them into data packets according to a preset data format, and use multi-channel parallel transmission to send them to the data management module. Multi-channel parallel transmission makes full use of the transmission channel resources. Compared with single-channel transmission, it can transmit more data in the same time, reduce the data transmission delay, improve the overall bandwidth and speed of data transmission, enable the data management module to obtain the operation parameter data of the GPU faster, and provide more timely data support for the BMC. After the BMC receives the query instruction sent by the data management module, the data management module extracts the operation parameter data cached locally based on the dynamic prefetching algorithm and sends it back to the BMC. The dynamic prefetching algorithm analyzes the historical data access sequence, predicts the data blocks that the BMC may access, and prefetches them into the cache in advance. This enables the BMC to obtain data without waiting for the data management module to read from the original storage, reduces the waiting time of the BMC for data, improves the operation performance of the BMC, saves the memory space of the BMC at the same time, and avoids the occupation of memory resources by frequent data reading operations. During the operation of the server, there is a data acquisition blank period in the traditional BMC monitoring GPU method before the system starts. In this system, the data management module obtains the GPU operation parameter data in advance and starts collecting data before the BMC starts. When the BMC starts and sends a query instruction, the data management module can immediately provide data, reduce the blank period of GPU-related parameter collection, ensure the continuity and integrity of the BMC's monitoring of the GPU operation status, and enable the BMC to more comprehensively and accurately grasp the operation of the GPU.

[0096] To clearly illustrate the embodiments of the present disclosure, this embodiment provides a schematic flowchart of another data interaction method.

[0097] As Figure 5 shown, this method includes the following steps:

[0098] Step 301, control the data management module to send a pre-query instruction to multiple image processors according to the sampling interval.

[0099] Specifically, in step 301, the data management module sends a pre-query instruction to multiple graphics processors according to a pre-set sampling interval. The setting of the sampling interval comprehensively considers the characteristics of the change of the operating parameters of the graphics processor and the real-time requirements of the system data, ensuring the timeliness of data acquisition and the rational use of system resources. The instruction is transmitted through a specific communication interface and protocol, accurately conveying key information such as the type and range of the required data, laying the foundation for subsequent data collection.

[0100] Step 302: divide the operating parameter data into a plurality of data blocks, and compress the data blocks.

[0101] Specifically, in step 302, after receiving the query instruction, the graphics processor processes the operating parameter data. A specific block algorithm is used to divide the data into fixed-size data blocks, generally in the range of 4KB-1MB, to reduce memory fragmentation and random access overhead. Subsequently, the data blocks are compressed using an algorithm such as LZ4 to reduce the data volume and improve subsequent transmission efficiency.

[0102] Step 303: encapsulate the compressed data into multiple parameter data packets in a preset data format.

[0103] Specifically, in step 303, the compressed data are divided into multiple parameter data packets according to the data format preset by the system. This preset format clarifies the data organization method, header and footer structure, etc., to ensure the integrity and identifiability of the data during transmission, and facilitate the data management module to receive and process. For example, a high-performance binary data serialization protocol such as Cap'n Proto is used.

[0104] Step 304: Transmit the multiple parameter data packets in parallel to the data management module via multiple channels.

[0105] Specifically, in step 304, multiple parameter data packets are transmitted to the data management module using multi-channel parallel transmission technology. This technology uses the multi-tasking engine of the GPU to divide the data packets into blocks and distribute them to multiple CUDA streams, achieving parallel transmission and calculation, greatly improving the data transmission speed, reducing transmission delay, enabling the data management module to quickly obtain data, and optimizing the data interaction process between the BMC and the GPU.

[0106] Step 305: Control the substrate management controller to send a query instruction to the data management module according to the sampling interval.

[0107] Specifically in step 305, the baseboard management controller sends a query instruction to the data management module at a sampling interval. The sampling interval is set according to the variation law of the graphics processor operation parameters and the system's requirement for data real-time performance, ensuring that the obtained data can timely reflect the operation state of the graphics processor while taking into account the reasonable utilization of system resources. The instruction is transmitted through the established communication link and protocol, carrying clear data request information, such as the required data type, range, etc. After receiving the instruction, the data management module prepares to provide the corresponding data. This step is a key link for obtaining the graphics processor operation parameter data and provides data support for subsequent system monitoring and management.

[0108] Step 306, based on the historical data access sequence, construct a probability matrix of the operation parameter data.

[0109] Step 307, according to the target data block identifier whose transition probability value in the probability matrix is greater than the preset threshold, preload the operation parameter data corresponding to the target data block identifier.

[0110] Specifically in steps 306 to 307, a probability matrix of the operation parameter data is constructed based on the historical data access sequence. The system continuously records the access behavior of the baseboard management controller to the graphics processor operation parameter data, forming a historical data access sequence. Based on this, a specific algorithm is used to statistically analyze the sequence and frequency of access to different data blocks. By calculating the transition probability between each data block, a probability matrix is constructed. This matrix quantifies the correlation between data blocks and reflects the likelihood of accessing from one data block to another, providing a basis for predicting future data access trends and being an important foundation for realizing data preloading optimization.

[0111] Operate according to the probability matrix constructed in step 306. Set a preset threshold as the screening criterion to judge the transition probability values in the probability matrix. Determine the data block identifier with a transition probability value greater than the preset threshold as the target data block identifier. For these target data block identifiers, the system starts the preloading mechanism and reads and loads the corresponding operation parameter data from the storage location into the cache in advance. The preloading operation takes advantage of the high-speed read and write characteristics of the cache to reduce the waiting time during subsequent actual data access, accelerate the data acquisition speed, improve the efficiency of the baseboard management controller in obtaining the graphics processor operation parameter data, and optimize the data interaction process between the BMC and the GPU.

[0112] Furthermore, as an implementable manner of the embodiment of the present disclosure, if the operation parameter data is stored in the memory of the data management module, the operation parameter data is cached based on the page-locked memory algorithm.

[0113] Specifically, the data management module is responsible for storing the operation parameter data obtained from multiple graphics processors. Since these data need to be frequently read and used in subsequent processing, their storage and access efficiency directly affects the performance of the entire data interaction process. The page-locked memory algorithm uses the cudaHostAlloc function of CUDA to allocate page-locked memory (Pinned Memory). Through this operation, the storage of data in memory bypasses the paging mechanism of the operating system and is directly bound to the physical address. This binding method effectively reduces the number of memory copies. In the traditional memory management mode, the transfer of data between different memory areas often requires multiple copies, which not only consumes time but also occupies system resources and reduces the data access speed. The page-locked memory algorithm enables data to be directly stored and read in physical memory, avoiding the additional overhead caused by paging and copying.

[0114] During the data storage process, the data management module stores the operation parameter data in an orderly manner in the page-locked memory area according to certain storage rules. Different types of data, such as temperature, usage rate, power consumption, etc., are stored in pre-planned memory locations for quick retrieval and access. This caching mechanism based on the page-locked memory algorithm greatly improves the storage and reading efficiency of the operation parameter data, ensures that the baseboard management controller can respond quickly when obtaining data, and further optimizes the data interaction process between the BMC and the GPU, meeting the technical requirements of this patent technical solution to solve the problem of slow data interaction.

[0115] It should be noted that the embodiments of the present disclosure may include multiple steps. For the convenience of description, these steps are numbered, but these numbers are not intended to limit the execution time slots and execution orders between the steps; these steps can be implemented in any order, and the embodiments of the present disclosure do not make any limitations in this regard.

[0116] The embodiments of the present application also provide an electronic device, including a memory and a processor. The memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the method embodiments of the above data interaction.

[0117] The embodiments of the present application also provide a computer-readable storage medium, which stores a computer program. The computer program is configured to execute the steps in any one of the method embodiments of the above data interaction when running.

[0118] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media that can store computer programs, such as USB flash drives, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disks, magnetic disks, or optical discs.

[0119] The embodiments of the present application also provide a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the method embodiments of the above data interaction.

[0120] The embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the method embodiments of the above data interaction.

[0121] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0122] The above has introduced in detail a data interaction system and method, an electronic device, and a storage medium provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A system for data interaction, characterized in that, The system includes: a baseboard management controller, a data management module, and multiple graphics processors; Before the baseboard management controller obtains the operating status parameter data from the multiple graphics processors, the data management module obtains the operating parameter data from the multiple graphics processors; The multiple graphics processors split the operating parameter data to generate multiple data blocks, encapsulate the multiple data blocks into multiple parameter data packets according to a preset data format, and transmit the multiple parameter data packets to the data management module through multi-channel parallel transmission; After receiving the query instruction for obtaining the operating status data of the graphics processor sent by the baseboard management controller, the data management module extracts the operating parameter data cached in the local storage area of the data management module based on the dynamic prefetching algorithm. The dynamic prefetching algorithm constructs a state transition probability matrix according to the reading order of various sensor data, and extracts the operating parameter data according to the state transition probability matrix; and sends the cached operating parameter data back to the baseboard management controller.

2. The data interaction system according to claim 1, wherein The data management module includes multiple data processing units, and one of the data processing units is arranged on each of the graphics processors; each data processing unit includes a data processing subunit and a storage subunit; The data processing subunit sends a pre-query instruction to the graphics processor at a sampling interval to pre-obtain the operating parameter data at different time periods, and sends the operating parameter data to the storage subunit; The storage subunit receives and caches the operating parameter data; The data processing subunit responds to the query instruction sent by the baseboard management controller, obtains the operating parameter data from the storage subunit, and sends it to the baseboard management controller.

3. The data interaction system according to claim 2, wherein The baseboard management controller is further configured to send the query instruction to the multiple data processing units respectively at the sampling interval to obtain the operating parameter data of the multiple graphics processors.

4. The data interaction system according to claim 1, wherein The data management module includes at least one processor management unit disposed in the baseboard management controller; The processor management unit sends a pre-query instruction to the multiple graphics processors at a sampling interval to pre-obtain the operating parameter data at different time periods, and caches the operating parameter data of the multiple graphics processors into the memory of the data management module based on the page-locked memory algorithm; The processor management unit responds to the query instruction sent by the baseboard management controller, extracts the operating parameter data of the multiple graphics processors from the memory of the data management module, and sends it to the baseboard management controller.

5. The data interaction system according to claim 4, wherein The operating parameter data of the multiple graphics processors is stored in the storage partition corresponding to each graphics processor.

6. The data interaction system according to claim 4, wherein The baseboard management controller is further configured to send the query instruction to the at least one processor management unit at the sampling interval to obtain the operating parameter data of the multiple graphics processors.

7. The data interaction system according to claim 1, wherein The baseboard management controller is further configured to store the status data in a preset database; and in response to a data display instruction, perform data display on the status data in the preset database.

8. The data interaction system according to claim 1, wherein The multiple graphics processors divide the operation parameter data to obtain multiple data chunks, and perform compression processing on the multiple data chunks.

9. A method for data interaction, characterized in that, The method is applied to the data interaction system according to any one of claims 1-8, and the method includes: Controlling a data management module to send a pre-query instruction to multiple image processors at a sampling interval; Respectively performing segmentation processing on the operation parameter data of the multiple graphics processors, and transmitting the segmented operation parameter data to the data management module through multi-channel parallel transmission, and caching the operation parameter data; Controlling a baseboard management controller to send the query instruction to the data management module at the sampling interval, and extracting the operation parameter data corresponding to the multiple image processors cached in the local storage area of the data management module; Sending the operation parameter data corresponding to the multiple image processors to the baseboard management controller.

10. The method for data interaction according to claim 9, wherein The respectively performing segmentation processing on the operation parameter data of the multiple graphics processors, and transmitting the segmented operation parameter data to the data management module through multi-channel parallel transmission includes: Dividing the operation parameter data into multiple data chunks, and performing compression processing on the data chunks; Encapsulating the multiple compressed data chunks into multiple parameter data packets according to a preset data format; Transmitting the multiple parameter data packets to the data management module through multi-channel parallel transmission.

11. The method for data interaction according to claim 9, wherein The extracting the operation parameter data corresponding to the multiple image processors cached in the local storage area of the data management module includes: Constructing a probability matrix of the operation parameter data based on a historical data access sequence; Preloading the operation parameter data corresponding to the target data block identifier according to the target data block identifier with a transition probability value greater than a preset threshold in the probability matrix.

12. The method for data interaction according to claim 9, wherein The method further includes: If the operation parameter data is stored in a target memory, caching the operation parameter data based on a page-locked memory algorithm.

13. An electronic device, characterized in that, Including: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the data interaction method according to any one of claims 9-12.

14. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the data interaction method according to any one of claims 9-12.

15. A computer program product, characterized in that, Including a computer program that, when executed by a processor, implements the data interaction method according to any one of claims 9-12.

Citation Information

Patent Citations

  • Method, system and device for monitoring GPU (graphics processing unit) by BMC (baseboard management controller) as well as storage medium

    CN108268361A

  • Graphic processor monitoring method, system and device and electronic equipment

    CN115543746A