Data interaction system and method, electronic equipment and storage medium

By using dynamic prefetch algorithms and multi-channel parallel transmission technology in the data management module, the problem that the substrate management controller cannot monitor the operating status of the graphics processor in a timely and accurate manner is solved, efficient data interaction is achieved, and the stability and performance of the server system are improved.

CN120029856AActive Publication Date: 2025-05-23INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510490580.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-05-23
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

The substrate management controller cannot monitor the operating status of the graphics processor in a timely and accurate manner, especially when the number of GPUs increases, the total interaction time increases significantly, affecting the stability and performance of the server system.

Method used

By introducing dynamic prefetch algorithms and multi-channel parallel transmission technology into the data management module, the operation parameter data is obtained from multiple graphics processors in advance, and divided into multiple data blocks, encapsulated into parameter data packets according to the preset data format, and transmitted to the data management module in parallel through multiple channels.

Benefits of technology

It greatly saves the time when the substrate management controller directly interacts with the graphics processor to obtain data, improves the data acquisition efficiency, ensures that the substrate management controller can obtain the operating status parameters of the graphics processor in a timely manner, and improves the stability and performance of the server system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029856A_ABST
    Figure CN120029856A_ABST
Patent Text Reader

Abstract

The invention provides a data interaction system and method, an electronic device and a storage medium, a data management module obtains operation parameter data from a plurality of graphics processors before a substrate management controller obtains operation state parameter data from the plurality of graphics processors; the plurality of graphics processors divide the operation parameter data into a plurality of data blocks, encapsulate the plurality of data blocks into a plurality of parameter data packets, and transmit the plurality of parameter data packets to the data management module in parallel through multiple channels; after the data management module receives a query instruction which is sent by the substrate management controller and is used for obtaining the running state data of the graphics processor, the data management module extracts the running parameter data cached locally based on a dynamic prefetching algorithm and sends the cached running parameter data back to the substrate management controller. Compared with the prior art, the time for directly interacting with the graphics processor by the baseboard management controller to obtain the data can be saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of server technology, and in particular to a data interaction system and method, an electronic device and a storage medium. Background Art

[0002] With the development of mobile Internet, the demand for servers is rising. As an independent system on the server, the baseboard management controller (BMC) is responsible for monitoring hardware status, remote management, fault alarm and other functions. The graphics processing unit (GPU) is an important computing resource, and its operating status is related to system stability and performance. BMC needs to monitor key GPU parameters, such as temperature, usage, power consumption, video memory usage, fan speed and other dynamic information, as well as static information such as serial number and bus topology information.

[0003] Currently, BMC obtains GPU data through polling, that is, it sends commands to each GPU separately, and the GPU returns data after receiving and processing. However, it takes a long time for the GPU to receive, process, and return data. When BMC monitors all GPUs, it needs to interact with each GPU in turn. When the number of GPUs increases, the total interaction time increases significantly, making it impossible for BMC to monitor the GPU operating status in a timely and accurate manner, seriously affecting the stability and performance of the server system. Summary of the invention

[0004] The present disclosure provides a data interaction system and method, an electronic device and a storage medium, which are mainly intended to solve the problem that a baseboard management controller cannot monitor the GPU operation status in a timely and accurate manner.

[0005] According to a first aspect of the present disclosure, there is provided a data interaction system, comprising: a baseboard management controller, a data management module, and a plurality of graphics processors; The data management module obtains the operation parameter data from the multiple graphics processors before the baseboard management controller obtains the operation status parameter data from the multiple graphics processors; The multiple graphics processors divide the running parameter data into multiple data blocks, encapsulate the multiple data blocks into multiple parameter data packets according to a preset data format, and transmit the multiple parameter data packets to the data management module in parallel through multiple channels; After receiving the query instruction for obtaining the operation status data of the graphics processor sent by the baseboard management controller, the data management module extracts the operation parameter data cached locally based on the dynamic pre-fetch algorithm, and sends the cached operation parameter data back to the baseboard management controller.

[0006] Optionally, the data management module includes a plurality of data processing units, each graphics processor is provided with a data processing unit; the data processing unit includes a data processing sub-unit and a storage sub-unit; The data processing subunit sends a pre-query instruction to the graphics processor according to the sampling interval to pre-acquire the operation parameter data of different time periods, and sends the operation parameter data to the storage subunit; The storage subunit receives and caches the operation parameter data; The data processing subunit responds to the query instruction sent by the baseboard management controller, obtains the operation parameter data from the storage subunit, and sends the data to the baseboard management controller.

[0007] Optionally, the baseboard management controller is further configured to send query instructions to the multiple data processing units respectively according to sampling intervals to obtain operating parameter data of the multiple graphics processors.

[0008] Optionally, the data management module includes at least one processor management unit provided in the baseboard management controller; the processor management unit includes a processor management unit; The processor management unit sends a pre-query instruction to the multiple graphics processors according to the sampling interval to pre-acquire the operation parameter data of different time periods, and caches the operation parameter data of the multiple graphics processors into the memory of the data management module based on the page-locked memory algorithm; The processor management unit responds to the query instruction sent by the baseboard management controller, extracts the operation parameter data of the plurality of graphics processors from the memory of the data management module, and sends the data to the baseboard management controller.

[0009] Optionally, the operating parameter data of the multiple graphics processors are stored in storage partitions corresponding to the respective graphics processors.

[0010] Optionally, the baseboard management controller is further configured to send a query instruction to at least one processor management unit according to a sampling interval to obtain operating parameter data of multiple graphics processors.

[0011] Optionally, the baseboard management controller is further configured to store the status data in a preset database; and in response to a data display instruction, display the status data in the preset database.

[0012] Optionally, after dividing the operating parameter data, the multiple graphics processors obtain multiple data blocks, and perform compression processing on the multiple data blocks.

[0013] According to a second aspect of the present disclosure, a data interaction method is provided, comprising: Control the data management module to send a pre-query instruction to the plurality of image processors according to the sampling interval; Separately process the operating parameter data of the multiple graphics processors, transmit the segmented operating parameter data to the data management module in parallel through multiple channels, and cache the operating parameter data; Controlling the substrate management controller to send a query instruction to the data management module according to the sampling interval, and extracting the operating parameter data corresponding to the multiple image processors cached in the data management module; The operating parameter data corresponding to the plurality of image processors are sent to the baseboard management controller.

[0014] Optionally, the operating parameter data of the plurality of graphics processors are segmented and processed respectively, and the segmented operating parameter data are transmitted to the data management module in parallel through multiple channels, including: Divide the operating parameter data into multiple data blocks, and compress the data blocks; Encapsulating the compressed data into multiple parameter data packets according to a preset data format; Multiple parameter data packets are transmitted in parallel to the data management module through multiple channels.

[0015] Optionally, extracting the operating parameter data corresponding to the plurality of image processors cached in the data management module includes: Based on the historical data access sequence, a probability matrix of operating parameter data is constructed; According to the target data block identifier whose transition probability value in the probability matrix is ​​greater than a preset threshold, the operating parameter data corresponding to the target data block identifier is preloaded.

[0016] Optionally, the method further includes: If the operating parameter data is stored in the target memory, the operating parameter data is cached based on a page-locked memory algorithm.

[0017] According to a third aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the data interaction method described in the second aspect above.

[0018] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the data interaction method described in the second aspect.

[0019] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the data interaction method as described in the second aspect above.

[0020] The present disclosure provides a data interaction system and method, an electronic device and a storage medium. The data interaction system includes: a baseboard management controller, a data management module, and multiple graphics processors; the data management module obtains operation parameter data from multiple graphics processors before the baseboard management controller obtains operation status parameter data from the multiple graphics processors; the multiple graphics processors divide the operation parameter data into multiple data blocks, encapsulate the multiple data blocks into multiple parameter data packets according to a preset data format, and transmit the multiple parameter data packets to the data management module in parallel through multiple channels; after the data management module receives the query instruction sent by the baseboard management controller to obtain the operation status data of the graphics processor, it extracts the operation parameter data cached locally based on the dynamic prefetch algorithm, and sends the cached operation parameter data back to the baseboard management controller. Compared with the related art, the data management module in the present disclosure obtains the operation parameter data from the GPU in advance before the BMC obtains the operation status parameter data from the multiple GPUs. The GPU divides the operation parameter data into multiple data blocks, encapsulates them into parameter data packets, and transmits them to the data management module in parallel through multiple channels. This method greatly saves the time for the BMC to directly interact with the GPU to obtain data. In the traditional way, it takes a long time for BMC to poll each GPU to obtain data. However, this system allows BMC to quickly obtain processed data from the data management module and obtain GPU operating status parameters in a timely manner, thereby improving data acquisition efficiency. Multiple graphics processors divide data into data blocks, encapsulate them into data packets according to the preset data format, and transmit them to the data management module in parallel using multiple channels. Multi-channel parallel transmission makes full use of transmission channel resources. Compared with single-channel transmission, it can transmit more data in the same time, reduce data transmission delay, and improve the overall bandwidth and speed of data transmission, so that the data management module can obtain GPU operating parameter data faster and provide more timely data support for BMC. After BMC receives the query instruction sent by the data management module, the data management module extracts the locally cached operating parameter data based on the dynamic prefetch algorithm and sends it back to BMC. The dynamic prefetch algorithm predicts the data blocks that BMC may access and prefetches them into the cache in advance by analyzing the historical data access sequence. This means that the BMC does not need to wait for the data management module to read from the original storage when obtaining data, which reduces the time the BMC waits for data, improves the BMC's operating performance, and saves the BMC's memory space, avoiding the frequent data reading operations from occupying memory resources. During server operation, the traditional BMC monitoring GPU method has a blank period for data acquisition before the system starts. In this system, the data management module obtains the GPU operating parameter data in advance and starts collecting data before the BMC starts. When the BMC starts and sends a query command, the data management module can provide data immediately, reducing the blank period for collecting GPU-related parameters, ensuring the continuity and integrity of the BMC's monitoring of the GPU's operating status, and enabling the BMC to more comprehensively and accurately grasp the GPU's operating status.

[0021] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure. Figure 1 A schematic diagram of the structure of a data interaction system provided by an embodiment of the present disclosure; Figure 2 A structural diagram of another data interaction system provided by an embodiment of the present disclosure; Figure 3 A structural diagram of another data interaction system provided by an embodiment of the present disclosure; Figure 4 A flowchart of a data interaction method provided by an embodiment of the present disclosure; Figure 5 A flowchart of another data interaction method provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0023] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0024] The following describes the data interaction system and method, electronic device and storage medium according to the embodiments of the present disclosure with reference to the accompanying drawings.

[0025] Figure 1 This is a schematic diagram of the structure of a data interaction system provided by an embodiment of the present disclosure. Figure 1 As shown, the system includes: a baseboard management controller 11, a data management module 12, and a plurality of graphics processors 13; The data management module 12 obtains the operation parameter data from the plurality of graphics processors 13 before the baseboard management controller 11 obtains the operation status parameter data from the plurality of graphics processors 13 .

[0026] In the embodiment of the present disclosure, before the baseboard management controller 11 starts the operation of obtaining the operating status parameter data from the multiple graphics processors 13, the data management module 12 actively initiates the data interaction process with the multiple graphics processors 13. Its purpose is to pre-acquire various parameter data during the operation of the graphics processor 13. These data cover dynamic information such as the temperature, usage rate, power consumption, video memory usage, fan speed, etc. of the graphics processor 13, as well as static information such as serial number, i2c topology information, vendor ID, device ID, etc., which are of great significance for comprehensively and accurately monitoring the operating status of the graphics processor 13. The data management module 12 has the hardware and software conditions to realize the above functions. From the hardware point of view, it has an interface that can interact with the graphics processor 13 for data. These interfaces can be high-speed data transmission interfaces such as PCIE interfaces according to different implementation methods to ensure the stability and efficiency of data transmission. From the software level, the data management module 12 is installed with a management system with specific functions. The management system has a polling mechanism, through which the data management module 12 can send commands to obtain data to each graphics processor 13 in turn according to the set rules. The data management module is connected to the plurality of graphics processors through the first interaction interface, and is connected to the baseboard management controller through the second interaction interface. During the data acquisition process, the data management module 12 and the graphics processor 13 follow a specific data interaction protocol.

[0027] The multiple graphics processors 13 divide the operation parameter data into multiple data blocks, encapsulate the multiple data blocks into multiple parameter data packets according to a preset data format, and transmit the multiple parameter data packets to the data management module 12 in parallel through multiple channels.

[0028] In the embodiment of the present disclosure, the graphics processor 13 performs specific data processing and transmission operations in the data transmission link, that is, the operating parameter data is segmented and packaged, and transmitted to the data management module 12 in parallel through multiple channels. This series of operations can achieve efficient data interaction.

[0029] From the perspective of technical principles, the amount of operating parameter data generated by multiple graphics processors 13 during operation is huge and complex. In order to ensure that the data can be efficiently and stably transmitted to the data management module 12, the graphics processor 13 first divides the operating parameter data into multiple data blocks according to the established block algorithm. The block algorithm is carefully designed to divide the large data into fixed-size blocks, such as a specific value between 4KB-1MB. The setting of this range is based on a comprehensive consideration of memory management and data transmission efficiency. Transmitting data in block order effectively reduces memory fragmentation and random access overhead, and ensures the stability and continuity of data storage and transmission. After completing the data block division, the graphics processor 13 encapsulates multiple data blocks into multiple parameter data packets according to the preset data format. This preset data format is clearly defined in the system design stage, and its purpose is to unify the structure and standard of the data, so as to facilitate the identification, parsing and processing of data between different components. Through standardized data encapsulation, whether in the data transmission process or when the data management module 12 receives and processes data, the data error rate can be reduced and the accuracy and efficiency of data processing can be improved.

[0030] In terms of data transmission, multiple graphics processors 13 use multi-channel parallel transmission to transmit multiple parameter data packets to the data management module 12. Multi-channel parallel transmission takes advantage of the multi-tasking engine characteristics of the GPU itself, and divides the compressed data into blocks and distributes them to multiple CUDA streams. This method breaks the limitations of traditional single-channel transmission, allowing data to be transmitted simultaneously on multiple channels, significantly improving the overall bandwidth and speed of data transmission. Through parallel processing, the time for data to be transmitted from the graphics processor 13 to the data management module 12 is greatly shortened, allowing the data management module 12 to obtain the operating parameter data of the graphics processor 13 more timely, providing strong support for subsequent data processing and feedback to the baseboard management controller 11.

[0031] From the perspective of the integrity and logic of the system operation, the data segmentation, packaging and multi-channel parallel transmission operations performed by the multiple graphics processors 13 are closely coordinated with the functions of the data management module 12 and the baseboard management controller 11. After receiving these parameter data packets, the data management module 12 stores and processes them, and then feeds the processed data back to the baseboard management controller 11 according to the query instructions of the baseboard management controller 11, thereby realizing comprehensive monitoring of the operating status of the graphics processors 13. This series of operations together constitutes a complete and efficient data interaction process, which can optimize the data interaction process and improve system performance.

[0032] After receiving the query instruction for obtaining the operation status data of the graphics processor 13 sent by the baseboard management controller 11 , the data management module 12 extracts the operation parameter data cached locally based on the dynamic pre-fetch algorithm, and sends the cached operation parameter data back to the baseboard management controller 11 .

[0033] In the embodiment of the present disclosure, when the data management module 12 receives a query instruction for obtaining the operation status data of the graphics processor 13 sent by the baseboard management controller 11, it extracts the local cache data through a dynamic pre-fetch algorithm and transmits it back.

[0034] During the operation of the system, the data management module 12 continuously collects and stores the operating parameter data obtained from the graphics processor 13. These data are cached in the local storage area of ​​the data management module 12 in an orderly manner. The storage area has the characteristics of fast reading and writing, so that it can respond quickly when receiving a query instruction. When the baseboard management controller 11 issues a query instruction to obtain the operating status data of the graphics processor 13, the data management module 12 first parses and identifies the instruction to confirm the data type and range requested by the instruction. Based on the dynamic prefetch algorithm, the data management module 12 starts to perform data extraction operations. The dynamic prefetch algorithm realizes the prefetching of target data through in-depth analysis of historical data access sequences. For example, by counting the reading order of various types of sensor data by the baseboard management controller 11 in the past, such as the reading order of sensor A→B→C, a state transition probability matrix is ​​constructed. This matrix can accurately reflect the probability relationship between different data blocks being accessed. Based on this matrix, the data management module 12 can predict the data block that the baseboard management controller 11 may access next time.

[0035] Based on the prediction, the data management module 12 pre-fetches the data blocks with high probability of being accessed into the local cache in advance. When a query instruction is received, the relevant operating parameter data is extracted from the cache first. This method greatly reduces the delay of data reading, because the cached data can be quickly obtained without reading from the relatively slow conventional storage medium. If all the data that meets the query instruction requirements exists in the cache, the data management module 12 will directly encapsulate and organize these cached operating parameter data according to the established data transmission protocol.

[0036] In the data return stage, the data management module 12 sends the sorted operating parameter data back to the baseboard management controller 11 through the communication link connected to the baseboard management controller 11. The selection of the communication link and the formulation of the data transmission protocol have been clarified in the system design stage to ensure the stability, accuracy and efficiency of data transmission. During the transmission process, the data management module 12 will verify and correct the data to ensure that the data received by the baseboard management controller 11 is complete and correct. If there is no data in the cache that fully meets the requirements of the query instruction, the data management module 12 will quickly read the remaining part from the local storage, merge it with the cache data, and send it back to the baseboard management controller 11.

[0037] The present disclosure provides a data interaction system, in which the data management module 12 obtains the operation parameter data from the GPU in advance before the BMC obtains the operation status parameter data from the multiple GPUs. The GPU divides the operation parameter data into multiple data blocks, encapsulates them into parameter data packets, and transmits them to the data management module 12 in parallel through multiple channels. This method greatly saves the time for the BMC to interact directly with the GPU to obtain data. In the traditional way, it takes a long time for the BMC to poll each GPU to obtain data, but this system allows the BMC to quickly obtain the processed data from the data management module 12, obtain the GPU operation status parameters in time, and improve the data acquisition efficiency. Multiple graphics processors 13 divide the data into data blocks, encapsulate them into data packets according to the preset data format, and transmit them to the data management module 12 in parallel using multiple channels. Multi-channel parallel transmission makes full use of the transmission channel resources. Compared with single-channel transmission, it can transmit more data in the same time, reduce data transmission delay, and improve the overall bandwidth and speed of data transmission, so that the data management module 12 can obtain the GPU operation parameter data faster and provide more timely data support for the BMC. After the BMC receives the query instruction sent by the data management module 12, the data management module 12 extracts the locally cached operating parameter data based on the dynamic prefetch algorithm and sends it back to the BMC. The dynamic prefetch algorithm predicts the data blocks that the BMC may access and prefetches them into the cache in advance by analyzing the historical data access sequence. This allows the BMC to obtain data without waiting for the data management module 12 to read from the original storage, reducing the time the BMC waits for data, improving the operating performance of the BMC, while saving the memory space of the BMC and avoiding the occupation of memory resources by frequent data reading operations. During server operation, the traditional BMC monitoring GPU method has a data acquisition blank period before the system starts. In this system, the data management module 12 obtains the GPU operating parameter data in advance and starts collecting data before the BMC starts. When the BMC starts and sends a query instruction, the data management module 12 can provide data immediately, reducing the blank period for collecting GPU-related parameters, ensuring the continuity and integrity of the BMC's monitoring of the GPU operating status, and enabling the BMC to more comprehensively and accurately grasp the GPU's operating conditions.

[0038] Furthermore, in a possible implementation of this embodiment, as Figure 2As shown, the data management module 12 includes multiple data processing units 121, and each graphics processor 13 is provided with a data processing unit 121; the data processing unit 121 includes a data processing sub-unit 1211 and a storage sub-unit 1212; the data processing sub-unit 1211 sends a pre-query instruction to the graphics processor 13 according to the sampling interval to pre-acquire the operating parameter data of different time periods, and sends the operating parameter data to the storage sub-unit 1212; the storage sub-unit 1212 receives and caches the operating parameter data; the data processing sub-unit 1211 responds to the query instruction sent by the baseboard management controller 11, obtains the operating parameter data from the storage sub-unit 1212, and sends it to the baseboard management controller 11.

[0039] Specifically, in the embodiment of the present disclosure, the data management module 12 is composed of a plurality of data processing units 121. This distributed architecture design enables the system to have better parallel processing capabilities and scalability when processing data of a plurality of graphics processors 13. Each graphics processor 13 is provided with a corresponding data processing unit 121. This one-to-one correspondence ensures the pertinence and efficiency of data processing, and can accurately obtain the operating parameter data of each graphics processor 13. In the data acquisition stage, the data processing subunit 1211 sends a pre-query instruction to the corresponding graphics processor 13 according to a preset sampling interval. The sampling interval is determined comprehensively based on factors such as the system's requirements for data real-time performance and the characteristics of the graphics processor 13 operating parameter changes. By periodically sending query instructions, the data processing subunit 1211 can pre-acquire the graphics processor 13 operating parameter data of different time periods, which covers dynamic information such as temperature, usage rate, power consumption, video memory usage, fan speed, and static information such as GPU serial number, i2c topology information, vendor ID, and device ID. After obtaining the operating parameter data, the data processing subunit 1211 sends it to the storage subunit 1212. The storage subunit 1212 serves as a temporary storage space for data and has a high-speed cache function. It receives and caches the operating parameter data sent by the data processing subunit 1211, and uses a specific storage algorithm to store the data in order to ensure fast retrieval and reading of the data. When the baseboard management controller 11 sends a query instruction, the data processing subunit 1211 is responsible for responding to this instruction. It first parses the specific content of the query instruction to clarify the type, range and other requirements of the required data. Then, based on these requirements, the corresponding operating parameter data is obtained from the storage subunit 1212. Due to the efficient storage and management of data by the storage subunit 1212, the data processing subunit 1211 can quickly and accurately retrieve the required data. After obtaining the data, the data processing subunit 1211 encapsulates and formats the operating parameter data in accordance with the data transmission protocol agreed with the baseboard management controller 11 to ensure the integrity and accuracy of the data during the transmission process. After that, the processed data is sent to the baseboard management controller 11 through the established communication link in the system.

[0040] Furthermore, in a possible implementation of this embodiment, as Figure 2 As shown, the baseboard management controller 11 is also used to send query instructions to the multiple data processing units 121 respectively according to the sampling interval to obtain the operating parameter data of the multiple graphics processors 13.

[0041] Specifically, in the embodiment of the present disclosure, the baseboard management controller 11 is a core component for monitoring and managing the server hardware status. In order to effectively monitor the operating status of multiple graphics processors 13, it is necessary to obtain their operating parameter data. In order to ensure the timeliness and accuracy of the data acquisition, the baseboard management controller 11 follows the pre-set sampling interval and sends query instructions to multiple data processing units 121 respectively. The sampling interval is not set arbitrarily, but takes into account many factors. From the operating characteristics of the graphics processor 13, the frequency of change of its operating parameters such as temperature, usage rate and other dynamic information varies. For example, the temperature may fluctuate in real time with the load situation, and the usage rate will also vary greatly in different task stages. At the same time, the system's requirements for data real-time should not be ignored. If the sampling interval is too long, the acquired data may lag behind and fail to reflect the actual operating status of the graphics processor 13 in time; if the sampling interval is too short, it may cause excessive consumption of system resources and affect the overall performance. Therefore, through the study of the operating law of the graphics processor 13 and the evaluation of system performance requirements, a suitable sampling interval is determined.

[0042] When sending a query command, the baseboard management controller 11 uses the communication link established between it and the multiple data processing units 121 to transmit data. The communication link adopts a specific communication protocol based on the system design, such as a high-speed serial computer expansion bus (PCIE) or other adapted communication standards to ensure the stability and efficiency of command transmission. The communication protocol clarifies key elements such as the format of the command, the transmission method, and the data verification rules.

[0043] When the data processing unit 121 receives the query instruction sent by the baseboard management controller 11, the data processing subunit 1211 inside it will parse the instruction. The data processing subunit 1211 obtains the corresponding graphics processor 13 operating parameter data from the storage subunit 1212 according to the data type and range specified in the instruction. The storage subunit 1212 has cached the operating parameter data of different time periods in accordance with the instructions of the data processing subunit 1211 in the early stage. Because the storage subunit 1212 adopts an efficient storage management strategy, such as a specific data index and cache algorithm, the data processing subunit 1211 can quickly locate and read the required data.

[0044] After acquiring the data, the data processing subunit 1211 encapsulates and processes the data according to the format and transmission mode agreed with the baseboard management controller 11. Subsequently, the data processing subunit 1211 transmits the processed data back to the baseboard management controller 11 through the original communication link.

[0045] Furthermore, in a possible implementation of this embodiment, as Figure 3As shown, the data management module 12 includes at least one processor management unit 122 arranged in the baseboard management controller 11; the processor management unit 122 sends a pre-query instruction to the multiple graphics processors 13 according to the sampling interval to pre-acquire the operating parameter data of different time periods, and caches the operating parameter data of the multiple graphics processors 13 into the memory of the data management module 12 based on the page-locked memory algorithm; the processor management unit 122 responds to the query instruction sent by the baseboard management controller 11, extracts the operating parameter data of the multiple graphics processors 13 from the memory of the data management module 12, and sends it to the baseboard management controller 11.

[0046] Specifically, in the embodiments of the present disclosure, the data management module 12 includes at least one processor management unit 122 disposed in the baseboard management controller 11. This architecture design is intended to fully utilize the hardware resources of the baseboard management controller 11 to achieve deep integration of data management functions with the overall functions of the BMC. As the direct executor of data interaction, the processor management unit 122 undertakes the important responsibilities of data acquisition, caching and response. Compared with the data processing unit 121 corresponding to each GPU in the above-mentioned embodiments, in this embodiment, a processor management unit 122 is directly disposed in the baseboard management controller 11, and the processor management unit 122 manages all graphics processors and performs functions similar to those of the data processing unit 121.

[0047] The processor management unit 122 sends a pre-query instruction to multiple graphics processors 13 according to a predetermined sampling interval. After acquiring the data, the processor management unit 122 caches the operating parameter data of multiple graphics processors 13 into the memory of the data management module based on the page-locked memory algorithm. The page-locked memory algorithm allocates page-locked memory (Pinned Memory) through CUDA's cudaHostAlloc, bypassing the operating system's paging mechanism and directly binding the physical address. The application of this algorithm effectively reduces the number of memory copies and improves the speed of data storage and reading. Since the amount of GPU operating parameter data is large, frequent memory copy operations will not only consume a lot of time, but may also cause data transmission delays, affecting the system's real-time monitoring of the GPU operating status. The page-locked memory algorithm directly stores the data in the physical memory and establishes a stable address mapping, so that the processor management unit 122 can quickly store and access the data, greatly improving the data processing efficiency.

[0048] When the baseboard management controller 11 sends a query instruction, the processor management unit 122 is responsible for responding. The processor management unit 122 first parses the query instruction to clarify the data type, range and other relevant parameters required by the instruction. Then, based on the parsing result, the corresponding operating parameter data of multiple graphics processors 13 are extracted from its memory. Since the page-locked memory algorithm was used for data caching, the processor management unit 122 can quickly locate and read the required data, avoiding the time overhead caused by memory paging and data search. After extracting the data, the processor management unit 122 encapsulates and formats the data according to the data transmission protocol pre-agreed with the baseboard management controller 11. This may include operations such as adding data verification information and packaging according to a specific data format to ensure the integrity and accuracy of the data during transmission. Finally, the processor management unit 122 sends the processed data to the baseboard management controller 11 through the communication link inside the system, so that the BMC can obtain the operating parameter data of the GPU in a timely manner and realize effective monitoring of the GPU operating status.

[0049] Furthermore, in a possible implementation of this embodiment, as Figure 3 As shown, the operating parameter data of the multiple graphics processors 13 are stored in the storage partitions corresponding to the respective graphics processors 13 .

[0050] Specifically, in the embodiments of the present disclosure, during the operation of the server, multiple graphics processors 13 each generate a large amount of operating parameter data, which is essential for monitoring their own operating status and ensuring the stable operation of the entire server system. In order to achieve effective management and rapid access to these data, each graphics processor 13 is provided with a corresponding storage partition. This correspondence is established based on specific rules of system design. From a hardware perspective, a storage partition can physically be part of an internal storage module of a graphics processor 13, or an external storage device space closely connected thereto. Logically, each storage partition is associated with a corresponding graphics processor 13 through a unique identifier, ensuring that data storage and reading are highly accurate and targeted.

[0051] Furthermore, in a possible implementation of this embodiment, the baseboard management controller 11 is further configured to send a query instruction to at least one processor management unit 122 according to a sampling interval to obtain operating parameter data of the plurality of graphics processors 13 .

[0052] Specifically, in the embodiment of the present disclosure, the baseboard management controller 11 sends a query instruction to at least one processor management unit 122 according to a predetermined sampling interval. The sampling interval is set according to the characteristics of the change in the operating parameters of the graphics processor 13 and the real-time requirements of the system data. After receiving the instruction, the processor management unit 122 extracts the operating parameter data corresponding to the multiple graphics processors 13 from its own memory cache. These data were previously obtained by the processor management unit 122 according to the sampling interval and cached using the page-locked memory algorithm. After the processor management unit 122 obtains the data, it encapsulates and processes it in accordance with the protocol agreed upon with the baseboard management controller 11, and then transmits it back to the baseboard management controller 11 through the system communication link, helping it to achieve effective monitoring of the operating status of multiple graphics processors 13, optimize the data interaction process between the BMC and the GPU, and improve system performance.

[0053] Furthermore, in a possible implementation of this embodiment, the baseboard management controller 11 is also used to store the status data in a preset database; and in response to a data display instruction, display the status data in the preset database.

[0054] Specifically, in the embodiment of the present disclosure, the baseboard management controller 11 stores the running status data of the graphics processor 13 obtained from the data management module 12 into a preset database according to the established data storage rules. The preset database can be a specific storage area inside the baseboard management controller 11, or a designated database space in an external storage device connected thereto. When storing, the status data is classified and indexed according to information such as data type and timestamp for subsequent quick query and call. When the baseboard management controller 11 receives the data display instruction, it will parse the instruction to clarify the data range, format and other requirements to be displayed. Subsequently, the corresponding status data is retrieved from the preset database according to the instruction requirements. During the retrieval process, the query optimization mechanism of the database, such as index query, is used to improve the data acquisition efficiency. After obtaining the data, the baseboard management controller 11 processes the status data according to the preset data display format. If it needs to be displayed through a web page, the data will be converted into a format suitable for web page display, such as HTML or JSON; if it is output through the redfish interface or ipmi command, the data is encapsulated according to the corresponding interface or command specification. Finally, the processed data is presented to the user or other related systems in an appropriate manner, so as to realize an intuitive display of the operating status of the graphics processor 13, so that the user can timely understand the operating status of the GPU and provide support for the management and maintenance of the server system.

[0055] Furthermore, in a possible implementation of this embodiment, the multiple graphics processors 13 obtain multiple data blocks after dividing the operating parameter data, and perform compression processing on the multiple data blocks.

[0056] Specifically, in the embodiments of the present disclosure, the amount of operating parameter data generated by multiple graphics processors 13 is large. In order to facilitate transmission and processing, it will be segmented to divide the large data into fixed-size data blocks, such as 4KB-1MB. This block method can reduce memory fragmentation and random access overhead and improve the stability of data processing. After the segmentation is completed, the graphics processor 13 uses a specific compression algorithm, such as the LZ4 algorithm, to compress the multiple data blocks. The compression process is intended to reduce the amount of data transmission and reduce the bandwidth and time required for transmission. Through compression, the originally larger data blocks are converted into a more compact format, which improves the efficiency of data transmission while maintaining the key information of the data. In the subsequent transmission process, the compressed data is encapsulated into a specific format (such as GPU-BIN format) and sent to the data management module 12 in a multi-stream parallel transmission mode through a binary protocol (such as Cap'n Proto). This processing method makes the transmission of data within the server system more efficient, which helps to solve the problem of slow data interaction between BMC and GPU.

[0057] Figure 4 A flowchart of a data interaction method provided by an embodiment of the present disclosure.

[0058] like Figure 4 As shown, the method comprises the following steps: Step 201: Control the data management module to send pre-query instructions to multiple image processors according to sampling intervals.

[0059] In the embodiment of the present disclosure, the data management module assumes the important responsibility of obtaining the operating parameter data of the graphics processor. Among them, the setting of the sampling interval is based on the comprehensive consideration of multiple factors such as the operating characteristics of the graphics processor and the system's requirements for data real-time performance. For example, the frequency of change of the operating parameters of the graphics processor in different application scenarios is different. In complex graphics rendering tasks, its temperature, usage rate and other parameters change rapidly, and a shorter sampling interval is required to ensure the timeliness of data acquisition; in relatively stable operating scenarios, the sampling interval can be appropriately extended to balance the real-time performance of data acquisition and system resource consumption. When the system is running, according to the established control logic, the data management module follows the set sampling interval and sends pre-query instructions to multiple graphics processors in an orderly manner. These instructions are transmitted through specific communication interfaces and protocols, such as PCIE interfaces combined with corresponding high-speed communication protocols to ensure the stability and accuracy of instruction transmission. The instructions clearly contain key information such as the type and range of the required data, so that the graphics processor can accurately know the data content that needs to be fed back. After receiving the instruction, the graphics processor organizes the operating parameter data according to the instruction requirements and prepares for subsequent transmission to the data management module. This process lays the foundation for efficient BMC and GPU data interaction, which is in line with the original design intention of the patented technical solution to solve the problem of slow data interaction.

[0060] Step 202 , segmenting the operating parameter data of the plurality of graphics processors respectively, transmitting the segmented operating parameter data to the data management module in parallel through multiple channels, and caching the operating parameter data.

[0061] In an embodiment of the present disclosure, after receiving the pre-query instruction, multiple graphics processors will segment the operating parameter data. The segmentation is based on a specific algorithm to divide the big data into fixed-size blocks, such as 4KB-1MB, which can reduce memory fragmentation and random access overhead and improve data processing efficiency. After the segmentation is completed, the data is sent to the data management module through multi-channel parallel transmission technology. Multi-channel parallel transmission uses the multi-tasking engine of the GPU to divide the data into blocks and distribute it to multiple CUDA streams, and transmit and calculate at the same time, which greatly improves the data transmission speed.

[0062] After receiving the data, the data management module will cache the data. Using the page-locked memory algorithm, the page-locked memory is allocated through CUDA's cudaHostAlloc, bypassing the operating system paging mechanism, directly binding the physical address, reducing the number of memory copies, achieving fast storage, and facilitating the subsequent baseboard management controller to obtain data, optimizing the data interaction process between BMC and GPU, and meeting the requirements of patented technology.

[0063] Step 203: Control the substrate management controller to send a query instruction to the data management module according to the sampling interval, and extract the operating parameter data corresponding to the multiple image processors cached in the data management module.

[0064] In an embodiment of the present disclosure, a baseboard management controller (BMC) sends a query instruction to a data management module according to a preset sampling interval. The setting of the sampling interval comprehensively considers the frequency of change of the operating parameters of the graphics processor and the system's demand for real-time data, ensuring that the data acquired by the BMC can both reflect the operating status of the graphics processor in a timely manner and not consume system resources excessively. After receiving the query instruction, the data management module parses the instruction to clarify the specific requirements of the data required by the BMC. Then, according to the instruction requirements, the operating parameter data corresponding to multiple graphics processors are extracted from its cache. The data management module adopts an efficient cache management strategy, such as an index-based data storage method, which can quickly locate and extract the data required by the BMC and improve data acquisition efficiency. In the data extraction process, the data management module may also adopt a dynamic pre-fetching algorithm, such as a pre-fetching model based on a Markov chain. By statistically analyzing the historical data access sequence to construct a state transition probability matrix, the data blocks that the BMC may access are predicted, and the data blocks are pre-fetched to the cache in advance, further shortening the data acquisition time. Finally, the data management module sends the extracted data to the BMC according to the established data transmission protocol, so that the BMC can obtain the operating parameter data of the graphics processor and provide data support for the server system to monitor and manage the graphics processor, which meets the requirements of this patented technical solution to optimize the data interaction between BMC and GPU.

[0065] Step 204: Send the operating parameter data corresponding to the plurality of image processors to the baseboard management controller.

[0066] In an embodiment of the present disclosure, after receiving the query instruction sent by the baseboard management controller and completing the data extraction, the data management module sends the operating parameter data corresponding to the multiple graphics processors to the baseboard management controller according to the established data transmission protocol. The data transmission protocol clearly stipulates key elements such as the packaging format, transmission method and verification rules of the data. In terms of the packaging format, the data management module will package the extracted operating parameter data in a specific format, such as arranging different types of data (such as temperature, usage rate, etc.) in a prescribed order and structure to ensure the integrity and accuracy of the data during transmission. In terms of transmission mode, a high-performance binary data serialization protocol such as Cap'n Proto is usually adopted, and combined with a multi-stream parallel transmission mode, the multi-tasking engine of the GPU is used to divide the data into blocks and distribute it to multiple CUDA streams for parallel transmission to improve the speed and efficiency of data transmission and reduce transmission delays. At the same time, in order to ensure the accuracy of the data, the data management module will add verification information to the transmitted data according to the verification rules, such as using a cyclic redundancy check (CRC) algorithm to generate a verification code and attach it to the data. After receiving the data, the baseboard management controller will verify the data according to the same verification rules. If the verification passes, the data will be received; if the verification fails, the data management module will be required to retransmit the data.

[0067] The present disclosure provides a method for data interaction. In the present disclosure, before the BMC obtains the operating status parameter data from multiple GPUs, the data management module obtains the operating parameter data from the GPU in advance. The GPU divides the operating parameter data into multiple data blocks, encapsulates them into parameter data packets, and transmits them to the data management module through multiple channels in parallel. This method greatly saves the time for the BMC to directly interact with the GPU to obtain data. In the traditional method, it takes a long time for the BMC to poll each GPU to obtain data, but this system allows the BMC to quickly obtain the processed data from the data management module and obtain the GPU operating status parameters in time, thereby improving the data acquisition efficiency. Multiple graphics processors divide the data into data blocks, encapsulate them into data packets according to a preset data format, and transmit them to the data management module in parallel using multiple channels. Multi-channel parallel transmission makes full use of the transmission channel resources. Compared with single-channel transmission, more data can be transmitted in the same time, which reduces the data transmission delay, improves the overall bandwidth and speed of data transmission, and enables the data management module to obtain the GPU operating parameter data faster, providing more timely data support for the BMC. After the BMC receives the query instruction sent by the data management module, the data management module extracts the locally cached operating parameter data based on the dynamic prefetching algorithm and sends it back to the BMC. The dynamic prefetch algorithm predicts the data blocks that the BMC may access by analyzing the historical data access sequence and prefetches them into the cache in advance. This allows the BMC to obtain data without waiting for the data management module to read from the original storage, reducing the time the BMC waits for data, improving the BMC's operating performance, and saving the BMC's memory space, avoiding the occupation of memory resources by frequent data reading operations. During server operation, the traditional BMC monitoring GPU method has a data acquisition blank period before the system starts. In this system, the data management module obtains the GPU operating parameter data in advance and starts collecting data before the BMC starts. When the BMC starts and sends a query command, the data management module can provide data immediately, reducing the blank period for collecting GPU-related parameters, ensuring the continuity and integrity of the BMC's monitoring of the GPU operating status, and enabling the BMC to more comprehensively and accurately grasp the GPU's operating status.

[0068] In order to clearly illustrate the embodiment of the present disclosure, this embodiment provides a flowchart of another data interaction method.

[0069] like Figure 5 As shown, the method comprises the following steps: Step 301: Control the data management module to send pre-query instructions to multiple image processors according to sampling intervals.

[0070] Specifically, in step 301, the data management module sends a pre-query instruction to multiple graphics processors according to a pre-set sampling interval. The setting of the sampling interval comprehensively considers the characteristics of the change of the operating parameters of the graphics processor and the real-time requirements of the system data to ensure the timeliness of data acquisition and the rational use of system resources. The instruction is transmitted through a specific communication interface and protocol to accurately convey key information such as the type and range of the required data, laying the foundation for subsequent data collection.

[0071] Step 302: divide the operating parameter data into a plurality of data blocks, and compress the data blocks.

[0072] Specifically, in step 302, after receiving the query instruction, the graphics processor processes the operating parameter data. A specific block algorithm is used to divide the data into fixed-size data blocks, generally in the range of 4KB-1MB, to reduce memory fragmentation and random access overhead. Subsequently, the data blocks are compressed using an algorithm such as LZ4 to reduce the data volume and improve subsequent transmission efficiency.

[0073] Step 303: encapsulate the compressed data into multiple parameter data packets in a preset data format.

[0074] Specifically, in step 303, the compressed data are divided into multiple parameter data packets according to the data format preset by the system. This preset format clarifies the data organization method, header and footer structure, etc., to ensure the integrity and identifiability of the data during transmission, and facilitate the data management module to receive and process. For example, a high-performance binary data serialization protocol such as Cap'n Proto is used.

[0075] Step 304: Transmit the multiple parameter data packets in parallel to the data management module via multiple channels.

[0076] Specifically, in step 304, multiple parameter data packets are transmitted to the data management module using multi-channel parallel transmission technology. This technology uses the multi-tasking engine of the GPU to divide the data packets into blocks and distribute them to multiple CUDA streams, achieving parallel transmission and calculation, greatly improving the data transmission speed, reducing transmission delay, enabling the data management module to quickly obtain data, and optimizing the data interaction process between the BMC and the GPU.

[0077] Step 305: Control the substrate management controller to send a query instruction to the data management module according to the sampling interval.

[0078] Specifically, in step 305, the baseboard management controller sends a query instruction to the data management module according to the sampling interval. The sampling interval is set according to the change law of the graphics processor operating parameters and the system's requirements for data real-time performance, ensuring that the acquired data can timely reflect the graphics processor operating status while taking into account the rational use of system resources. The instruction is transmitted through the established communication link and protocol, carrying clear data request information, such as the required data type and range. After receiving the instruction, the data management module is ready to provide the corresponding data. This step is a key step in obtaining the graphics processor operating parameter data, providing data support for subsequent system monitoring and management.

[0079] Step 306: construct a probability matrix of the operating parameter data based on the historical data access sequence.

[0080] Step 307 , pre-load the operation parameter data corresponding to the target data block identifier according to the target data block identifier whose transition probability value in the probability matrix is ​​greater than a preset threshold.

[0081] Specifically, in steps 306 to 307, a probability matrix of operating parameter data is constructed based on the historical data access sequence. The system continuously records the baseboard management controller's access behavior to the graphics processor's operating parameter data to form a historical data access sequence. Based on this, a specific algorithm is used to count the order and frequency of access to different data blocks. By calculating the transition probability between each data block, a probability matrix is ​​constructed. This matrix quantifies the correlation between data blocks, reflects the probability of accessing and transferring from one data block to another, provides a basis for predicting future data access trends, and is an important basis for achieving data preloading optimization.

[0082] The operation is performed according to the probability matrix constructed in step 306. A preset threshold is set as a screening criterion to judge the transition probability value in the probability matrix. The data block identifier whose transition probability value is greater than the preset threshold is determined as the target data block identifier. For these target data block identifiers, the system starts the preloading mechanism to read and load the corresponding operating parameter data from the storage location into the cache in advance. The preloading operation utilizes the high-speed read and write characteristics of the cache to reduce the waiting time during subsequent actual data access, speed up data acquisition, improve the efficiency of the baseboard management controller in obtaining the graphics processor operating parameter data, and optimize the data interaction process between the BMC and the GPU.

[0083] Further, as an implementable manner of the embodiment of the present disclosure, if the operating parameter data is stored in the memory of the data management module, the operating parameter data is cached based on a page-locked memory algorithm.

[0084] Specifically, the data management module is responsible for storing the operating parameter data obtained from multiple graphics processors. Since this data needs to be frequently read and used in subsequent processing, its storage and access efficiency directly affect the performance of the entire data interaction process. The page-locked memory algorithm uses CUDA's cudaHostAlloc function to allocate page-locked memory (PinnedMemory). Through this operation, the storage of data in memory bypasses the paging mechanism of the operating system and is directly bound to the physical address. This binding method effectively reduces the number of memory copies. In the traditional memory management mode, the transmission of data between different memory areas often requires multiple copies, which not only consumes time, but also occupies system resources and reduces data access speed. The page-locked memory algorithm allows data to be stored and read directly in physical memory, avoiding the additional overhead caused by paging and copying.

[0085] During the data storage process, the data management module stores the operating parameter data in an orderly manner in the page-locked memory area according to certain storage rules. Different types of data, such as temperature, usage rate, power consumption, etc., are stored in pre-planned memory locations for quick retrieval and access. This cache mechanism based on the page-locked memory algorithm greatly improves the storage and reading efficiency of operating parameter data, ensures that the baseboard management controller can respond quickly when acquiring data, and further optimizes the data interaction process between the BMC and the GPU, which meets the technical requirements of this patented technical solution to solve the problem of slow data interaction.

[0086] It should be noted that the embodiments of the present disclosure may include multiple steps. For the convenience of description, these steps are numbered, but these numbers do not limit the execution time slots or execution order between the steps; these steps can be implemented in any order, and the embodiments of the present disclosure do not limit this.

[0087] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above data interaction method embodiments.

[0088] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned data interaction method embodiments when running.

[0089] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0090] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned data interaction method embodiments are implemented.

[0091] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned data interaction method embodiments are implemented.

[0092] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0093] The above is a detailed introduction to a data interaction system and method, electronic device and storage medium provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of ​​the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A data interaction system, characterized in that: The system includes: a baseboard management controller, a data management module, and a plurality of graphics processors; The data management module obtains the operating parameter data from the multiple graphics processors before the baseboard management controller obtains the operating status parameter data from the multiple graphics processors; The multiple graphics processors divide the operation parameter data into multiple data blocks, encapsulate the multiple data blocks into multiple parameter data packets according to a preset data format, and transmit the multiple parameter data packets to the data management module in parallel through multiple channels; After receiving the query instruction for obtaining the operation status data of the graphics processor sent by the baseboard management controller, the data management module extracts the operation parameter data cached locally based on the dynamic pre-fetch algorithm, and sends the cached operation parameter data back to the baseboard management controller.

2. The data interaction system according to claim 1, characterized in that: The data management module includes a plurality of data processing units, each of which is provided with one data processing unit; the data processing unit includes a data processing sub-unit and a storage sub-unit; The data processing subunit sends a pre-query instruction to the graphics processor according to a sampling interval to pre-acquire the operation parameter data of different time periods, and sends the operation parameter data to the storage subunit; The storage subunit receives and caches the operating parameter data; The data processing subunit responds to the query instruction sent by the baseboard management controller, obtains the operating parameter data from the storage subunit, and sends the data to the baseboard management controller.

3. The data interaction system according to claim 2, characterized in that: The baseboard management controller is further configured to send the query instructions to the multiple data processing units respectively according to the sampling interval to obtain the operating parameter data of the multiple graphics processors.

4. The data interaction system according to claim 1, characterized in that: The data management module includes at least one processor management unit arranged in the baseboard management controller; The processor management unit sends a pre-query instruction to the multiple graphics processors according to a sampling interval to pre-acquire the operating parameter data of different time periods, and caches the operating parameter data of the multiple graphics processors into the memory of the data management module based on a page-locked memory algorithm; The processor management unit responds to the query instruction sent by the baseboard management controller, extracts the operating parameter data of the multiple graphics processors from the memory of the data management module, and sends the data to the baseboard management controller.

5. The data interaction system according to claim 4, characterized in that: The operating parameter data of the multiple graphics processors are stored in storage partitions corresponding to the respective graphics processors.

6. The data interaction system according to claim 4, characterized in that: The baseboard management controller is further configured to send the query instruction to the at least one processor management unit according to the sampling interval to obtain the operating parameter data of the multiple graphics processors.

7. The data interaction system according to claim 1, characterized in that: The baseboard management controller is further used to store the status data in a preset database; and in response to a data display instruction, display the status data in the preset database.

8. The data interaction system according to claim 1, characterized in that: The multiple graphics processors obtain multiple data blocks after dividing the operating parameter data, and perform compression processing on the multiple data blocks.

9. A data interaction method, characterized in that: The method is applied to a data interaction system according to any one of claims 1 to 8, and the method comprises: Control the data management module to send a pre-query instruction to the plurality of image processors according to the sampling interval; Separately process the operating parameter data of the multiple graphics processors, transmit the segmented operating parameter data to the data management module in parallel through multiple channels, and cache the operating parameter data; Controlling the substrate management controller to send the query instruction to the data management module according to the sampling interval, and extracting the operating parameter data corresponding to the multiple image processors cached in the data management module; The operating parameter data corresponding to the plurality of image processors are sent to the baseboard management controller.

10. The data interaction method according to claim 9, characterized in that: The step of segmenting the operating parameter data of the plurality of graphics processors respectively and transmitting the segmented operating parameter data to the data management module in parallel through multiple channels includes: Dividing the operating parameter data into a plurality of data blocks, and compressing the data blocks; Encapsulating the compressed data into multiple parameter data packets according to a preset data format; The multiple parameter data packets are transmitted to the data management module in parallel through multiple channels.

11. The data interaction method according to claim 9, characterized in that: The extracting the operating parameter data corresponding to the plurality of image processors cached in the data management module includes: constructing a probability matrix of the operating parameter data based on the historical data access sequence; According to the target data block identifier whose transition probability value in the probability matrix is ​​greater than a preset threshold, the operating parameter data corresponding to the target data block identifier is preloaded.

12. The data interaction method according to claim 9, characterized in that: The method further comprises: If the operating parameter data is stored in the target memory, the operating parameter data is cached based on a page-locked memory algorithm.

13. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the data interaction method according to any one of claims 9 to 12.

14. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the data interaction method according to any one of claims 9-12.

15. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the method for data interaction according to any one of claims 9 to 12.

Citation Information

Patent Citations

  • Method, system and device for monitoring GPU (graphics processing unit) by BMC (baseboard management controller) as well as storage medium

    CN108268361A

  • Graphic processor monitoring method, system and device and electronic equipment

    CN115543746A

  • Information unvarnished transmission method, device and equipment of monitoring system and storage medium

    CN115878414A

  • Operation state monitoring method and device, electronic equipment and storage medium

    CN117493120A

  • Composition of improving ectoine production and Method for improved ectoine production using the same

    KR1020230013959A

Cited By

  • Monitoring method and device of graphics processor, baseboard management controller and medium

    CN120687327A

  • Information processing method and electronic equipment

    CN120821634A

  • Information processing method and electronic device

    CN120821634B

  • Data expansion acceleration processing system

    CN122367717A