A GPU server data security management method and system
Patent Information
- Application Number
- CN202611035223.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-13
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]传统的GPU服务器数据安全管理方法虽然能够通过访问控制和恶意代码防护降低未授权访问及恶意入侵风险,但是,在不同计算任务连续复用显存页面时,前一计算任务写入或使用后遗留的数据仍可能保存在已经回收的显存页面中;当显存页面的残留数据清除状态未被准确区分,且待清除显存页面缺少与任务安全等级和等待时间相适应的处理顺序时,尚未完成残留数据清除的显存页面可能被后续计算任务重新使用;同时,当显存页面清除过程未结合当前计算任务能够承受的等待时间进行控制时,还存在清除过程与显存资源分配时限不协调的问题
本发明根据显存页面的占用解除状态和遗留数据清除完成状态,对待清除显存页面与已经完成清除的显存页面进行区分,能够避免清除状态不明的显存页面直接进入再次分配过程;随后,结合前一计算任务安全等级、显存页面回收队列等待时长和显存页面历史选取次数确定残留数据清除执行顺序,使需要优先处理的显存页面按照对应顺序进入紧急清除队列,从而使显存页面的安全属性与清除先后次序保持协调;在此基础上,根据当前计算任务的页面申请数量、紧急清除队列中尚未完成清除的显存页面数量以及平均残留数据清除耗时控制显存页面的选取时机,使残留数据清除过程能够适应当前计算任务的最大可容忍等待时间;此外,在显存页面完成写零处理或覆盖清除处理并形成页面清除完成记录后,再分配给当前计算任务,既能够阻断前一计算任务遗留数据被后续计算任务读取,也能够限制未完成清除的显存页面重新参与分配,并兼顾显存资源分配的时效性。
Smart Images

Figure CN122839447A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data management technology, and in particular to a method and system for data security management of GPU servers. Background Technology
[0002] GPU server data security management methods refer to the protection processes implemented in computing centers equipped with graphics processors during the execution of computing and storage tasks, which are used to prevent unauthorized access and the risk of malicious code intrusion.
[0003] While traditional GPU server data security management methods can reduce the risk of unauthorized access and malicious intrusion through access control and malware protection, there are still some drawbacks. When different computing tasks reuse GPU memory pages consecutively, data left over from previous computing tasks may still be stored in reclaimed GPU memory pages. When the residual data clearing status of GPU memory pages is not accurately distinguished, and the processing order of GPU memory pages to be cleared lacks an appropriate level of task security and waiting time, GPU memory pages that have not yet completed residual data clearing may be reused by subsequent computing tasks. At the same time, when the GPU memory page clearing process is not controlled in conjunction with the waiting time that the current computing task can tolerate, there is also the problem of inconsistency between the clearing process and the GPU memory resource allocation time limit. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing a GPU server data security management method and system.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: a GPU server data security management method, comprising the following steps: Obtain the current computing task, the previous computing task corresponding to the current computing task, and the current memory page recycling queue. Collect memory page information for each memory page of the previous computing task contained in the current memory page recycling queue. Each video memory page is divided according to the video memory page information to obtain a set of pages to be cleared and a first set of cleared pages; Determine the execution order for clearing residual data in each video memory page of the set of pages to be cleared; Based on the residual data removal execution order, separate the video memory pages from the set of pages to be removed, and push the separated video memory pages into the preset emergency removal queue; Clear residual data left over from the previous computation task in the emergency clearing queue of video memory pages, obtain all video memory pages after clearing residual data, obtain the second cleared page set, and integrate the first cleared page set and the second cleared page set to obtain the allocatable page set. Assign the set of allocatable pages to the current computing task to obtain the data management results.
[0006] As a further aspect of the present invention, the step of obtaining the current computing task, the previous computing task corresponding to the current computing task, and the current video memory page reclamation queue, and collecting video memory page information of each video memory page of the previous computing task contained in the current video memory page reclamation queue includes: Get the GPU server's current memory resource allocation request, and take the computing task that sent the memory resource allocation request as the current computing task. At the same time, get the previous computing task corresponding to the current computing task and the current memory page reclamation queue corresponding to the current computing task. Collect all video memory pages in the current video memory page reclamation queue, as well as the video memory page information for each video memory page. The video memory page information includes the previous computing task's release flag, the previous computing task's legacy data clearing completion flag, the previous computing task's security level, the video memory page reclamation queue waiting time, and the video memory page's historical selection count. Among them, the previous computing task's release flag indicates that the video memory page is no longer occupied by the previous computing task of the current computing task and has been returned to the current video memory page reclamation queue. The previous computing task's legacy data clearing completion flag indicates that the data left in the video memory page after being written or used by the previous computing task has been written to zero or overwritten and cleared.
[0007] As a further aspect of the present invention, the step of dividing each video memory page according to the video memory page information to obtain a set of pages to be cleared and a first set of cleared pages includes: Referring to the aforementioned video memory page information, each video memory page that has a previous computing task's unoccupation flag but does not have a previous computing task's data clearing completion flag is grouped into a set of pages to be cleared. Referring to the aforementioned video memory page information, each video memory page with an identifier indicating that the data left over from the previous computing task has been cleared is grouped into a first set of cleared pages.
[0008] As a further aspect of the present invention, determining the execution order of residual data removal for each video memory page in the set of pages to be removed includes: Obtain the security level of the previous computing task, the waiting time of the video memory page recycling queue, and the number of times the video memory page has been selected in the history of each video memory page in the set of pages to be cleared; Normalization is performed on the security level of the previous computing task, the waiting time of the video memory page reclamation queue, and the number of times the video memory page was selected in history, respectively, to generate normalized security level of the previous computing task, normalized waiting time of the video memory page reclamation queue, and normalized number of times the video memory page was selected in history. Assign weights to the normalized security level of the previous computation task, the normalized waiting time of the memory page reclamation queue, and the normalized historical selection count of the memory page. Calculate the weighted sum of the normalized security level of the previous computation task, the normalized waiting time of the memory page reclamation queue, and the normalized historical selection count of the memory page to obtain the residual data removal priority score for each memory page. Determine the residual data removal execution order for each memory page in the set of pages to be removed based on the residual data removal priority score.
[0009] As a further aspect of the present invention, the step of separating video memory pages from the set of pages to be cleared according to the residual data clearing execution order, and pushing the separated video memory pages into a preset emergency clearing queue includes: Get the memory resource allocation request sent by the current computing task, and the number of pages requested in the memory resource allocation request, as the clearing batch for the current task; Obtain the emergency cleanup queue pre-set by the GPU server for the current computing task. Read the page cleanup records of the video memory pages corresponding to the current computing task in the emergency cleanup queue according to the current task cleanup batch. The page cleanup records include page cleanup completed and page cleanup incomplete. Filter the page cleanup records as video memory pages with incomplete page cleanup to obtain the total number of incomplete residual data cleanup pages. Obtain the average residual data cleanup time for each video memory page. Calculate the product of the sum of the number of page requests and the total number of incomplete residual data cleanup pages with the average residual data cleanup time for each video memory page to obtain the estimated total processing time for each video memory page. The estimated total processing time for each video memory page is compared with the preset maximum tolerable waiting time threshold. When the estimated total processing time is greater than the maximum tolerable waiting time threshold, the page clearing records in the emergency clearing queue are monitored. Each time a page clearing record is generated that has been cleared, the total number of pages with incomplete residual data clearing is reduced by one. The estimated total processing time is then recalculated based on the reduced total number of pages with incomplete residual data clearing until the estimated total processing time is less than or equal to the maximum tolerable waiting time threshold. When the estimated total processing time is less than or equal to the maximum tolerable waiting time threshold, memory pages are selected sequentially from the first position in the residual data clearing execution order until the number of selected pages reaches the page request number. The selected memory pages are then separated from the set of pages to be cleared and pushed into the emergency clearing queue.
[0010] As a further aspect of the present invention, the process of clearing residual data left over from the previous computation task in the emergency clearing queue of video memory pages, obtaining all video memory pages after clearing the residual data, obtaining a second cleared page set, and integrating the first cleared page set and the second cleared page set to obtain an allocatable page set includes: The data left over from the previous computing task after being written or used in the video memory page of the emergency clearing queue is obtained as residual data. Zero-writing or overwrite clearing is performed on the residual data. After the zero-writing or overwrite clearing is completed, an emergency queue clearing completion event is fed back, and the emergency queue clearing completion confirmation information of the corresponding video memory page attached to the emergency queue clearing completion event is obtained. The emergency queue clearing completion confirmation information is used as the page clearing completion record of the corresponding video memory page. All video memory pages with page clearing completion records are combined to obtain the second cleared page set. The first set of cleared pages and the second set of cleared pages are combined to construct an allocatable set of pages, indicating that the video memory pages in both the first set of cleared pages and the second set of cleared pages have completed the residual data removal process and can continue to be allocated and used.
[0011] As a further aspect of the present invention, allocating the set of allocatable pages to the current computing task to obtain data management results includes: Select a number of video memory pages corresponding to the number of page requests from the first cleared page set and the second cleared page set of the allocable page set, and use them as safe and available video memory pages; Based on the video memory resource allocation request, the secure and available video memory pages are allocated to the current computing task to obtain the video memory allocation result; The memory allocation results are used as data management results of the GPU server to provide page allocation information of safe and available memory pages to the current computing task, and to provide the memory allocator with registration data of memory pages occupied by the current computing task that have undergone residual data clearing.
[0012] A GPU server data security management system, the system comprising: The task page acquisition module obtains the current computing task, the previous computing task corresponding to the current computing task, and the current video memory page recycling queue. It also collects video memory page information for each video memory page of the previous computing task contained in the current video memory page recycling queue. The page partitioning module partitions each video memory page according to the video memory page information, resulting in a set of pages to be cleared and a first set of cleared pages. The clearing order determination module determines the execution order of clearing residual data in each video memory page of the page set to be cleared; The emergency queue push module separates video memory pages from the set of pages to be cleared according to the execution order of residual data clearing, and pushes the separated video memory pages into the preset emergency clearing queue. The allocatable page construction module clears residual data left over from the previous computing task in the emergency clearing queue of video memory pages, obtains all video memory pages after clearing residual data, and obtains the second cleared page set. The first cleared page set and the second cleared page set are integrated to obtain the allocatable page set. The page allocation management module assigns the set of allocatable pages to the current computing task, thus obtaining the data management results.
[0013] Compared with the prior art, the advantages and positive effects of the present invention are as follows: This invention distinguishes between memory pages to be cleared and those that have already been cleared based on their memory page occupancy release status and residual data clearing completion status, preventing memory pages with unclear clearing status from directly entering the reallocation process. Subsequently, it determines the residual data clearing execution order by combining the security level of the previous computing task, the memory page reclamation queue waiting time, and the historical selection count of memory pages. This ensures that memory pages requiring priority processing enter the emergency clearing queue in the corresponding order, thus maintaining coordination between the security attributes and clearing order of memory pages. Furthermore, it controls the timing of memory page selection based on the number of pages requested by the current computing task, the number of memory pages not yet cleared in the emergency clearing queue, and the average residual data clearing time, allowing the residual data clearing process to adapt to the maximum tolerable waiting time of the current computing task. In addition, after a memory page completes zero-writing or overwrite clearing and forms a page clearing completion record, it is reassigned to the current computing task. This not only prevents residual data from the previous computing task from being read by subsequent computing tasks but also restricts the re-allocation of memory pages that have not yet been cleared, while also considering the timeliness of memory resource allocation. Attached Figure Description
[0014] Figure 1 This is a schematic diagram of the steps of the method of the present invention; Figure 2 This is a flowchart illustrating the steps and principles of the method of the present invention. Figure 3 This is a block diagram of the system of the present invention; Figure 4 This is a block diagram illustrating the module flow principle of the system of the present invention. Detailed Implementation
[0015] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0016] Please see Figure 1-2 This invention provides a technical solution, a method for data security management of a GPU server, comprising the following steps: S1: Obtain the current computation task, the previous computation task corresponding to the current computation task, and the current memory page reclamation queue. Collect memory page information for each memory page of the previous computation task contained in the current memory page reclamation queue, including: Get the GPU server's current memory resource allocation request, and take the computing task that sent the memory resource allocation request as the current computing task. At the same time, get the previous computing task corresponding to the current computing task and the current memory page reclamation queue corresponding to the current computing task. First, the GPU server reads the memory allocation request received from the GPU server. A memory allocation request is a request submitted by a computing task to the GPU server to request a specified number of memory pages. The request includes the computing task identifier, the request receipt time, and the number of pages requested. Based on the computing task identifier, the corresponding computing task registration information is read, and the corresponding computing task is identified as the current computing task. The current computing task refers to the computing task that has submitted the memory page request and is waiting to receive the memory pages. Next, according to the graphics processing unit (GPU) device used by the current computing task, the memory page occupancy record and memory page release record stored by the memory allocator are read. The memory allocator stores memory page addresses, memory page states, and the relationship between computing tasks and memory page occupancy, and handles memory page reclamation and reallocation. A memory page refers to a contiguous storage area of GPU memory divided into fixed-capacity blocks. The computational tasks that have finished occupying video memory pages are arranged in ascending order of their memory page release time. The computational task whose memory page release time is closest to the time the video memory resource allocation request is received is identified as the previous computational task corresponding to the current computational task. The previous computational task refers to the computational task that occupied and released the corresponding video memory page before the current computational task made a video memory page request. Subsequently, video memory pages whose page status has changed from occupied to reclaimed and have not yet been reallocated are read and arranged in ascending order of the time the video memory pages were returned to the video memory allocator. The current video memory page reclamation queue refers to the sequence of video memory pages that have been released from occupation but have not yet been reallocated, arranged in order of their return time. This yields the current computational task, the previous computational task corresponding to the current computational task, and the current video memory page reclamation queue.
[0017] Collect all memory pages in the current memory page reclamation queue, as well as the memory page information for each memory page. The memory page information includes the previous computing task's release flag, the previous computing task's data clearing completion flag, the previous computing task's security level, the memory page reclamation queue waiting time, and the historical selection count of the memory page. Among them, the previous computing task's release flag indicates that the memory page is no longer occupied by the previous computing task of the current computing task and has been returned to the current memory page reclamation queue. The previous computing task's data clearing completion flag indicates that the data left in the memory page after being written or used by the previous computing task has been written to zero or overwritten and cleared. All video memory pages are read in the current order of their place in the memory page reclamation queue. For each page, the memory page occupancy record, residual data clearing record, previous computation task registration information, the time the page returned to the current memory page reclamation queue, and the page selection record are read based on its memory page identifier. Memory page information refers to associated data indicating the memory page's occupancy release status, residual data clearing status, data security level, reclamation wait time, and historical selection status. When the occupancy status in the memory page occupancy record changes from being occupied by a previous computation task to being released, a previous computation task occupancy release identifier is written to the corresponding memory page record. This identifier indicates that the memory page has finished being occupied by the previous computation task and has re-entered the current memory page reclamation queue. When the residual data clearing record records that a memory page has completed write-to-zero or overwrite clearing, a previous computation task's residual data clearing completion identifier is written to the corresponding memory page record. This identifier indicates that the data left in the memory page after being written or used by the previous computation task has been completely overwritten. Retrieve the security level value from the registration information of the previous computation task based on the previous computation task identifier, and record the corresponding security level value as the security level of the previous computation task. The security level of the previous computation task refers to the security classification corresponding to the data processed by the previous computation task. Read the current time recorded by the GPU server clock, and the waiting time in the video memory page reclamation queue is equal to the current time. The time it takes for a video memory page to return to the current video memory page recycling queue is counted. The video memory page recycling queue waiting time refers to the cumulative time a video memory page has waited to be processed again after entering the current video memory page recycling queue. The cumulative number of video memory page selection records is counted according to the video memory page identifier, and this cumulative number is recorded as the historical selection count of the video memory page. The historical selection count of the video memory page refers to the cumulative number of times the video memory page has been selected based on the residual data cleanup execution order. This provides the video memory page information for each video memory page in the current video memory page recycling queue.
[0018] S2: Based on the video memory page information, each video memory page is divided to obtain the set of pages to be cleared and the first set of cleared pages, including: Referring to the video memory page information, each video memory page that has the previous computing task's unoccupation release flag but does not have the previous computing task's residual data clearing completion flag is grouped into a set of pages to be cleared; Based on the obtained video memory page information, the previous computational task's release flag and the previous computational task's residual data clearing completion flag are read according to the current order in the video memory page recycling queue. Video memory page flags with a previous computational task release flag but no previous computational task's residual data clearing completion flag are written to the same page record, along with the corresponding video memory page's position in the current video memory page recycling queue. The existence of a previous computational task release flag indicates that the corresponding video memory page has been released from the previous computational task's occupation; the absence of a previous computational task's residual data clearing completion flag indicates that the corresponding video memory page has not yet completed residual data clearing. The set of pages to be cleared refers to the set of video memory pages that have been released from the previous computational task's occupation, have returned to the current video memory page recycling queue, and have not yet completed residual data clearing. The video memory pages written are arranged from front to back according to their position in the current video memory page recycling queue, thus obtaining the set of pages to be cleared.
[0019] Referring to the video memory page information, each video memory page with a marker indicating that the data left over from the previous computing task has been cleared is grouped into a first set of cleared pages; For memory pages that have already been marked as having completed clearing residual data from the previous computation task, the corresponding memory page identifiers are read according to their order in the current memory page reclamation queue. The corresponding memory page identifier and its position in the current memory page reclamation queue are then written to the same page record. The first set of cleared pages refers to the set of memory pages that have completed zero-writing or overwrite clearing before the start of this emergency clearing process and do not need to re-enter the emergency clearing queue. The first set of cleared pages is obtained by arranging the written memory pages from front to back according to their position in the current memory page reclamation queue.
[0020] S3: Determine the execution order of residual data removal for each video memory page in the set of pages to be cleared, including: Get the security level of the previous computing task, the waiting time of the video memory page recycling queue, and the number of times the video memory page has been selected in the history of each video memory page in the set of pages to be cleared; The memory page identifier of each memory page is read according to the sorted order of the memory pages in the set of pages to be cleared, and the corresponding memory page information is located based on the memory page identifier. The security level of the previous computation task, the waiting time in the memory page recycling queue, and the historical selection count of the memory page are extracted from the memory page information. The security level of the previous computation task remains the original security level value in the registration information of the previous computation task; the waiting time in the memory page recycling queue is uniformly converted to milliseconds; and the historical selection count of the memory page remains the cumulative integer value formed by the statistical analysis of the memory page selection records. The security level of the previous computation task, the waiting time in the memory page recycling queue, and the historical selection count of the memory page are written into the same memory page sorting data record according to the memory page identifier, so that each memory page corresponds to one security level of the previous computation task, one waiting time in the memory page recycling queue, and one historical selection count of the memory page. This yields the security level of the previous computation task, the waiting time in the memory page recycling queue, and the historical selection count of the memory page corresponding to each memory page in the set of pages to be cleared.
[0021] Normalize the security level of the previous computing task, the waiting time of the video memory page reclamation queue, and the number of times the video memory page was selected in history, and generate normalized security level of the previous computing task, normalized waiting time of the video memory page reclamation queue, and normalized number of times the video memory page was selected in history. Sort the previous computation task security level, video memory page reclamation queue wait time, and video memory page historical selection count in ascending order of value. Normalization refers to converting values with different ranges and units of measurement to a unified range of 0 to 1. Read the smallest previous computation task security level at the top of the list and the largest previous computation task security level at the bottom; if the largest previous computation task security level is greater than the smallest, the normalized previous computation task security level = (previous computation task security level) / (previous computation task security level). (Minimum previous calculation task security level) ÷ (Maximum previous calculation task security level) The minimum and maximum security levels of the previous computational task are equal. When the maximum and maximum security levels of the previous computational task are equal, the normalized security level of the previous computational task is recorded as 0. Next, the shortest and longest video memory page reclamation queue wait times are read. When the longest video memory page reclamation queue wait time is greater than the shortest, the normalized video memory page reclamation queue wait time is equal to (video memory page reclamation queue wait time). (Shortest memory page reclamation queue wait time) ÷ (Longest memory page reclamation queue wait time) (Shortest memory page reclamation queue wait time); when the longest memory page reclamation queue wait time equals the shortest memory page reclamation queue wait time, the normalized memory page reclamation queue wait time is recorded as 0. Then, read the minimum and maximum historical memory page selection counts; when the maximum historical memory page selection count is greater than the minimum historical memory page selection count, the normalized historical memory page selection count = (historical memory page selection count) (Minimum number of times video memory pages were selected in history) ÷ (Maximum number of times video memory pages were selected in history) The minimum number of historical memory page selections is calculated; when the maximum number of historical memory page selections equals the minimum number of historical memory page selections, the normalized number of historical memory page selections is recorded as 0. This yields the normalized security level of the previous computation task, the normalized waiting time for the memory page reclamation queue, and the normalized number of historical memory page selections.
[0022] Weights are assigned to the normalized security level of the previous computing task, the normalized waiting time of the memory page reclamation queue, and the normalized historical selection count of the memory page. A weighted sum of these three factors is calculated to obtain the residual data removal priority score for each memory page. The execution order of residual data removal for each memory page in the set of pages to be removed is determined based on these residual data removal priority scores. The weights include a level weight assigned to the normalized security level of the previous computing task, a duration weight assigned to the normalized waiting time of the memory page reclamation queue, and a number weight assigned to the normalized historical selection count of the memory page. The magnitudes of the level weight, duration weight, and number weight are determined based on their respective proportions in the sum of the three normalized values. After normalization, for each video memory page in the set of pages to be cleared, the normalized security level of the previous computation task, the normalized waiting time in the video memory page reclamation queue, and the normalized historical selection count of the video memory page are read, and the sum of these three normalized values is calculated. The sum of the three normalized values = normalized security level of the previous computation task + normalized waiting time in the video memory page reclamation queue + normalized historical selection count of the video memory page. When the sum of the three normalized values is greater than 0, the grade weight = the normalized security level of the previous calculation task ÷ the sum of the three normalized values; the duration weight = the normalized waiting time in the video memory page recycling queue ÷ the sum of the three normalized values; and the number of times weight = the normalized number of times the video memory page was selected in history ÷ the sum of the three normalized values. The greater the proportion of the normalized security level of the previous calculation task in the sum of the three normalized values, the greater the grade weight; the greater the proportion of the normalized waiting time in the video memory page recycling queue in the sum of the three normalized values, the greater the duration weight; and the greater the proportion of the normalized number of times the video memory page was selected in history, the greater the number of times weight. When the sum of the three normalized values is equal to 0, the grade weight, duration weight, and number of times weight are set to the same value, so that the sum of the grade weight, duration weight, and number of times weight equals 1. The weighted sum refers to calculating the product of each of the three normalized values and its corresponding weight, and then summing the three products. The residual data removal priority score is calculated as follows: Normalized security level of the previous computation task × level weight + Normalized waiting time in the video memory page reclamation queue × duration weight + Normalized historical selection count of video memory pages × count weight. A higher residual data removal priority score corresponds to a higher position of the corresponding video memory page in the set of pages to be removed. Video memory pages are arranged from highest to lowest residual data removal priority score. If residual data removal priority scores are the same, they are arranged from front to back according to their original position in the set of pages to be removed. This determines the execution order of residual data removal for each video memory page in the set of pages to be removed.
[0023] S4: Separate video memory pages from the set of pages to be cleared according to the residual data clearing execution order, and push the separated video memory pages into the preset emergency clearing queue, including: Get the memory resource allocation request sent by the current computing task, and the number of pages requested in the memory resource allocation request, as the clearing batch for the current task; After the current computing task sends a memory resource allocation request, the corresponding memory resource allocation request is read according to the current computing task identifier, and the page allocation number field in the memory resource allocation request is located. The page allocation number refers to the total number of memory pages requested by the current computing task in this memory resource allocation request, which is directly read according to the integer number of pages registered in the memory resource allocation request. The page allocation number, the current computing task identifier, and the memory resource allocation request identifier are written into the same cleanup batch record. The current task cleanup batch refers to the number of memory pages corresponding to this memory resource allocation request that need to be selected from the set of pages to be cleaned and the residual data cleaned. Thus, the current task cleanup batch is obtained.
[0024] Obtain the emergency cleanup queue pre-set by the GPU server for the current computing task. Read the page cleanup records of the video memory pages corresponding to the current computing task in the emergency cleanup queue according to the cleanup batch of the current task. The page cleanup records include page cleanup completed and page cleanup incomplete. Filter the page cleanup records as video memory pages with incomplete page cleanup to obtain the total number of incomplete residual data cleanup pages. Obtain the average residual data cleanup time for each video memory page. Calculate the product of the sum of the number of page requests and the total number of incomplete residual data cleanup pages with the average residual data cleanup time for each video memory page to obtain the estimated total processing time for each video memory page. The system reads the emergency cleanup queue pre-established by the GPU server for the current computing task based on the current computing task identifier. The emergency cleanup queue receives and stores the sequence of memory pages requiring immediate cleanup according to the residual data cleanup execution order. It reads the page cleanup records of the memory pages corresponding to the current computing task in the emergency cleanup queue according to the current task's cleanup batch. Page cleanup records are data records of the memory page cleanup status, start time, and completion time; page cleanup completion indicates that the corresponding memory page has completed zero-writing or overwrite cleanup processing, while page cleanup incomplete indicates that the corresponding memory page has not yet completed zero-writing or overwrite cleanup processing. It filters memory pages with a cleanup status of incomplete page cleanup and counts the number of corresponding page cleanup records. The total number of pages with incomplete residual data cleanup equals the number of page cleanup records with a cleanup status of incomplete page cleanup. The total number of pages with incomplete residual data cleanup refers to the number of memory pages in the emergency cleanup queue corresponding to the current computing task that have not yet completed residual data cleanup. The system records the cleanup start time when a memory page begins residual data cleanup and the cleanup completion time when a memory page completes residual data cleanup. The residual data cleanup time for a single memory page equals the cleanup completion time. The cleanup start time is calculated as follows: The average residual data cleanup time per memory page refers to the average cleanup time of all memory pages that have been cleaned. The average residual data cleanup time per memory page = the sum of the residual data cleanup times of all individual memory pages ÷ the number of records completed for page cleanup. The estimated total processing time per memory page refers to the estimated total time required to complete the cleanup of the memory pages corresponding to the number of page requests and the remaining memory pages in the emergency cleanup queue that have not yet been cleaned. The estimated total processing time per memory page = (number of page requests + total number of pages with incomplete residual data cleanup) × the average residual data cleanup time per memory page. This gives the estimated total processing time for each memory page.
[0025] The estimated total processing time for each video memory page is compared with the preset maximum tolerable waiting time threshold. When the estimated total processing time exceeds the maximum tolerable waiting time threshold, the page clearing records in the emergency clearing queue are monitored. For each page clearing record that has been cleared, the total number of pages with incomplete residual data clearing is reduced by one. The estimated total processing time is then recalculated based on the reduced total number of pages with incomplete residual data clearing until the estimated total processing time is less than or equal to the maximum tolerable waiting time threshold. The process for setting the maximum tolerable waiting time threshold is as follows: the latest video memory occupancy time and the video memory resource allocation request reception time of the current computing task are obtained. The time difference between the latest video memory occupancy time and the video memory resource allocation request reception time is calculated to obtain the task's waiting time. The waiting reduction coefficient corresponding to the current computing task priority is obtained. The task's waiting time is multiplied by the waiting reduction coefficient to obtain the maximum tolerable waiting time threshold. The higher the current computing task priority, the smaller the waiting reduction coefficient; the lower the current computing task priority, the larger the waiting reduction coefficient. The latest memory occupancy time is read from the task registration information of the current computing task, and the memory allocation request reception time is read from the memory allocation request. The latest memory occupancy time refers to the latest time the current computing task should begin occupying memory pages, and the memory allocation request reception time refers to the time recorded by the GPU server when it receives the memory allocation request sent by the current computing task. Task wait time = Latest memory occupancy time The video memory resource allocation request is received. Next, the corresponding waiting reduction coefficient is read according to the task priority of the current computing task. The waiting reduction coefficient is the proportion used to shorten the waiting time of the task according to the current computing task priority; the waiting reduction coefficient is greater than 0 and less than or equal to 1. All task priorities are pre-arranged in descending order, with the highest priority task number recorded as 1, and the priority numbers of the remaining tasks increasing sequentially. The total number of task priorities is counted. The waiting reduction coefficient = current computing task priority number ÷ total number of task priorities. The higher the current computing task priority, the smaller the current computing task priority number, and the smaller the waiting reduction coefficient; the lower the current computing task priority, the larger the current computing task priority number, and the larger the waiting reduction coefficient. The maximum tolerable waiting time threshold = task waiting time × waiting reduction coefficient. The maximum tolerable waiting time threshold refers to the maximum time allowed for the current computing task to wait for the video memory page to complete the clearing of residual data. Subsequently, the estimated total processing time of each video memory page is compared with the maximum tolerable waiting time threshold. When the estimated total processing time exceeds the maximum tolerable waiting time threshold, newly generated page clearing records in the emergency clearing queue are continuously read according to the generation time of the page clearing record; for each page clearing record that has been cleared, the total number of pages with incomplete residual data clearing after the update equals the total number of pages with incomplete residual data clearing before the update. 1. The recalculated total processing time = (number of page requests + total number of pages with incomplete residual data clearing after the update) × average residual data clearing time per video memory page, and continues to compare the recalculated total processing time with the maximum tolerable waiting time threshold. When the total processing time is less than or equal to the maximum tolerable waiting time threshold, the reduction of the total number of pages with incomplete residual data clearing is stopped, thus obtaining a total processing time that is less than or equal to the maximum tolerable waiting time threshold.
[0026] When the total processing time is less than or equal to the maximum tolerable waiting time threshold, the video memory pages are selected sequentially from the first position in the residual data clearing execution order until the number of selected pages reaches the page request number. The selected video memory pages are then separated from the set of pages to be cleared and pushed into the emergency clearing queue. Once the estimated total processing time for each video memory page is less than or equal to the preset maximum tolerable waiting time threshold, the first video memory page identifier in the residual data cleanup execution order is read. For each video memory page identifier read, the selection count is increased by 1; the selection count refers to the number of video memory pages that have been determined from the set of pages to be cleaned and need to be pushed into the emergency cleanup queue. When the selection count is less than the page allocation count, the next video memory page in the residual data cleanup execution order is read; when the selection count reaches the page allocation count, reading stops. Subsequently, the corresponding video memory pages are removed from the set of pages to be cleaned according to the selected video memory page identifiers, and the removed video memory pages are written into the emergency cleanup queue sequentially according to the residual data cleanup execution order. Separation of a video memory page from the set of pages to be cleaned means that the video memory page is no longer retained in the set of pages to be cleaned and is transferred to the emergency cleanup queue to await residual data cleanup. This yields the video memory pages separated from the set of pages to be cleaned and pushed into the emergency cleanup queue.
[0027] S5: Clear residual data left over from the previous computation task in the emergency clearing queue of video memory pages. Obtain all video memory pages after clearing residual data to obtain the second cleared page set. Integrate the first cleared page set and the second cleared page set to obtain the allocatable page set, which includes: Get the data left over from the previous computing task in the video memory page of the emergency clearing queue as residual data, perform write zeroing or overwrite clearing on the residual data, and after completing write zeroing or overwrite clearing, send back an emergency queue clearing completion event, and get the emergency queue clearing completion confirmation information of the corresponding video memory page attached to the emergency queue clearing completion event. When performing residual data erasure on memory pages in the emergency erasure queue, the physical start address and page capacity of the memory page are read. The physical start address refers to the first storage address of the memory page in the graphics processor's physical memory, and the page capacity refers to the length of storage continuously allocated to the memory page starting from the physical start address. Residual data refers to the bytes of content written or used by the previous computing task within the aforementioned continuous storage range that have not yet been erased. Write-zero processing refers to rewriting the byte content of all storage locations of the memory page to the value 0; starting from the physical start address, the value 0 is written to continuous memory addresses. After each write, the number of bytes written is read, and the next write address = current write address + number of bytes written. Writing continues as long as the cumulative number of bytes written is less than the page capacity, and write-zero processing is completed when the cumulative number of bytes written equals the page capacity. Overwrite and clearing refers to continuously rewriting the original byte content of all memory locations in a video memory page using overwrite data. Starting from the physical start address, overwrite data is written to consecutive video memory addresses. After each overwrite, the number of bytes overwritten is read, and the next overwrite address equals the current overwrite address plus the number of bytes overwritten. Overwriting continues until the cumulative number of bytes overwritten equals the page capacity, completing the overwrite and clearing process. After completing either zero-writing or overwrite and clearing, the video memory page identifier, clearing method, and clearing completion time are written to the emergency queue clearing completion event. The emergency queue clearing completion event is a clearing completion notification generated after all addresses of the video memory page have completed zero-writing or overwrite and clearing. The video memory page identifier and clearing completion status attached to the emergency queue clearing completion event are read. The emergency queue clearing completion confirmation information is data generated along with the emergency queue clearing completion event to confirm that the corresponding video memory page has completed the clearing of residual data. This yields the emergency queue clearing completion confirmation information for the corresponding video memory page.
[0028] The emergency queue clearing completion confirmation information is used as the page clearing completion record of the corresponding video memory page. All video memory pages with page clearing completion records are combined to obtain the second cleared page set. After obtaining the emergency queue clearing completion confirmation information, a one-to-one correspondence is established between the emergency queue clearing completion confirmation information and the video memory pages in the emergency clearing queue according to the video memory page identifier. A page clearing completion record refers to a data record formed based on the emergency queue clearing completion confirmation information, which records the video memory page identifier, clearing method, clearing completion time, and clearing completion status. The video memory page identifier, clearing method, clearing completion time, and clearing completion status from the emergency queue clearing completion confirmation information are written into the page clearing completion record of the corresponding video memory page. Subsequently, all video memory pages with page clearing completion records are read according to their order in the emergency clearing queue, and the corresponding video memory page identifiers are sequentially written into the same page set. The second cleared page set refers to the set of video memory pages that have been processed by the emergency clearing queue and for which page clearing completion records have been formed. This yields the second cleared page set.
[0029] Combine the first set of cleared pages and the second set of cleared pages to construct an allocatable set of pages, indicating that the video memory pages in both the first set of cleared pages and the second set of cleared pages have completed the residual data removal process and can continue to be allocated and used; When integrating the first and second cleared page sets, all video memory page identifiers are first read in the order listed in the first cleared page set, and then in the order listed in the second cleared page set. The first cleared page set corresponds to video memory pages whose residual data was cleared before entering the emergency clearing process, and the second cleared page set corresponds to video memory pages whose residual data was cleared after the emergency clearing process. The two sets of video memory page identifiers are sequentially written into the same page set record, while retaining the page clearing completion record and source set identifier for each video memory page. The allocatable page set refers to the set of video memory pages that have completed residual data clearing and are currently available for reallocation to computing tasks. This yields the allocatable page set.
[0030] S6: Allocate the set of allocatable pages to the current computation task, and obtain the data management results, including: Select a number of video memory pages corresponding to the number of page requests from the first cleared page set and the second cleared page set of the allocatable page set, and use them as safe and available video memory pages; When selecting video memory pages from the allocatable page set, the number of page requests is read, and then video memory page identifiers are read sequentially from the first in the allocatable page set. Safe and available video memory pages refer to those that have undergone residual data clearing, are currently allocatable, and have been selected for allocation to the current computing task. Each time a video memory page identifier is read, the selection count is incremented by 1, and the corresponding video memory page identifier is retained. If the selected count is less than the number of page requests, the next video memory page identifier is read; when the selected count reaches the number of page requests, reading stops. This yields the number of safe and available video memory pages corresponding to the number of page requests.
[0031] Based on the video memory resource allocation request, allocate safe and available video memory pages to the current computing task to obtain the video memory allocation result; Based on the current computing task identifier in the video memory resource allocation request, all secure and available video memory page identifiers are read, and the occupancy mapping between the current computing task identifier and the secure and available video memory page identifiers is written to the video memory allocator item by item. The occupancy mapping refers to the association record formed between the current computing task and the video memory pages allocated to it. For each occupancy mapping written, the page state of the corresponding secure and available video memory page is changed from allocable to occupied by the current computing task, and the corresponding secure and available video memory page is removed from the allocable page set. After writing the occupancy mappings for all secure and available video memory pages, the number of allocated secure and available video memory pages is counted, and the current computing task identifier, the number of allocated secure and available video memory page identifiers, and the number of allocated secure and available video memory pages are summarized. The video memory allocation result refers to the data recording the secure and available video memory page identifiers and the number of video memory pages actually acquired by the current computing task. This yields the video memory allocation result.
[0032] The video memory allocation results are used as the data management results of the GPU server to provide page allocation information of safe and available video memory pages to the current computing task, and to provide the video memory allocator with the registration data of video memory pages occupied by the current computing task that have completed the residual data clearing process. Finally, the current computing task identifier, the allocated secure available memory page identifiers, and the number of allocated secure available memory pages are read from the memory allocation results. The GPU server's data management results refer to the result data formed by the page allocation information obtained by the current computing task and the memory page registration data saved by the memory allocator. The page allocation information points to the secure available memory page identifiers provided by the current computing task and the number of allocated secure available memory pages; the registration data refers to the occupancy record between the current computing task and the memory pages that have undergone residual data clearing processing, saved by the memory allocator. According to the current computing task identifier, the allocated secure available memory page identifiers and the number of allocated secure available memory pages are written into the page allocation information corresponding to the current computing task. Subsequently, according to the allocated secure available memory page identifiers, the occupancy correspondence between the current computing task identifier and the secure available memory page identifiers is read item by item, and the current computing task identifier, secure available memory page identifiers, page clearing completion status, and current computing task occupancy status are written into the memory page registration data of the memory allocator. This yields the GPU server's data management results.
[0033] Please see Figure 3-4 A GPU server data security management system, the system comprising: The task page acquisition module obtains the current computing task, the previous computing task corresponding to the current computing task, and the current video memory page recycling queue. It also collects video memory page information for each video memory page of the previous computing task contained in the current video memory page recycling queue. The page partitioning module partitions each video memory page according to the video memory page information, resulting in a set of pages to be cleared and a first set of cleared pages. The clearing order determination module determines the execution order of clearing residual data in each video memory page of the page set to be cleared; The emergency queue push module separates video memory pages from the set of pages to be cleared according to the execution order of residual data clearing, and pushes the separated video memory pages into the preset emergency clearing queue. The allocatable page construction module clears residual data left over from the previous computing task in the emergency clearing queue of video memory pages, obtains all video memory pages after clearing residual data, and obtains the second cleared page set. The first cleared page set and the second cleared page set are integrated to obtain the allocatable page set. The page allocation management module assigns the set of allocatable pages to the current computing task, thus obtaining the data management results.
[0034] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A method for data security management of a GPU server, characterized in that, Includes the following steps: Obtain the current computing task, the previous computing task corresponding to the current computing task, and the current memory page recycling queue. Collect memory page information for each memory page of the previous computing task contained in the current memory page recycling queue. Each video memory page is divided according to the video memory page information to obtain a set of pages to be cleared and a first set of cleared pages; Determine the execution order for clearing residual data in each video memory page of the set of pages to be cleared; Based on the residual data removal execution order, separate the video memory pages from the set of pages to be removed, and push the separated video memory pages into the preset emergency removal queue; Clear residual data left over from the previous computation task in the emergency clearing queue of video memory pages, obtain all video memory pages after clearing residual data, obtain the second cleared page set, and integrate the first cleared page set and the second cleared page set to obtain the allocatable page set. Assign the set of allocatable pages to the current computing task to obtain the data management results.
2. The GPU server data security management method according to claim 1, characterized in that, The step of obtaining the current computing task, the previous computing task corresponding to the current computing task, and the current video memory page reclamation queue, and collecting video memory page information for each video memory page of the previous computing task contained in the current video memory page reclamation queue includes: Get the GPU server's current memory resource allocation request, and take the computing task that sent the memory resource allocation request as the current computing task. At the same time, get the previous computing task corresponding to the current computing task and the current memory page reclamation queue corresponding to the current computing task. Collect all video memory pages in the current video memory page reclamation queue, as well as the video memory page information for each video memory page. The video memory page information includes the previous computing task's release flag, the previous computing task's legacy data clearing completion flag, the previous computing task's security level, the video memory page reclamation queue waiting time, and the video memory page's historical selection count. Among them, the previous computing task's release flag indicates that the video memory page is no longer occupied by the previous computing task of the current computing task and has been returned to the current video memory page reclamation queue. The previous computing task's legacy data clearing completion flag indicates that the data left in the video memory page after being written or used by the previous computing task has been written to zero or overwritten and cleared.
3. The GPU server data security management method according to claim 2, characterized in that, The process of dividing each video memory page based on the reference video memory page information to obtain a set of pages to be cleared and a first set of cleared pages includes: Referring to the aforementioned video memory page information, each video memory page that has a previous computing task's unoccupation flag but does not have a previous computing task's data clearing completion flag is grouped into a set of pages to be cleared. Referring to the aforementioned video memory page information, each video memory page with an identifier indicating that the data left over from the previous computing task has been cleared is grouped into a first set of cleared pages.
4. The GPU server data security management method according to claim 3, characterized in that, The process of determining the execution order of residual data removal for each video memory page in the set of pages to be removed includes: Obtain the security level of the previous computing task, the waiting time of the video memory page recycling queue, and the number of times the video memory page has been selected in the history of each video memory page in the set of pages to be cleared; Normalization is performed on the security level of the previous computing task, the waiting time of the video memory page reclamation queue, and the number of times the video memory page was selected in history, respectively, to generate normalized security level of the previous computing task, normalized waiting time of the video memory page reclamation queue, and normalized number of times the video memory page was selected in history. Assign weights to the normalized security level of the previous computation task, the normalized waiting time of the memory page reclamation queue, and the normalized historical selection count of the memory page. Calculate the weighted sum of the normalized security level of the previous computation task, the normalized waiting time of the memory page reclamation queue, and the normalized historical selection count of the memory page to obtain the residual data removal priority score for each memory page. Determine the residual data removal execution order for each memory page in the set of pages to be removed based on the residual data removal priority score.
5. The GPU server data security management method according to claim 4, characterized in that, The step of separating video memory pages from the set of pages to be cleared according to the residual data clearing execution order, and pushing the separated video memory pages into a preset emergency clearing queue includes: Get the memory resource allocation request sent by the current computing task, and the number of pages requested in the memory resource allocation request, as the clearing batch for the current task; Obtain the emergency cleanup queue pre-set by the GPU server for the current computing task. Read the page cleanup records of the video memory pages corresponding to the current computing task in the emergency cleanup queue according to the current task cleanup batch. The page cleanup records include page cleanup completed and page cleanup incomplete. Filter the page cleanup records as video memory pages with incomplete page cleanup to obtain the total number of incomplete residual data cleanup pages. Obtain the average residual data cleanup time for each video memory page. Calculate the product of the sum of the number of page requests and the total number of incomplete residual data cleanup pages with the average residual data cleanup time for each video memory page to obtain the estimated total processing time for each video memory page. The estimated total processing time for each video memory page is compared with the preset maximum tolerable waiting time threshold. When the estimated total processing time is greater than the maximum tolerable waiting time threshold, the page clearing records in the emergency clearing queue are monitored. Each time a page clearing record is generated that has been cleared, the total number of pages with incomplete residual data clearing is reduced by one. The estimated total processing time is then recalculated based on the reduced total number of pages with incomplete residual data clearing until the estimated total processing time is less than or equal to the maximum tolerable waiting time threshold. When the estimated total processing time is less than or equal to the maximum tolerable waiting time threshold, memory pages are selected sequentially from the first position in the residual data clearing execution order until the number of selected pages reaches the page request number. The selected memory pages are then separated from the set of pages to be cleared and pushed into the emergency clearing queue.
6. The GPU server data security management method according to claim 5, characterized in that, The process involves clearing residual data left over from the previous computation task in the emergency clearing queue of video memory pages, obtaining all video memory pages after clearing the residual data, and obtaining a second cleared page set. The first and second cleared page sets are then integrated to obtain an allocatable page set, which includes: The data left over from the previous computing task after being written or used in the video memory page of the emergency clearing queue is obtained as residual data. Zero-writing or overwrite clearing is performed on the residual data. After the zero-writing or overwrite clearing is completed, an emergency queue clearing completion event is fed back, and the emergency queue clearing completion confirmation information of the corresponding video memory page attached to the emergency queue clearing completion event is obtained. The emergency queue clearing completion confirmation information is used as the page clearing completion record of the corresponding video memory page. All video memory pages with page clearing completion records are combined to obtain the second cleared page set. The first set of cleared pages and the second set of cleared pages are combined to construct an allocatable set of pages, indicating that the video memory pages in both the first set of cleared pages and the second set of cleared pages have completed the residual data removal process and can continue to be allocated and used.
7. The GPU server data security management method according to claim 6, characterized in that, The process of allocating the set of allocatable pages to the current computing task to obtain data management results includes: Select a number of video memory pages corresponding to the number of page requests from the first cleared page set and the second cleared page set of the allocable page set, and use them as safe and available video memory pages; Based on the video memory resource allocation request, the secure and available video memory pages are allocated to the current computing task to obtain the video memory allocation result; The memory allocation results are used as data management results of the GPU server to provide page allocation information of safe and available memory pages to the current computing task, and to provide the memory allocator with registration data of memory pages occupied by the current computing task that have undergone residual data clearing.
8. The GPU server data security management method according to claim 4, characterized in that, The weights include a level weight assigned to the normalized security level of the previous computing task, a duration weight assigned to the normalized memory page reclamation queue waiting time, and a frequency weight assigned to the normalized historical selection count of memory pages. The level weight, duration weight, and frequency weight are determined based on the proportions of the normalized security level of the previous computing task, the normalized memory page reclamation queue waiting time, and the normalized historical selection count of memory pages in the sum of the three normalized values.
9. The GPU server data security management method according to claim 5, characterized in that, The process of setting the maximum tolerable waiting time threshold is as follows: obtain the latest memory occupancy time and memory resource allocation request reception time of the current computing task; calculate the time difference between the latest memory occupancy time and the memory resource allocation request reception time to obtain the task's waiting time; obtain the waiting reduction coefficient corresponding to the current computing task's priority; multiply the task's waiting time by the waiting reduction coefficient to obtain the maximum tolerable waiting time threshold. The higher the current computing task's priority, the smaller the waiting reduction coefficient; the lower the current computing task's priority, the larger the waiting reduction coefficient.
10. A GPU server data security management system, characterized in that, The system includes: The task page acquisition module obtains the current computing task, the previous computing task corresponding to the current computing task, and the current video memory page recycling queue. It also collects video memory page information for each video memory page of the previous computing task contained in the current video memory page recycling queue. The page partitioning module partitions each video memory page according to the video memory page information, resulting in a set of pages to be cleared and a first set of cleared pages. The clearing order determination module determines the execution order of clearing residual data in each video memory page of the page set to be cleared; The emergency queue push module separates video memory pages from the set of pages to be cleared according to the execution order of residual data clearing, and pushes the separated video memory pages into the preset emergency clearing queue. The allocatable page construction module clears residual data left over from the previous computing task in the emergency clearing queue of video memory pages, obtains all video memory pages after clearing residual data, and obtains the second cleared page set. The first cleared page set and the second cleared page set are integrated to obtain the allocatable page set. The page allocation management module assigns the set of allocatable pages to the current computing task, thus obtaining the data management results.