Method for controlling multi-user concurrent access of server

By analyzing task characteristics and server resource status in real time, performing space-time coupling matching and quantum sharding scheduling, the shortcomings of traditional task scheduling methods in multi-user concurrent access control are solved, and more efficient computing resource allocation and data access optimization are achieved.

CN120144313AActive Publication Date: 2025-06-13WUHAN SPARK ZHONGDA INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510323830.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-13
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

Traditional task scheduling methods are difficult to effectively control multi-user concurrent access in modern server cluster environments, resulting in improper allocation of computing resources, increased data access latency, reduced task throughput and high resource conflict rates.

Method used

By analyzing the task type and resource demand characteristics requested by the user in real time, a triple of calculation pattern identifiers, data topology fingerprints and time constraint factors are generated, and space-time coupling matching is performed based on the real-time resource topology status of the server cluster, space-time weight scores are calculated, and binding execution units are generated. Then, quantized sharding scheduling is implemented on the binding execution unit, and the calculation pattern matching is adjusted through the execution trajectory feeding mechanism.

Benefits of technology

Effectively reduce the computing resource adaptation bias of tasks, reduce remote data access overhead, improve task throughput and execution stability, and improve computing concurrency and resource scheduling flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144313A_ABST
    Figure CN120144313A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of access control, in particular to a method for controlling multi-user concurrent access of a server, which comprises the following steps of: analyzing task types and resource demand characteristics requested by users in real time, and generating a triple comprising a calculation mode identifier, a data topology fingerprint and a time constraint factor; performing space-time coupling matching on the triad and a server node, calculating a space-time weight score, and generating a binding execution unit based on the space-time weight score; quantum fragmentation scheduling is carried out on the bound execution unit, a reverse dependence barrier is inserted between adjacent atomic operation sequences, and execution track data is generated in execution; and adjusting calculation mode matching through an execution track feedback mechanism, and performing calculation mode identifier remapping of an unexecuted task by utilizing the path efficiency of the completed task. According to the method, the computing resource adaptation deviation of the task is effectively reduced, the remote data access overhead is reduced, and the task throughput and the execution stability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of access control, and particularly to a control method for multi-user concurrent access to a server. Background Art

[0002] In a modern server cluster environment, multi-user concurrent access control is a core issue in task scheduling and computing resource management. With the increasing growth of high-performance computing, cloud computing, and artificial intelligence training tasks, server resource allocation faces challenges from various task requirements such as compute-intensive and storage-intensive tasks. Most traditional task scheduling methods adopt static policies, such as fixed resource allocation based on task types or simple scheduling rules based on load balancing. These methods have many deficiencies in terms of computing resource adaptability, data local access optimization, and task execution stability, and are difficult to meet the requirements of complex concurrent computing environments.

[0003] There is insufficient accurate identification and matching of task computing patterns. Existing task scheduling usually only performs static matching based on the computing resource requirements declared in the user request, without deeply analyzing the actual computing patterns of tasks, such as characteristics like vectorized computing (SIMD), GPU parallel computing, and storage access patterns. This results in low resource utilization. Especially in SIMD-intensive tasks, GPU computing tasks, and large-scale memory operation tasks, the deviation in computing pattern recognition may lead to a decline in the computing efficiency of task scheduling.

[0004] Secondly, the matching between data access topology and computing tasks is insufficient. Data in a server cluster is usually distributed across different storage levels, such as local SSD, NVMe, cache, distributed storage (such as HDFS), etc. Traditional scheduling methods do not fully consider the data access paths of tasks, resulting in tasks possibly being scheduled to nodes with mismatched storage topologies, thereby increasing remote data access latency, reducing task throughput, and even causing data transmission bottlenecks, affecting the overall computing performance.

[0005] In addition, the scheduling strategy for task splitting and synchronization control is relatively crude. Existing parallel task scheduling usually splits using fixed time slices or task blocks, without dynamically adjusting according to the current load state of the server. This leads to overly large task granularity in high-load situations, blocking computing resources; and overly small task granularity in low-load situations, increasing scheduling overhead and reducing throughput. At the same time, traditional task synchronization methods usually insert barriers based on static dependency rules, without intelligently inserting them according to the probability of resource conflicts, which may lead to excessive synchronization or resource competition, affecting task execution efficiency. Summary of the Invention

[0006] The present invention provides a control method for multi-user concurrent access to a server.

[0007] Control method for concurrent access of multiple users to a server, including the following steps:

[0008] S1: Real-time analyze the task type and resource requirement characteristics of the user request, and generate a triple including a computing mode identifier, a data topology fingerprint, and a time constraint factor;

[0009] S2: Based on the real-time resource topology status of the server cluster, perform spatio-temporal coupling matching between the triple and the server nodes, calculate the spatio-temporal weight score, and generate a bound execution unit based on the spatio-temporal weight score. The bound execution unit includes a combination of server nodes that meet the screening constraint conditions. The screening constraint conditions include (filtering candidate nodes in the following priority order):

[0010] Computing mode matching layer: Match the computing mode identifier with the server characteristics;

[0011] Data topology adaptation layer: Match the coincidence degree of the storage level access paths;

[0012] Time window verification layer: The resource reservation time slice corresponding to the time constraint factor covers the expected execution window of the task;

[0013] S3: Implement quantization sharding scheduling for the bound execution unit, split each task into multiple atomic operation sequences, and insert reverse dependency barriers between adjacent atomic operation sequences according to the dynamic load rate of the server nodes. Execution trajectory data is also generated during execution;

[0014] S4: Adjust the computing mode matching through the execution trajectory feedback mechanism, and remap the computing mode identifier of the unexecuted task using the path efficiency of the completed tasks.

[0015] Optionally, the S1 specifically includes:

[0016] S11: Based on the maximum score method, identify the computing mode identifier M;

[0017] M = argmax(S cpu , S gpu , S mem ), where M is the computing mode identifier of the task, which is used to represent the main computing characteristics of the task in server resource allocation. S cpu is the CPU computing intensity score, S gpu is the GPU computing intensity score, S mem is the memory-intensive score;

[0018] S12: Extract the storage level distribution and data access paths of the task-related data, and construct a data topology fingerprint. The data topology fingerprint is measured by the coincidence degree C of the storage level access paths:

[0019] Among them, is the storage location identifier of the i-th data block of the task, is the local storage identifier of server node j, δ is the location matching function, and n represents the number of data blocks accessed by the task;

[0020] S13: Dynamically calculate the time constraint factor, which is reflected based on the dynamic urgency:

[0021] Parse the absolute deadline T declared by the task deadline , and calculate the dynamic urgency U(t):

[0022] Among them, T resrve is the benchmark reserved time window (taking the average value) statistically based on the execution time of historical similar tasks, t current is the current time, that is, the server system time when the task is parsed, and α(t) is the dynamic adjustment coefficient based on the current resource competition state.

[0023] Optionally, in the calculation mode matching layer:

[0024] If the calculation mode identifier is GPU-intensive, filter server nodes with CUDA core count ≥ task requirements and PCIE bandwidth ≥ 16GB / s;

[0025] If it is CPU-intensive, filter server nodes that support the AVX-512 instruction set and L3 cache capacity ≥ 1.2 times the task data volume;

[0026] If it is memory-intensive, filter server nodes with memory bandwidth utilization rate < 60% within the NUMA node;

[0027] In the data topology adaptation layer, based on the storage hierarchy access path coincidence degree C, only retain server nodes with C ≥ 80% to reduce the remote data access overhead.

[0028] Optionally, in the time window verification layer, the task expected execution window T required is calculated based on the estimated execution time T exec est combined with the dynamic urgency U(t), and search and filter whether there are continuous resource reservation time slices [t start , t end in the node resource reservation table that satisfy t end -t start ≥ T required of the server node.

[0029] Optionally, S2 includes calculating the spatio-temporal weight score for the server nodes passed the screening:

[0030] Among them, Score is the spatio-temporal weight score of the current computing node, α is the data topology weight, γ is the time adaptation weight, λ is the load balancing weight, C is the coincidence degree of the storage hierarchy access path, and T reserve is the benchmark reserved time window, that is, the length of the reservable time slice of the current node, and T required is the task expected execution window, and Q pending is the length of the task queue to be processed by the current node at present, and Q max is the maximum task queue capacity of the current node;

[0031] Select the node with the highest spatio-temporal weight score to generate a bound execution unit.

[0032] Optionally, the bound execution unit includes:

[0033] (a) Hardware resource lock: Bind the exclusive resource combination identifier, including:

[0034] GPU resources (GPU card serial numbers, such as GPU-1, GPU-3).

[0035] CPU resources (range of CPU physical core IDs, such as CPUcore4-7).

[0036] Storage resources (NVMe storage ID, such as NVMe-2).

[0037] (b) Spatio-temporal contract: The time and space allocation protocol between the task and the server node, including:

[0038] The maximum execution window allowed for the task: that is, the time constraint for the task to be completed at the latest;

[0039] Data access path mapping: including whether data migration is allowed and whether the local cache acceleration strategy is adopted.

[0040] Optionally, the S3 specifically includes:

[0041] S31, define the quantization splitting rule of the atomic operation sequence:

[0042] S311, each atomic operation satisfies the minimum resource exclusivity;

[0043] S312, the splitting granularity G is adaptively adjusted according to the node dynamic load rate η;

[0044] S32. Insertion strategy for constructing reverse dependency barriers: Predict the resource conflict risk between adjacent atomic operation sequences. If the time when a previous operation releases a certain resource is later than the time when a subsequent operation requests the same resource, mark that resource as conflicting. Calculate the probability of conflict occurrence based on historical conflict data and the current load situation. When the conflict probability exceeds the set threshold, insert a reverse dependency barrier between the previous operation and the subsequent operation to ensure that the subsequent operation can only be executed within the protection time interval after the previous operation releases the resource.

[0045] Optionally, the execution trace data is generated during task execution by recording the spatio-temporal distribution information of the atomic operation sequence, including the unique identifier of the operation, the server resources occupied, the start time and end time of execution. For tasks with dependency relationships, record the number of the pre-dependency barrier, and at the same time predict the post-conflict.

[0046] Optionally, S4 includes evaluating the benefit of the storage level access path for the execution trace data of the completed tasks and calculating the storage path benefit value E path ;

[0047] Monitor the storage path benefit of the consecutive task execution paths of the same user. When the storage path benefit values of 3 consecutive tasks satisfy: E path <0.6, then forcefully switch the calculation mode identifier of the user task to memory-intensive.

[0048] Optionally, the storage path benefit value is calculated as:

[0049] where D local represents the number of local storage accesses, D remote represents the number of cross-node accesses, and τ avg represents the average access latency.

[0050] Advantages of the present invention:

[0051] The present invention adopts the calculation mode identifier parsing technology to identify the calculation type of the task, and combines the data topology fingerprint to construct the coincidence degree of the storage level access path, ensuring that the task is allocated to nodes with a high data localization rate as much as possible. At the same time, combined with the dynamic time constraint factor, it ensures that the time window of task allocation can cover the expected execution time and is adaptively adjusted according to the historical timeout rate. Compared with the traditional static scheduling method, it effectively reduces the deviation of the calculation resource adaptation of the task, reduces the remote data access overhead, and improves the task throughput and execution stability.

[0052] In the present invention, by introducing a dynamic load-adaptive quantization task splitting strategy, the task splitting granularity is adjusted according to the server load. A finer-grained splitting is adopted in a high-load environment to enhance the flexibility of resource scheduling and avoid large tasks from blocking global computing resources. In addition, a probabilistic reverse-dependency barrier insertion strategy is adopted, and synchronization barriers are inserted only when the probability of resource conflicts is relatively high, avoiding the problem of reduced computing throughput caused by excessive insertion of barriers in traditional scheduling methods. Compared with the fixed time slice scheduling method, it can improve the flexibility of task scheduling, increase the computing concurrency in a high-load environment, and reduce the resource conflict rate.

[0053] In the present invention, through the automatic analysis of execution trace data, the benefit value of the storage hierarchy access path is calculated to evaluate the storage access mode of tasks. When the path benefit values of multiple consecutive tasks are relatively low, the system can automatically adjust the computing mode identifier to make tasks more likely to be scheduled to servers with high-bandwidth memory or better data access paths, thereby reducing the overhead of cross-node data access. Brief Description of the Drawings

[0054] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only those of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0055] Figure 1 It is a schematic flowchart of the control method according to an embodiment of the present invention;

[0056] Figure 2 It is a schematic diagram of the screening constraint conditions according to an embodiment of the present invention. Detailed Embodiments

[0057] The following will describe the present invention in detail with reference to the drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; moreover, the drawings are only for more specifically describing the embodiments and are not intended to specifically limit the present invention.

[0058] As Figure 1 - Figure 2 shown, the control method for multi-user concurrent access of a server includes the following steps:

[0059] S1: Real-time analyze the task type and resource requirement characteristics of the user request, and generate a triple including a computing mode identifier, a data topology fingerprint, and a time constraint factor;

[0060] S2: Based on the real-time resource topology status of the server cluster, perform spatio-temporal coupling matching between the triples and the server nodes, calculate the spatio-temporal weight score, and generate a bound execution unit based on the spatio-temporal weight score. The bound execution unit includes a combination of server nodes that meet the screening constraint conditions, and the screening constraint conditions include (filtering candidate nodes in the following priority order):

[0061] Computing pattern matching layer: Calculate the matching between the computing pattern identifier and the server characteristics;

[0062] Data topology adaptation layer: Match the coincidence degree of the storage layer access paths;

[0063] Time window verification layer: The resource reservation time slice corresponding to the time constraint factor covers the expected execution window of the task;

[0064] The screening constraint conditions are used to filter candidate nodes → Determine candidate server nodes;

[0065] Calculating the spatio-temporal weight score is used to obtain a score → Select the optimal server node;

[0066] S3: Implement quantization sharding scheduling on the bound execution unit, split each task into multiple atomic operation sequences, and insert reverse dependency barriers between adjacent atomic operation sequences according to the dynamic load rate of the server nodes. Execution trajectory data is also generated during execution;

[0067] S4: Adjust the computing pattern matching through the execution trajectory feedback mechanism, and remap the computing pattern identifier of the unexecuted task using the path efficiency of the completed tasks.

[0068] S1 specifically includes:

[0069] S11: Based on the maximum method by score, identify the computing pattern identifier M;

[0070] M = argmax(S cpu , S gpu , S mem ), where M is the computing pattern identifier of the task, which is used to represent the main computing characteristics of the task in server resource allocation, and S cpu is the CPU computing intensity score, S gpu is the GPU computing intensity score, and S mem is the memory-intensive score;

[0071] CPU computing intensity score: w cpu is the CPU weight coefficient, taking 1.0;

[0072] GPU computing intensity score: w gpu is the GPU weight coefficient, taking 1.5;

[0073] Memory-intensive score: w mem is the memory weight coefficient, taking 1.2;

[0074] If S cpu is the largest, then M = S cpu , then it is identified as CPU-intensive;

[0075] If S gpu is the largest, then M = S gpu , then it is identified as GPU-intensive;

[0076] If S mem is the largest, then M = S mem , then it is identified as memory-intensive;

[0077] S12: Extract the storage level distribution and data access path of the task-related data, construct a data topology fingerprint, and the data topology fingerprint is measured by the storage level access path overlap degree C:

[0078] Among them, is the storage location identifier of the i-th data block of the task, is the local storage identifier of server node j, δ is the location matching function, and n represents the number of data blocks accessed by the task. is defined as:

[0079]

[0080] S13: Dynamically calculate the time constraint factor, and the time constraint factor is reflected based on the dynamic urgency:

[0081] Parse the absolute deadline T deadline declared by the task, and calculate the dynamic urgency U(t):

[0082] Among them, T resrve is the benchmark reserved time window (taking the average value) statistically based on the execution time of historical similar tasks, t current is the current time, that is, the server system time when the task is parsed, and α(t) is the dynamic adjustment coefficient based on the current resource competition state, and the value range is 0.5 ≤ □ ≤ 2.0.

[0083] In the calculation mode matching layer:

[0084] If the computing mode identifier is GPU-intensive, filter server nodes with CUDA core count ≥ task requirements and PCIe bandwidth ≥ 16 GB / s. The core advantage of GPU computing lies in large-scale parallel computing. If the CUDA core count is insufficient, the computing will be restricted, affecting throughput. Data transfer is required between the CPU and GPU. If the PCIe bandwidth is too low, data transfer will become a bottleneck, resulting in underutilization of GPU resources;

[0085] If it is CPU-intensive, filter server nodes that support the AVX-512 instruction set and L3 cache capacity ≥ 1.2 times the task data volume. CPU computing (such as matrix operations, numerical simulations, etc.) relies on SIMD (Single Instruction Multiple Data). AVX-512 is the most powerful instruction set for CPU vectorization currently, which can significantly improve computing efficiency. CPU computing is sensitive to cache hit rate. Task data should be placed in the L3 cache as much as possible to reduce memory access. If the data volume far exceeds the L3 cache, the CPU needs to access the main memory frequently, resulting in a decline in computing performance;

[0086] If it is memory-intensive, filter server nodes with memory bandwidth utilization within the NUMA node < 60%. Under NUMA (Non-Uniform Memory Access Architecture), cross-NUMA access will increase memory latency and reduce task throughput. Therefore, prefer NUMA nodes with non-overloaded memory bandwidth to ensure that tasks can efficiently utilize local memory bandwidth;

[0087] In the data topology adaptation layer, based on the storage hierarchy access path overlap degree C, only retain server nodes with C ≥ 80% to reduce remote data access overhead.

[0088] In the time window verification layer, the task expected execution window T required Based on the estimated execution time T exec ext Combined with the dynamic urgency U(t) calculation, expressed as:

[0089] T required =T exec est ×max(1,U(t))×(1 + β), where β is the fault tolerance coefficient (default 0.2). Check and filter whether there is a continuous resource reservation time slice [t start ,t end in the node resource reservation table that satisfies t end -t start ≥T required of the server nodes.

[0090] S2 includes calculating the spatio-temporal weight score for the filtered server nodes:

[0091] Among them, Score is the spatio-temporal weight score of the current computing node, α is the data topology weight, α = 0.6, which measures the matching degree between the data and the node storage to ensure that the data is stored locally as much as possible. γ is the time adaptation weight, γ = 0.3, which measures whether the reserved time window is sufficient to prevent the task from being interrupted. λ is the load balancing weight, λ = 0.1, which measures the current load situation of the server to prevent tasks from concentrating on a few high-load nodes. C is the coincidence degree of the storage layer access path, T reserve is the reference reserved time window, that is, the length of the reservable time slice of the current node, T required is the expected execution window of the task, Q pending is the length of the task queue to be processed by the current node at present, Q max is the maximum task queue capacity of the current node;

[0092] Select the node with the highest spatio-temporal weight score to generate a bound execution unit.

[0093] The bound execution unit includes:

[0094] (a) Hardware resource lock: Bind the exclusive resource combination identifier, including:

[0095] GPU resources (GPU card serial numbers, such as GPU-1, GPU-3).

[0096] CPU resources (range of CPU physical core IDs, such as CPUcore4-7).

[0097] Storage resources (NVMe storage ID, such as NVMe-2).

[0098] (b) Spatio-temporal contract: The time and space allocation protocol between the task and the server node, including:

[0099] The maximum execution window allowed for the task: that is, the time constraint for the latest completion of the task (based on the result of the time window verification layer);

[0100] Data access path mapping: including whether data migration is allowed and whether to adopt a local cache acceleration strategy (for server nodes with a data access path coincidence degree C≥80% based on the data topology adaptation layer to ensure that the data is as local as possible).

[0101] S3 specifically includes:

[0102] S31, defining the quantization splitting rule of the atomic operation sequence:

[0103] S311, each atomic operation satisfies the minimum resource exclusivity:

[0104] Occupy the integer / floating-point operation unit of a single physical core;

[0105] Exclusive PCIe channel or exclusive contiguous memory block;

[0106] S312, the splitting granularity G is adaptively adjusted according to the node dynamic load rate η:

[0107] where η is the current CPU / GPU / memory comprehensive load rate of the node;

[0108] S32, construct an insertion strategy for reverse dependency barriers: predict the resource conflict risk between adjacent atomic operation sequences. If the time when the previous operation releases a certain resource is later than the time when the subsequent operation requests the same resource, then mark that the resource has a conflict. Calculate the probability of the conflict occurrence based on historical conflict data and the current load situation. When the conflict probability exceeds the set threshold, insert a reverse dependency barrier between the previous operation and the subsequent operation to ensure that the subsequent operation can only be executed within the protection time interval after the previous operation releases the resource, so as to avoid task anomalies or reduced execution efficiency caused by resource preemption;

[0109] S321, predict the resource conflict risk of adjacent atomic operation sequences:

[0110] If the previous operation O i releases the resource R k at time t release which is later than the time t j when the subsequent operation O k requests R acquire , then mark it as a conflict;

[0111] Calculate the conflict probability:

[0112] S322, when P c > 0.8, insert a barrier between O i and O j to force satisfaction of: where Δt guard is the protection interval, and the default value is 10 μs.

[0113] The execution trace data is generated during task execution by recording the spatio-temporal distribution information of the atomic operation sequence, including the unique identifier of the operation, the server resources occupied, the start time and end time of execution. For tasks with dependency relationships, record the number of the pre-dependency barrier, and at the same time predict the post-conflict.

[0114] Operation ID: The unique number that identifies the current atomic operation sequence;

[0115] Resource lock: Record the server resources exclusive to the task during execution;

[0116] Start time: Record the actual execution start point of this atomic operation;

[0117] End time: Record the actual execution end point of this atomic operation;

[0118] Dependency barrier: Preceding barrier ID: If this operation must wait for a certain barrier to be released before it can be executed, record the ID of this barrier; Post - conflict prediction: Predict operations that may cause resource conflicts subsequently.

[0119] S4 includes performing a storage - level access path benefit evaluation on the execution trace data of the completed tasks, and calculating the storage path benefit value E path ;

[0120] Monitor the storage path benefits of consecutive task execution paths of the same user. When the storage path benefit values of three consecutive tasks satisfy: E path <0.6, then force the calculation mode identifier of this user's task to be switched to memory - intensive, adjust the resource matching strategy of the task, and preferentially schedule it to a server node with higher memory bandwidth and lower data access latency.

[0121] The storage path benefit value is calculated as:

[0122] where D local represents the number of local storage accesses, D remote represents the number of cross - node accesses, and τ avg represents the average access latency.

[0123] The present invention covers any substitutions, modifications, equivalent methods, and solutions made within the spirit and scope of the present invention. To enable the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention. However, those skilled in the art can fully understand the present invention even without the description of these details. Additionally, to avoid unnecessary confusion to the essence of the present invention, well - known methods, processes, flows, components, and circuits are not described in detail.

[0124] The above - mentioned are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for controlling concurrent access of multiple users to a server, characterized in that: The following steps are involved: S1: Analyze the task type and resource requirement characteristics of the user's request in real time, and generate a triple including the computing mode identifier, data topology fingerprint, and time constraint factor; S2: Based on the real-time resource topology state of the server cluster, the triplet is matched with the server node in time and space, the time and space weight score is calculated, and a binding execution unit is generated based on the time and space weight score, wherein the binding execution unit includes a server node combination that satisfies the screening constraint condition, and the screening constraint condition includes: Computational pattern matching layer: Computational pattern identifiers are matched with server characteristics; Data topology adaptation layer: matches the access path overlap of storage layers; Time window verification layer: The resource reservation time slice corresponding to the time constraint factor covers the expected execution window of the task; S3: Implement quantized sharding scheduling for bound execution units, split each task into multiple atomic operation sequences, and insert reverse dependency barriers between adjacent atomic operation sequences based on the dynamic load rate of the server node. Execution trace data is also generated during execution. S4: Adjust the computational pattern matching through the execution trajectory feedback mechanism, and use the path efficiency of completed tasks to remap the computational pattern identifiers of unexecuted tasks.

2. The method for controlling concurrent access of multiple users to a server according to claim 1, characterized in that: The S1 specifically includes: S11: Based on the maximum score method, identify the calculation mode identifier M; M = arg max (S cpu ,S gpu ,S mem ), where M is the computing mode identifier of the task, which is used to represent the main computing characteristics of the task in server resource allocation, and S cpu Calculates the CPU intensity score, S gpu S is the GPU computing intensity score. mem Score for memory intensive; S12: Extract the storage level distribution and data access path of task-related data, and construct a data topology fingerprint. The data topology fingerprint is measured by the overlap C of the storage level access path: in, is the storage location identifier of the i-th data block of the task, is the local storage identifier of server node j, δ is the location matching function, and n represents the number of data blocks accessed by the task; S13: Dynamically calculate the time constraint factor, which is based on the dynamic urgency: The absolute deadline T for parsing the task statement deadline , calculate the dynamic urgency U(t): Among them, T reserve To reserve a time window based on the historical execution time statistics of similar tasks, t current is the current time, and α(t) is the dynamic adjustment coefficient based on the current resource competition status.

3. The method for controlling concurrent access of multiple users to a server according to claim 2, characterized in that: In the calculation pattern matching layer: If the computing mode identifier is GPU intensive, select server nodes with CUDA core number ≥ task requirement and PCIE bandwidth ≥ 16GB / s; If it is CPU-intensive, select server nodes that support the AVX-512 instruction set and whose L3 cache capacity is ≥ 1.2 times the task data volume; If it is memory-intensive, select server nodes with memory bandwidth utilization less than 60% within the NUMA node; In the data topology adaptation layer, based on the storage level access path overlap C, only server nodes with C≥80% are retained.

4. The method for controlling concurrent access of multiple users to a server according to claim 3, characterized in that: In the time window verification layer, the task expected execution window T required Based on the estimated execution time T exec est Combined with the dynamic urgency U(t) calculation, find and filter whether there are continuous resource reservation time slices [t start ,t end ]Satisfy t end -t start ≥T required server node.

5. The method for controlling concurrent access of multiple users to a server according to claim 1, characterized in that: S2 includes calculating the spatiotemporal weight score for the server nodes that pass the screening: Among them, Score is the spatiotemporal weight score of the current computing node, α is the data topology weight, γ is the time adaptation weight, λ is the load balancing weight, C is the storage level access path overlap, T reserve Reserve a time window for the benchmark, T required is the expected execution window of the task, Q pending is the length of the current queue of pending tasks at the current node, Q max is the maximum task queue capacity of the current node; The node with the highest spatiotemporal weight score is selected to generate the bound execution unit.

6. The method for controlling concurrent access of multiple users to a server according to claim 5, characterized in that: The binding execution unit comprises: (a) Hardware resource lock: binds exclusive resource combination identifiers, including: GPU resources; CPU resources; storage resources; (b) Space-time contract: The time and space allocation protocol between tasks and server nodes, including: The maximum execution window allowed for a task: that is, the time constraint for the latest completion of the task; Data access path mapping: including whether data migration is allowed and whether to adopt a local cache acceleration strategy.

7. The method for controlling concurrent access of multiple users to a server according to claim 1, characterized in that: The S3 specifically includes: S31, defines the quantization splitting rules of atomic operation sequences: S311, each atomic operation satisfies the minimum resource exclusivity; S312, the splitting granularity is adaptively adjusted according to the dynamic load rate of the node; S32, construct a reverse dependency barrier insertion strategy: predict the resource conflict risk between adjacent atomic operation sequences. When the time when the previous operation releases a certain resource is later than the time when the subsequent operation applies for the same resource, mark the resource as conflicting. Calculate the probability of conflict based on historical conflict data and current load conditions. When the conflict probability exceeds the set threshold, insert a reverse dependency barrier between the previous operation and the subsequent operation to ensure that the subsequent operation can be executed within the protection time interval after the previous operation releases the resource.

8. The method for controlling concurrent access of multiple users to a server according to claim 1, characterized in that: The execution trajectory data is generated during task execution by recording the spatiotemporal distribution information of the atomic operation sequence, including the unique identifier of the operation, the occupied server resources, the start time and the end time of the execution. For tasks with dependencies, the number of the preceding dependency barrier is recorded, and the subsequent conflicts are predicted.

9. The method for controlling concurrent access of multiple users to a server according to claim 1, characterized in that: S4 includes evaluating the storage level access path benefit of the completed task execution trajectory data and calculating the storage path benefit value E path ; Monitor the execution path benefits of consecutive tasks of the same user. When the storage path benefit values ​​of three consecutive tasks meet the following conditions: path <0.6, the calculation mode identifier of the user task is forced to switch to memory intensive.

10. The method for controlling concurrent access of multiple users to a server according to claim 1, characterized in that: The storage path benefit value is calculated as: Among them, D local Indicates the number of local storage accesses, D remote represents the number of cross-node visits, τ avg Indicates the average access latency.

Citation Information

Patent Citations

  • Task processing method based on thread resources and related device

    CN110018892A

  • Cluster file system client multi-core concurrent load implementation method

    CN118034892A

  • Assigning resources among multiple task groups in a database system

    US20150113540A1

Cited By

  • Computing resource scheduling method based on user demands and task priorities

    CN120353583A

  • Cross-level linkage comprehensive law enforcement business management and decision-making auxiliary method and platform

    CN122243067A