Computing power transmission cooperation technology based on VRB and network slicing
By acquiring computing power demand information and constructing a resource mapping model, the VRB protocol is used to directly allocate physical GPU server resources, solving the problem of low resource allocation efficiency in existing technologies, achieving efficient and stable resource allocation, and improving network support capabilities for cloud computing and artificial intelligence scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-19
- Publication Date
- 2026-03-27
AI Technical Summary
In existing technologies, resource allocation needs to be frequently adjusted according to changes in customer traffic, resulting in low allocation efficiency and an inability to effectively cope with the growth of network resources in high-demand scenarios such as cloud computing and artificial intelligence.
By acquiring computing power demand information, the target network slice is determined, and a resource mapping model is constructed. The VRB protocol is used to directly allocate physical GPU server resources to the target network slice, avoiding dynamic adjustment latency and achieving rapid matching.
It improves the efficiency and stability of resource allocation, reduces transmission latency, enhances the computing network's ability to support cloud computing and artificial intelligence scenarios, and adapts to the rapidly growing demand for network resources.
Smart Images

Figure CN121750474A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a method, apparatus, device, and storage medium for resource allocation in network slicing. Background Technology
[0002] With the rapid development of emerging technologies such as cloud computing and artificial intelligence, computing networks, as the infrastructure supporting these technologies, are increasingly demanding network resources. In existing technologies, resource allocation strategies often need to be frequently adjusted based on changes in customer traffic, which reduces allocation efficiency. Summary of the Invention
[0003] In view of the above problems, embodiments of the present invention are proposed to provide a method, apparatus, device and storage medium for resource allocation of network slices that overcomes or at least partially solves the above problems.
[0004] To address the aforementioned problems, this invention discloses a resource allocation method for network slicing, the method comprising: Obtain computing power demand information; Based on the computing power requirement information, the target network slice is determined; Based on the target network slice, a resource mapping model is constructed, which includes the mapping relationship between different network slices and physical GPU servers; Based on the resource mapping model, resources of the corresponding physical GPU server are allocated to the target network slice according to the VRB protocol.
[0005] Optionally, the computing power requirement information includes at least one of bandwidth requirement, latency requirement, and computing power resource requirement, and determining the target network slice based on the computing power requirement information includes: The target network slice is determined based on at least one of the bandwidth requirements, latency requirements, and computing resource requirements.
[0006] Optionally, constructing a resource mapping model based on the target network slice includes: Based on the target network slice, query the resource pool for a matching target physical GPU server; The resource mapping model is constructed based on the correspondence between the target slice and the target GPU server.
[0007] Optionally, querying the resource pool for a matching target physical GPU server based on the target slice includes: The physical GPU servers in the resource pool whose computing power meets the computing power requirements of the target network slice are identified as candidate GPU servers. The target physical GPU server is determined based on the candidate GPU servers.
[0008] Optionally, the step of allocating the corresponding physical GPU server resources to the target network slice based on the resource mapping model and the VRB protocol includes: Determine the allocation priority of the target network slice; Based on the allocation priority and the resource mapping model, the corresponding physical GPU server resources are allocated to the target network slice according to the VRB protocol.
[0009] Optionally, it also includes: Obtain the resource utilization and latency information of the physical GPU server; Based on the resource utilization and latency information of the physical GPU server, determine whether the resource allocation strategy needs to be adjusted; If the resource allocation strategy needs to be adjusted, a new target physical GPU server is selected from the resource pool to allocate resources to the target network slice.
[0010] Optionally, it also includes: Determine whether congestion has occurred in the target network slice; If congestion occurs in the target network slice, the rate at which resources are allocated to the target network slice is reduced.
[0011] The present invention also discloses a resource allocation device for network slicing, the device comprising: The first acquisition module is used to acquire computing power demand information; The determination module is used to determine the target network slice based on the computing power requirement information; A construction module is used to construct a resource mapping model based on the target network slice, wherein the resource mapping model includes the mapping relationship between different network slices and physical GPU servers; The allocation module is used to allocate the resources of the corresponding physical GPU server to the target network slice according to the resource mapping model and the VRB protocol.
[0012] Optionally, the computing power requirement information includes at least one of bandwidth requirement, latency requirement, and computing power resource requirement, and the determining module includes: The first determining submodule is used to determine the target network slice based on at least one of the bandwidth requirements, latency requirements, and computing resource requirements.
[0013] Optionally, the building module includes: The query submodule is used to query the resource pool for a matching target physical GPU server based on the target network slice; A submodule is constructed to build the resource mapping model based on the correspondence between the target slice and the target GPU server.
[0014] Optionally, the query submodule includes: The first determining unit is used to determine the physical GPU servers in the resource pool whose computing power meets the computing power requirements of the target network slice as candidate GPU servers; The second determining unit is used to determine the target physical GPU server based on the candidate GPU servers.
[0015] Optionally, the allocation module includes: The second determining submodule is used to determine the allocation priority of the target network slice; The allocation submodule is used to allocate the corresponding physical GPU server resources to the target network slice based on the VRB protocol, according to the allocation priority and the resource mapping model.
[0016] Optionally, it also includes: The second acquisition module is used to acquire the resource utilization and latency information of the physical GPU server; The first judgment module is used to determine whether the resource allocation strategy needs to be adjusted based on the resource utilization and latency information of the physical GPU server. The selection module is used to reselect the target physical GPU server in the resource pool to allocate resources to the target network slice if the resource allocation strategy needs to be adjusted.
[0017] Optionally, it also includes: The second judgment module is used to determine whether congestion has occurred in the target network slice; The reduction module is used to reduce the rate at which resources are allocated to the target network slice if congestion occurs in the target network slice.
[0018] The present invention also discloses an electronic device, comprising: a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program, when executed by the processor, implements the steps of the resource allocation method for network slicing as described above.
[0019] The present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the resource allocation method for network slicing as described above.
[0020] The embodiments of the present invention have the following advantages: This invention discloses a method, apparatus, device, and storage medium for resource allocation in network slicing. By acquiring computing power demand information and determining the target network slice, it can pre-establish the mapping relationship between different network slices and physical GPU servers. This allows for rapid matching directly based on the resource mapping model during resource allocation, avoiding latency caused by dynamic adjustments, improving the efficiency and stability of resource allocation, and reducing the overhead caused by frequent strategy adjustments. The resource allocation mechanism based on the VRB protocol reduces the number of data copies, significantly reduces transmission latency, and improves network transmission efficiency. At the same time, through the dynamic mapping capability of network slices, it enhances the computing power network's support capability for high-demand scenarios such as cloud computing and artificial intelligence, thereby better adapting to the rapidly growing demand for network resources from emerging technologies. Attached Figure Description
[0021] Figure 1 This is a flowchart illustrating the steps of a resource allocation method for network slicing provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of communication between GPU servers provided in an embodiment of the present invention; Figure 3 This is a structural block diagram of a network slicing resource allocation device provided in an embodiment of the present invention. Detailed Implementation
[0022] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0023] One of the core concepts of this invention is that by acquiring computing power demand information and determining the target network slice, a mapping relationship between different network slices and physical GPU servers can be pre-established. This allows for rapid matching directly based on the resource mapping model during resource allocation, avoiding delays caused by dynamic adjustments, improving the efficiency and stability of resource allocation, and reducing overhead caused by frequent strategy adjustments. The resource allocation mechanism based on the VRB protocol reduces the number of data copies, significantly reduces transmission latency, and improves network transmission efficiency. At the same time, the dynamic mapping capability of network slices enhances the computing power network's support capability for high-demand scenarios such as cloud computing and artificial intelligence, thereby better adapting to the rapidly growing demand for network resources from emerging technologies.
[0024] Reference Figure 1 The diagram illustrates a flowchart of a resource allocation method for network slices provided by an embodiment of the present invention. The method may include the following steps: Step 101: Obtain computing power requirement information.
[0025] In this embodiment of the invention, computing power requirement information refers to the performance requirements of a business or application for computing resources. This typically includes key indicators such as computing power, storage, network, and real-time performance. It may also include indicators such as the type of computing task, required computing precision, latency sensitivity, duration, peak computing power requirements, and data throughput. This information can be obtained through user-submitted service level agreements, real-time monitoring systems, or historical data analysis. For example, autonomous driving scenarios may require low-latency, high-reliability computing servers, while scientific computing may focus more on high-throughput GPU computing servers. Step 102: Determine the target network slice based on the computing power requirement information.
[0026] This invention can select or create the most matching slice instance from a predefined or dynamically generated network slice pool by analyzing computing power demand information. A network slice is a virtualized logical network whose characteristics can be aligned with computing power demand. That is, different demands are matched with different types of network slices. For example, tasks that require high-performance GPUs will be mapped to slices equipped with the corresponding hardware, while lightweight tasks may be assigned to shared slices to save resources. The decision-making process can be based on a policy engine or algorithm. If the existing slices cannot meet the demand, the orchestration process of new slices can be triggered.
[0027] Step 103: Construct a resource mapping model based on the target network slice.
[0028] In this embodiment of the invention, the resource mapping model is a key component for decoupling virtual slices from physical resources. Essentially, it is a dynamically maintained distributed database entry. Each record in the table is associated with a slice ID and the attributes of the physical GPU server, such as server IP, GPU model (e.g., NVIDIA V100 / A40), memory capacity, current load, rack topology, etc. The mapping relationship can adopt a one-to-one (exclusive GPU), one-to-many (shared GPU), or many-to-one (aggregated resources) model. The construction process can combine real-time resource monitoring data and scheduling strategies, and ensure data synchronization through consistent hashing or distributed locks. The table can also contain failover rules, such as automatically pointing to the backup node when the primary GPU node fails.
[0029] Step 104: Based on the resource mapping model, allocate the corresponding physical GPU server resources to the target network slice according to the VRB protocol.
[0030] In this embodiment of the invention, VRB (V2V RDMA BAND, Remote Direct Memory Access Bandwidth Based on VideoNet Protocol) protocol, V2V refers to the VideoNet Protocol, a network protocol used for video communication; RDMA stands for Remote Direct Memory Access, which allows a computer to directly access the memory of other computers without operating system intervention; BAND stands for bandwidth, and the whole refers to on-chip remote direct memory access bandwidth technology based on the VideoNet Protocol.
[0031] This invention can first parse the resource mapping model to locate the specific resource pool of the target GPU server, and then allocate the physical resources in the GPU to the target network slice through the VRB protocol.
[0032] like Figure 2 The diagram illustrates a communication method between GPU servers provided by the present invention. In the diagram, GPU servers transmit resources based on the VRB transmission channel, which reduces the number of data copies, significantly reduces transmission latency, and improves network transmission efficiency.
[0033] This invention discloses a resource allocation method for network slices. By acquiring computing power demand information and determining the target network slice, a mapping relationship between different network slices and physical GPU servers can be pre-established. This allows for rapid matching directly based on the resource mapping model during resource allocation, avoiding latency caused by dynamic adjustments, improving the efficiency and stability of resource allocation, and reducing overhead caused by frequent strategy adjustments. The resource allocation mechanism based on the VRB protocol reduces the number of data copies, significantly lowers transmission latency, and improves network transmission efficiency. At the same time, the dynamic mapping capability of network slices enhances the support capability of the computing network for high-demand scenarios such as cloud computing and artificial intelligence, thereby better adapting to the rapid growth in network resource demands of emerging technologies.
[0034] In one embodiment of the present invention, the computing power requirement information includes at least one of bandwidth requirement, latency requirement, and computing power resource requirement. Determining the target network slice based on the computing power requirement information includes: determining the target network slice based on at least one of bandwidth requirement, latency requirement, and computing power resource requirement.
[0035] In this embodiment of the invention, in network slice resource allocation, the computing power requirement information covers at least one of bandwidth requirement, latency requirement, and computing power resource requirement. Determining the target network slice based on these requirements is the core operation for achieving accurate resource allocation. Bandwidth requirement reflects the service's capacity to transmit data per unit time. For example, for high-definition video live streaming, to ensure smooth and uninterrupted video playback, high bandwidth is needed to transmit large amounts of video data. Based on the bandwidth configuration of different network slices, slices that can meet the bandwidth requirements of this service are selected as candidate targets. Latency requirement reflects the service's tolerance for data transmission delay. For real-time applications such as autonomous driving and telemedicine... For applications with extremely high performance requirements, extremely low latency is crucial. Network slices with latency parameters that meet the requirements can be prioritized to ensure rapid response to business commands and avoid accidents or operational errors caused by delays. Computing resource requirements involve the computing power required for business operations. For example, deep learning training tasks require powerful computing resources to process massive amounts of data and complex algorithms. The computing resources such as CPU and GPU provided by each network slice can be comprehensively evaluated, and network slices that match the computing resource requirements of the business can be matched. By comprehensively considering and screening one or more of these three requirements, the most suitable target network slice for the current business can be determined, achieving a precise match between network resources and business needs.
[0036] In one example, network slices corresponding to services can be divided into high-bandwidth slices, medium-bandwidth slices, or low-bandwidth slices based on bandwidth requirements. High-bandwidth slices have bandwidth requirements greater than 50Gbps, medium-bandwidth slices have bandwidth requirements between 10Gbps and 50Gbps, and low-bandwidth slices have bandwidth requirements less than 10Gbps.
[0037] In another example, the network slices corresponding to the service are divided into low-latency slices, medium-latency slices, or high-latency slices according to latency requirements. The latency requirement for low-latency slices is less than 10 milliseconds, the latency requirement for medium-latency slices is between 10 and 100 milliseconds, and the latency requirement for high-latency slices is greater than 100 milliseconds.
[0038] In another example, the network slices corresponding to the service can be divided into high-computing-power slices, low-computing-power slices, and medium-computing-power slices according to the computing power resource requirements. For example, the computing power resources of high-computing-power slices are greater than 100 TFLOPS, the computing power resources of medium-computing-power slices are 50~100 TFLOPS, and the computing power resources of low-computing-power slices are less than 50 TFLOPS.
[0039] In another example, the network slices corresponding to the service can be divided into real-time slices or high real-time slices according to the real-time requirements, wherein the latency requirement of real-time slices is less than 5 microseconds and the latency requirement of high real-time slices is less than 1 microsecond.
[0040] This invention determines the target network slice based on computing power demand information, avoiding over-allocation or under-allocation of resources. This method of dynamically determining the target network slice according to computing power demand enables the network to quickly adapt to changes in different services. When new services are launched or existing service requirements are adjusted, the system can quickly reassess computing power requirements and select appropriate network slices, enhancing the network's flexibility and scalability.
[0041] In one embodiment of the present invention, constructing a resource mapping model based on a target network slice includes: querying a matching target physical GPU server in a resource pool based on the target network slice; and constructing a resource mapping model based on the correspondence between the target slice and the target GPU server.
[0042] In this embodiment of the invention, firstly, based on computing power requirement information (such as bandwidth, latency, and computing resource requirements), a matching target physical GPU server can be directly queried from the resource pool. The query process here is based on a precise comparison between the real-time computing power parameters of the servers in the resource pool, such as the number of GPU cores, memory capacity, and floating-point operation capability, and the requirements of the target slice. For example, if the target network slice needs to process high-resolution image rendering, physical GPU servers with memory bandwidth ≥ 400 GB / s and floating-point operation capability ≥ 10 TFLOPS can be searched in the resource pool. At the same time, combined with the current load status of the server (such as CPU / GPU utilization < 50%) and network connection quality (such as latency with the node where the slice is located < 1 ms), the most suitable physical GPU server can be directly located.
[0043] After identifying the target server, the correspondence between the target network slice's identifier, requirement parameters, and the target physical GPU server's device ID, hardware configuration, location information, etc., can be recorded in a structured table as a resource mapping model. This provides a direct index for subsequent resource allocation. For example, a correspondence can be established between high-computing-power slices and the adapted physical GPU server 1 and stored in the resource mapping model. This invention can directly match the target physical GPU server through computing power parameters and real-time status, reducing intermediate decision-making processes and making the construction of the resource mapping model more efficient.
[0044] In one embodiment of the present invention, querying a matching target physical GPU server in a resource pool based on a target slice includes: determining physical GPU servers in the resource pool whose computing power meets the computing power requirements of the target network slice as candidate GPU servers; and determining the target physical GPU server based on the candidate GPU servers.
[0045] In this embodiment of the invention, the computing power requirements of the target network slice can be analyzed, including but not limited to computing power (such as floating-point operations per second, FLOPS), memory capacity requirements, parallel processing capabilities (such as the number of CUDA cores), and support for specific computing instruction sets (such as TensorCore acceleration). At the same time, the hardware parameters and operating status data of each physical GPU server in the resource pool can be collected in real time, including GPU model (such as NVIDIA A100, V100), memory size, current load rate, temperature, power status, etc. By establishing a multi-dimensional parameter matching model, the computing power of the GPU server and the slice requirements are quantitatively compared, and all servers that meet the basic requirements are selected as candidate GPU servers. For example, for a network slice that requires large-scale deep learning model training, GPU servers with ≥32GB of memory, TensorCore performance ≥100TFLOPS, and current load rate <60% can be selected as candidate GPU servers.
[0046] This invention comprehensively considers multiple optimization objectives to ensure that the final selected server not only meets functional requirements but also achieves high efficiency in overall resource utilization. First, it assesses the resource availability of candidate servers, including remaining computing resources, network bandwidth, and storage I / O capabilities, prioritizing servers with sufficient resource reserves to handle peak business periods. Second, it considers the server's geographical location and network topology, selecting servers with the lowest network latency (e.g., <1ms) and highest link stability (e.g., packet loss rate <0.01%) to reduce data transmission overhead. Furthermore, it can evaluate server reliability by combining historical operating data, including hardware failure rate and system stability, prioritizing servers with redundant designs for critical business operations. Finally, it uses a comprehensive scoring model, such as weighted computing resource utilization, network latency, and reliability indicators, to centrally determine the highest-scoring GPU server from the candidate GPU servers as the target physical GPU server.
[0047] In one application scenario, the business requirement is to support real-time analysis of 4K video at 1000+ frames per second, with target detection accuracy ≥99% and end-to-end processing latency <100ms. First, based on the computing power requirements of this target network slice, candidate GPU servers are screened from the resource pool. Through parameter comparison, three servers in the resource pool are found to meet the basic conditions: Server A is an NVIDIA A100 80GB, currently at 30% load, located in a city data center; Server B is an NVIDIA T4 16GB, currently at 45% load, located in an edge data center; Server C is an NVIDIA V100 32GB, currently at 20% load, located in a regional cloud center.
[0048] Subsequently, the target server determination phase began. Considering the extreme sensitivity of vehicle-road cooperative services to low latency, the network latency between each candidate server and the RSU was first evaluated: Server A, connected via a dedicated 5G slice, had a latency of approximately 15ms; Server B, directly connected via local fiber optic, had a latency of approximately 8ms; Server C required cross-regional transmission, with a latency of approximately 40ms. Simultaneously, analysis of the service's computational characteristics revealed that video analytics tasks require extensive parallel computation but have moderate GPU memory requirements (approximately 10GB). Therefore, Server B's 16GB GPU memory was more economical while meeting the requirements. Furthermore, the edge data center where Server B is located employs a dual-power redundant design with a historical failure rate of less than 0.01%, ensuring reliability that meets service requirements. Taking all these factors into account, Server B (NVIDIA T4 16GB, edge data center) was ultimately determined as the target physical GPU server, and a resource mapping model was constructed to record this correspondence.
[0049] This invention is based on dynamic evaluation of real-time load and resource availability to ensure that new network slice requirements are preferentially allocated to low-load, highly available servers, thereby achieving load balancing within the resource pool. It can comprehensively consider multiple optimization objectives to ensure that the final selected server can not only meet functional requirements but also achieve high efficiency in overall resource utilization.
[0050] In one embodiment of the present invention, allocating resources of a corresponding physical GPU server to a target network slice based on the VRB protocol according to the resource mapping model includes: determining the allocation priority of the target network slice; and allocating resources of a corresponding physical GPU server to the target network slice based on the VRB protocol according to the allocation priority and the resource mapping model.
[0051] In this embodiment of the invention, determining the allocation priority of target network slices is a crucial preliminary step in resource allocation. This process requires comprehensive analysis of multiple factors. First, the service type is a core consideration: services with extremely high real-time requirements are assigned the highest priority to ensure their resource needs are met first; services with low latency sensitivity but high computational demands are assigned a medium priority; and general services are assigned a basic priority. Second, the real-time status of the service can also be considered: for network slices during peak service periods, their priority can be dynamically increased; for slices with resource utilization consistently below a threshold (e.g., <30%), their priority can be appropriately reduced to free up resources for more urgently needed services. By integrating the above factors, a dynamic priority value can be calculated for each target network slice, serving as the core basis for subsequent resource allocation.
[0052] Furthermore, the physical GPU server corresponding to the target network slice can be located according to the resource mapping model. In one example, the target network slices are slice A, slice B, and slice C. The physical GPU server corresponding to slice A is GPU2, the physical GPU server corresponding to slice B is GPU4, and the physical GPU server corresponding to slice C is GPU1. Slice A has a higher priority than slice C, and slice C has a higher priority than slice B. When allocating resources, the resources of GPU2 can be allocated to slice A first, then the resources of GPU1 can be allocated to slice C, and finally the resources of GPU4 can be allocated to slice B.
[0053] This invention determines the allocation priority of target network slices and sorts the resource allocation weights of different slices according to business needs, ensuring that high-priority slices get physical GPU server resources first, thereby guaranteeing the service quality of critical businesses.
[0054] In one embodiment of the present invention, the method further includes: obtaining the resource utilization and latency information of the physical GPU server; determining whether the resource allocation strategy needs to be adjusted based on the resource utilization and latency information of the physical GPU server; if the resource allocation strategy needs to be adjusted, reselecting the target physical GPU server in the resource pool to allocate resources to the target network slice.
[0055] In this embodiment of the invention, the operating status data of the physical GPU server can be collected in real time through a multi-dimensional monitoring mechanism. For resource utilization, GPU core utilization, such as the real-time utilization of CUDA cores, video memory utilization, memory bandwidth utilization, and power consumption (reflecting the overall load of the server) can be collected. For latency information, GPU computing latency, data transmission latency, and I / O latency (data read and write latency with storage devices) can be monitored.
[0056] Furthermore, a comprehensive evaluation can be conducted based on preset dynamic thresholds and machine learning models. First, multiple thresholds can be set for resource utilization. For example, GPU core utilization exceeding 85% is considered "high load," and exceeding 95% is considered "overload." Memory utilization exceeding 90% is considered "stressful," and exceeding 98% is considered "critical." Second, business-specific thresholds can be set for latency information. For example, for real-time rendering, a computation latency exceeding 50ms and a data transmission latency exceeding 20ms trigger adjustment conditions. The thresholds for batch computing can be relaxed to 200ms and 100ms, respectively. At the same time, historical data analysis and prediction models can be used to determine whether the current state is a short-term peak or a long-term trend. For example, if GPU core utilization exceeds 90% for 3 consecutive minutes and is predicted to show no downward trend in the next 10 minutes, it is determined that adjustment is needed. If it is only a momentary spike (duration < 30 seconds) and quickly falls back to the normal range, it is determined that no adjustment is needed.
[0057] Furthermore, a new round of screening can be initiated in the resource pool based on new constraints. First, based on the original computing power requirements and current priority of the target network slice, all candidate servers that meet the basic conditions (such as GPU model ≥ target model, video memory ≥ 120% of the requirement) are screened from the resource pool. Then, the candidate servers are evaluated in multiple dimensions, calculating resource utilization score (the lower the resource utilization, the higher the score; for example, a server with 50% GPU core utilization scores 90 points, while one with 80% scores 60 points), latency score (the shorter the latency, the higher the score; for example, a server with 10ms latency scores 95 points, while one with 50ms latency scores 70 points), and network topology score (the fewer the network hops and the higher the bandwidth to the target network slice access point, the higher the score). Finally, a comprehensive score is obtained through weighted calculation (e.g., resource utilization weight 0.4, latency weight 0.4, network topology weight 0.2). The server with the highest score is selected as the new target physical GPU server. After the new GPU server is determined, resource migration can be initiated through the VRB protocol: first, an equal amount of resources are reserved on the new GPU server, and then the computing tasks of the target network slice are gradually migrated to ensure uninterrupted service. After the migration is completed, the resource mapping model is updated to record the correspondence between the target network slice and the new server, while releasing the resources on the original server.
[0058] This invention can detect server performance bottlenecks in advance by monitoring resource utilization and latency information in real time, avoiding task failures or service quality degradation caused by resource overload. The dynamic adjustment mechanism greatly improves the overall utilization efficiency of the resource pool: by migrating slices on high-load servers to low-load servers, load balancing within the resource pool is achieved, thereby improving the utilization rate of cluster resources.
[0059] In one embodiment of the present invention, the method further includes: determining whether congestion occurs in the target network slice; if congestion occurs in the target network slice, reducing the rate at which resources are allocated to the target network slice.
[0060] In this embodiment of the invention, determining whether congestion has occurred in a target network slice can involve real-time collection and analysis of multiple network status indicators. These indicators include, but are not limited to, the bandwidth utilization of the network slice, end-to-end latency, packet loss rate, and queue length. These indicators can be compared with preset congestion thresholds. For example, when the bandwidth utilization consistently exceeds 80% and the end-to-end latency increases by more than 50% compared to the baseline value, or the packet loss rate exceeds 1%, congestion can be determined. To avoid false positives, a sliding window mechanism can be used for trend analysis. Congestion is only confirmed when at least one of the multiple indicators exceeds the threshold and persists for a certain period of time. This multi-dimensional, dynamic threshold judgment mechanism can effectively distinguish between temporary network fluctuations and real congestion states.
[0061] Furthermore, a tiered rate adjustment strategy can be adopted based on the severity of congestion and the type of service. First, the physical GPU server corresponding to the target network slice and its currently allocated resource parameters are located through the resource mapping model. Then, a resource adjustment request is sent based on the VRB protocol. This request can carry new resource allocation rate parameters. For services with extremely high real-time requirements, the transmission rate of critical data can be prioritized, while only reducing the resource allocation rate of non-critical business data. For example, the bandwidth share of non-critical data can be reduced from 30% to 10%, while maintaining the bandwidth of critical control signals unchanged. For services that can tolerate a certain amount of latency, the overall resource allocation rate can be reduced proportionally. For example, the GPU computing power share can be reduced from 100% to 60%, while the data transmission bandwidth can be adjusted to 50% of the original rate. During the rate adjustment process, the resource negotiation mechanism of the VRB protocol can be used to interact with the physical GPU server to ensure that the adjusted resource allocation parameters take effect on the server side and update the real-time status in the resource mapping model. The dynamic rate adjustment mechanism effectively alleviates network overload pressure and improves the stability of the network environment.
[0062] Taking the vehicle-road cooperative system in intelligent transportation as an example, multiple autonomous vehicles in a certain area simultaneously transmit real-time traffic data and control commands through 5G network slices. The target network slice carries key data transmissions such as vehicle location information, obstacle detection results, and braking commands. The network status of the slice can be monitored in real time. When a sudden traffic flow peak occurs during a certain period, multiple vehicles enter the intersection at the same time, causing the bandwidth utilization of the network slice to surge from 40% to 85% of the normal state, the end-to-end latency to rise from 20ms to 45ms (exceeding 50% of the baseline value), and the packet loss rate to rise from 0.1% to 1.2%. The system determines that the target slice is experiencing mild congestion.
[0063] At this point, a congestion control mechanism can be immediately activated. The resource mapping model locates the edge GPU server cluster corresponding to the slice, and a tiered rate adjustment is performed based on the VRB protocol. For critical control signals such as braking commands, the original bandwidth allocation and GPU computing power priority are maintained to ensure the real-time performance of vehicle braking response. For periodically updated data such as vehicle location information, its transmission bandwidth is reduced from 30% to 15% of the total bandwidth, and the GPU computing power share is adjusted from 40% to 25% to reduce the processing pressure on non-critical data. For redundant image data from obstacle detection, transmission is temporarily stopped, retaining only the transmission of key feature parameters. After the adjustment, the bandwidth utilization of the target network slice drops to 75%, the end-to-end latency falls back to 30ms, and the packet loss rate drops to 0.5%, effectively alleviating the congestion.
[0064] This invention monitors congestion status in real time using multi-dimensional indicators, enabling early warning and activation of preventative mechanisms in the early stages of congestion to avoid service interruptions caused by severe congestion. It also ensures differentiated service quality for different types of services through a tiered rate adjustment strategy.
[0065] This invention discloses a resource allocation method for network slices. By acquiring computing power demand information and determining the target network slice, a mapping relationship between different network slices and physical GPU servers can be pre-established. This allows for rapid matching directly based on the resource mapping model during resource allocation, avoiding latency caused by dynamic adjustments, improving the efficiency and stability of resource allocation, and reducing overhead caused by frequent strategy adjustments. The resource allocation mechanism based on the VRB protocol reduces the number of data copies, significantly lowers transmission latency, and improves network transmission efficiency. At the same time, the dynamic mapping capability of network slices enhances the support capability of the computing network for high-demand scenarios such as cloud computing and artificial intelligence, thereby better adapting to the rapid growth in network resource demands of emerging technologies.
[0066] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.
[0067] Reference Figure 3 This illustration shows a network slicing resource allocation device provided by an embodiment of the present invention, which may include the following modules: The first acquisition module 201 is used to acquire computing power demand information; The determination module 202 is used to determine the target network slice based on the computing power requirement information; Module 203 is used to build a resource mapping model based on the target network slice. The resource mapping model includes the mapping relationship between different network slices and physical GPU servers. The allocation module 204 is used to allocate the resources of the corresponding physical GPU server to the target network slice according to the resource mapping model and the VRB protocol.
[0068] This invention discloses a resource allocation device for network slicing. By acquiring computing power demand information and determining the target network slice, it can pre-establish the mapping relationship between different network slices and physical GPU servers. This allows for rapid matching directly based on the resource mapping model during resource allocation, avoiding the latency caused by dynamic adjustments, improving the efficiency and stability of resource allocation, and reducing the overhead caused by frequent strategy adjustments. The resource allocation mechanism based on the VRB protocol reduces the number of data copies, significantly reduces transmission latency, and improves network transmission efficiency. At the same time, through the dynamic mapping capability of network slices, it enhances the computing power network's support capability for high-demand scenarios such as cloud computing and artificial intelligence, thereby better adapting to the rapidly growing demand for network resources from emerging technologies.
[0069] In one embodiment of the present invention, the computing power requirement information includes at least one of bandwidth requirement, latency requirement, and computing power resource requirement; the determining module includes: The first determination submodule is used to determine the target network slice based on at least one of bandwidth requirements, latency requirements, and computing resource requirements.
[0070] In one embodiment of the present invention, the construction module includes: The query submodule is used to query the resource pool for matching target physical GPU servers based on the target network slice; The construction submodule is used to build a resource mapping model based on the correspondence between target slices and target GPU servers.
[0071] In one embodiment of the present invention, the query submodule includes: The first determining unit is used to determine the physical GPU servers in the resource pool whose computing power meets the computing power requirements of the target network slice as candidate GPU servers. The second determining unit is used to determine the target physical GPU server based on the candidate GPU servers.
[0072] In one embodiment of the present invention, the allocation module includes: The second determining submodule is used to determine the allocation priority of the target network slice; The allocation submodule is used to allocate the corresponding physical GPU server resources to the target network slice based on the allocation priority and resource mapping model and the VRB protocol.
[0073] In one embodiment of the present invention, it further includes: The second acquisition module is used to acquire the resource utilization and latency information of the physical GPU server; The first judgment module is used to determine whether the resource allocation strategy needs to be adjusted based on the resource utilization and latency information of the physical GPU server. The selection module is used to reselect the target physical GPU server in the resource pool to allocate resources to the target network slice if the resource allocation strategy needs to be adjusted.
[0074] In one embodiment of the present invention, it further includes: The second judgment module is used to determine whether congestion has occurred in the target network slice; The reduction module is used to reduce the rate at which resources are allocated to the target network slice if congestion occurs in the target network slice.
[0075] This invention discloses a resource allocation device for network slicing. By acquiring computing power demand information and determining the target network slice, it can pre-establish the mapping relationship between different network slices and physical GPU servers. This allows for rapid matching directly based on the resource mapping model during resource allocation, avoiding the latency caused by dynamic adjustments, improving the efficiency and stability of resource allocation, and reducing the overhead caused by frequent strategy adjustments. The resource allocation mechanism based on the VRB protocol reduces the number of data copies, significantly reduces transmission latency, and improves network transmission efficiency. At the same time, through the dynamic mapping capability of network slices, it enhances the computing power network's support capability for high-demand scenarios such as cloud computing and artificial intelligence, thereby better adapting to the rapidly growing demand for network resources from emerging technologies.
[0076] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.
[0077] This invention also provides an electronic device, comprising: It includes a processor, a memory, and a computer program stored in the memory and capable of running on the processor. When the computer program is executed by the processor, it implements the various processes of the above-described network slicing resource allocation method embodiment and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0078] This invention also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described network slicing resource allocation method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0079] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0080] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, embodiments of the present invention can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0081] Embodiments of the present invention are described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0082] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0083] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0084] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.
[0085] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes the element.
[0086] The resource allocation method, apparatus, device, and storage medium for network slicing provided by the present invention have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A resource allocation method for network slicing, characterized in that, The method includes: Obtain computing power demand information; Based on the computing power requirement information, the target network slice is determined; Based on the target network slice, a resource mapping model is constructed, which includes the mapping relationship between different network slices and physical GPU servers; Based on the resource mapping model, resources of the corresponding physical GPU server are allocated to the target network slice according to the VRB protocol.
2. The resource allocation method according to claim 1, characterized in that, The computing power requirement information includes at least one of bandwidth requirement, latency requirement, and computing resource requirement. Determining the target network slice based on the computing power requirement information includes: The target network slice is determined based on at least one of the bandwidth requirements, latency requirements, and computing resource requirements.
3. The resource allocation method according to claim 1, characterized in that, The step of constructing a resource mapping model based on the target network slice includes: Based on the target network slice, query the resource pool for a matching target physical GPU server; The resource mapping model is constructed based on the correspondence between the target slice and the target GPU server.
4. The resource allocation method according to claim 3, characterized in that, The step of querying the resource pool for a matching target physical GPU server based on the target slice includes: The physical GPU servers in the resource pool whose computing power meets the computing power requirements of the target network slice are identified as candidate GPU servers. The target physical GPU server is determined based on the candidate GPU servers.
5. The resource allocation method according to claim 1, characterized in that, The step of allocating resources of the corresponding physical GPU server to the target network slice according to the resource mapping model and the VRB protocol includes: Determine the allocation priority of the target network slice; Based on the allocation priority and the resource mapping model, the corresponding physical GPU server resources are allocated to the target network slice according to the VRB protocol.
6. The resource allocation method according to claim 1, characterized in that, Also includes: Obtain the resource utilization and latency information of the physical GPU server; Based on the resource utilization and latency information of the physical GPU server, determine whether the resource allocation strategy needs to be adjusted; If the resource allocation strategy needs to be adjusted, a new target physical GPU server is selected from the resource pool to allocate resources to the target network slice.
7. The resource allocation method according to claim 1, characterized in that, Also includes: Determine whether congestion has occurred in the target network slice; If congestion occurs in the target network slice, the rate at which resources are allocated to the target network slice is reduced.
8. A resource allocation device for network slicing, characterized in that, The device includes: The acquisition module is used to obtain computing power demand information; The determination module is used to determine the target network slice based on the computing power requirement information; A construction module is used to construct a resource mapping model based on the target network slice, wherein the resource mapping model includes the mapping relationship between different network slices and physical GPU servers; The allocation module is used to allocate the resources of the corresponding physical GPU server to the target network slice according to the resource mapping model and the VRB protocol.
9. An electronic device, characterized in that, include: A processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program, when executed by the processor, implements the steps of the resource allocation method for network slicing as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the steps of the resource allocation method for network slicing as described in any one of claims 1-7.