A hierarchical resource orchestration method and system for intent guidance
Patent Information
- Application Number
- CN202611349186.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-09-02
- Publication Date
- 2026-09-29
AI Technical Summary
当两者共享同一逻辑带宽池与边缘算力池时,若仅采用静态切片或面向单一目标的调度策略,则容易导致服务质量恶化
[0020]本领域技术人员将会理解的是,能够用本发明实现的目的和优点不限于以上具体所述,并且根据以下详细说明将更清楚地理解本发明能够实现的上述和其他目的。
Smart Images

Figure CN122838115A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of resource scheduling technology, and in particular to an intention-guided hierarchical resource orchestration method and system. Background Technology
[0002] With the application of extended reality (XR) and large language model technologies, interactive services under edge computing architectures place higher demands on data transmission. These services typically involve multiple stages, including uploading perceptual data, model inference and decision-making, and delivering rendered content. In practical applications, the system needs to simultaneously ensure real-time processing of uplink semantic / contextual data and low-latency transmission of downlink rendering streams to maintain the smoothness of closed-loop interaction. However, due to the different sensitivities of uplink and downlink data streams to bandwidth and latency, balancing bidirectional service flows with limited edge computing resources has become a key technical challenge for improving the interactive experience.
[0003] Most existing technologies design XR rendering optimization and large-model inference optimization in isolation, lacking unified modeling and collaborative scheduling capabilities for the coupled competitive relationship between communication and computing resources. Specifically, on the one hand, XR services are typically highly sensitive to motion-to-photon (MTP) latency, requiring rendering, encoding, transmission, and display processes to be completed within a short time. On the other hand, large-model inference services rely on the collection of uplink multimodal contexts and edge-side computation execution, exhibiting significant bursty computational demands and strong uplink bandwidth consumption characteristics. When both share the same logical bandwidth pool and edge computing power pool, using only static slicing or single-target-oriented scheduling strategies can easily lead to service quality degradation. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide an intent-guided hierarchical resource orchestration method and system to eliminate or improve one or more defects existing in the prior art.
[0005] One aspect of the present invention provides an intent-guided hierarchical resource orchestration method, the steps of which include: Acquire multimodal data collected by the terminal, construct a first sensor sequence based on the multimodal data, and input the first sensor sequence into a lightweight student model to obtain low-dimensional semantic labels; At the start of the current macrocycle, the running statistics of the previous macrocycle are aggregated. A global context vector is constructed based on the running statistics and low-dimensional semantic labels. The global context vector is then input into a pre-trained lightweight intent prediction network to obtain a multi-dimensional intent vector. The multidimensional intent vector is distributed to a distributed decision entity, so that the distributed decision entity inputs its local observation state and the multidimensional intent vector into a multi-agent reinforcement learning policy network to obtain a joint resource scheduling policy. The network transmission resources, logical computing resources, and session layer service parameters are adjusted according to the joint resource scheduling strategy, and the task execution results are collected.
[0006] The above scheme employs a lightweight student model for semantic compression perception, extracting data into low-dimensional semantic labels, thus reducing data complexity. While preserving the semantics of the business scenario, it reduces the amount of data transmitted and processed subsequently, thereby reducing the impact of uplink semantic or contextual data on communication bandwidth and edge computing power. The low-dimensional semantic labels are combined with the operational statistics of the previous macrocycle to form a global context vector, which can uniformly represent the current business scenario, business operation status, and the usage status of shared communication and computing resources. During the online phase, the lightweight intent prediction network is used to determine the scheduling focus, resource consumption constraints, and resource adjustment range for different services. Each distributed decision entity further combines the multi-dimensional intent vector and local observation status to generate a joint resource scheduling strategy for network transmission resources, logical computing resources, and session layer service parameters. Therefore, this solution can determine the resource orchestration intention based on the business scenario and historical operating status within the slow-scale macro cycle, and coordinately adjust communication and computing resources based on the real-time status within the fast-scale control time slot. This enables resource scheduling to adapt to the low MTP latency requirements of XR services and the sudden uplink bandwidth and computing requirements of large language model inference services, thereby alleviating resource competition when the two types of services share the logical bandwidth pool and edge computing power pool, and reducing the risk of service quality degradation caused by static slicing or scheduling for a single objective.
[0007] In some embodiments of the present invention, the method further includes the following steps: determining XR session quality, large model service quality, total resource consumption estimate, and latency constraint default data based on the task execution result of a new macro cycle; combining the XR session quality, the large model service quality, and the total resource consumption estimate based on the XR service weight, the large model service weight, and the resource consumption penalty coefficient in the multidimensional intent vector to obtain an instantaneous reward for intent guidance; statistically calculating the XR service latency default rate and the large model inference service latency default rate within a sliding window based on the latency constraint default data; updating the corresponding constraint penalty multiplier based on the difference between each latency default rate and the corresponding allowed default probability threshold; correcting the instantaneous reward for intent guidance using the updated constraint penalty multiplier and the latency default result of the current fast-scale control slot to obtain a reward with a constraint penalty term; and updating the multi-agent reinforcement learning policy network based on the reward with the constraint penalty term.
[0008] In some embodiments of the present invention, the lightweight student model includes a temporal convolutional network and a gated recurrent unit. In the training step of the lightweight student model, a first sensor sequence and a second sensor sequence are constructed based on the multimodal data in the training data. The first sensor sequence is input into the lightweight student model, and the second sensor sequence is input into the teacher model. The lightweight student model is trained based on the output results of the lightweight student model and the teacher model, as well as the real labels in the training data.
[0009] In some embodiments of the present invention, in the step of aggregating the operational statistics of the previous macrocycle at the start of the current macrocycle and constructing a global context vector based on the operational statistics and low-dimensional semantic labels, Extract motion-to-photon latency, XR session quality, large model inference latency, GPU utilization, and bandwidth resource utilization from the runtime statistics to obtain the runtime data set; The low-dimensional semantic tags are encoded to obtain tag codes; The global context vector is obtained by combining the running data group with the tag encoding.
[0010] In some embodiments of the present invention, in the step of inputting the global context vector into a pre-trained lightweight intent prediction network to obtain a multi-dimensional intent vector, the lightweight intent prediction network includes an input layer, a first hidden layer, a second hidden layer and an output layer including a weight head and a boundary head arranged in sequence, wherein the first hidden layer and the second hidden layer both adopt the ReLU activation function. The global context vector is input into the input layer. The weight header is used to output the scheduling weights corresponding to XR services and large model inference services. The boundary header is used to output the upper limit of bandwidth resource utilization, the upper limit of GPU utilization, and the adjustment range of various resource scheduling actions. The outputs of the weight header and the boundary header are combined to obtain a multi-dimensional intent vector.
[0011] In some embodiments of the present invention, the local observation state includes the logical link quality observation matrix of the current time slot, downlink XR data queue state, uplink control and semantic information queue state, GPU utilization, XR rendering queue state, large model inference queue state, user-side estimated instantaneous latency, user-side experience quality estimate, session buffer state, and session-related semantic state; the distributed decision entity includes a wireless resource decision entity, a computing resource decision entity, and a session execution decision entity, all of which employ reinforcement learning models; the joint resource scheduling strategy includes logical resource block allocation actions, logical transmit power control actions, logical computing power allocation actions, task offloading control actions, encoding or rendering quality level adjustment actions, and service parameter adjustment actions. In the step of distributing the multidimensional intent vector to a distributed decision entity, so that the distributed decision entity inputs its local observation state and the multidimensional intent vector into a multi-agent reinforcement learning policy network to obtain a joint resource scheduling policy: The wireless resource decision entity uses the logical link quality observation matrix of the current time slot, the downlink XR data queue status, the uplink control and semantic information queue status, and the multidimensional intent vector as the first observation state. Based on the first observation state, the wireless resource decision entity outputs logical resource block allocation action and logical transmit power control action. The computing resource decision-making entity uses GPU utilization, XR rendering queue status, large model inference queue status and the multidimensional intent vector as the second observation status. The computing resource decision-making entity outputs logical computing power allocation actions and task unloading control actions based on the second observation status. The session execution decision entity uses the instantaneous latency estimated by the user side, the estimated user experience quality, the session buffer state, the session-related semantic state, and the multidimensional intent vector as the third observation state. The session execution decision entity outputs session execution actions based on the third observation state. The session execution actions include encoding quality level adjustment actions or rendering quality level adjustment actions, as well as business parameter adjustment actions.
[0012] In some embodiments of the present invention, in the step of adjusting network transmission resources, logical computing resources, and session layer service parameters according to the joint resource scheduling strategy... The logical resource block allocation action and the logical transmit power control action correspond to network transmission resources. Based on the logical resource block allocation action, logical resource blocks are allocated from the logical bandwidth pool for the downlink XR data queue and the uplink control and semantic information queue. Based on the logical transmit power control action, the logical transmit power of the corresponding logical link is set. The encoding or rendering quality level adjustment action and the service parameter adjustment action correspond to the session layer service parameters. The level determined based on the encoding or rendering quality level adjustment action is set as the encoding quality level or rendering quality level of the next XR frame. The service parameter adjustment action adjusts the service ratio between the local and edge layers. The service ratio between the local and edge layers can be the task execution ratio, data processing ratio, or rendering ratio. The logical computing power allocation action and the task unloading control action correspond to logical computing resources. Based on the logical computing power allocation action, corresponding computing power is allocated from the logical computing power pool to the XR rendering task and the LAM (Large AI Model) inference task, respectively. Based on the task unloading control action, it is determined whether the task is retained for execution in the current computing domain or transferred to other computing domains for execution.
[0013] In some embodiments of the present invention, in the step of determining delay constraint violation data and resource utilization data based on the task execution results, and updating the multi-agent reinforcement learning policy based on the delay constraint violation data and resource utilization data: Based on the operation of a new macro cycle, a multi-dimensional intent vector is determined, and from the multi-dimensional intent vector, the scheduling weight, bandwidth resource utilization limit, and GPU utilization limit corresponding to each service are determined. Resource utilization data is determined based on the value by which the bandwidth resource utilization exceeds the upper limit of the bandwidth resource utilization and the value by which the GPU utilization exceeds the upper limit of the GPU utilization. Based on the scheduling weights corresponding to each service, the delay constraint default data and resource utilization over-limit data are combined to obtain a feedback reward value; Based on the local observation state before the execution of the joint resource scheduling strategy, the multi-dimensional intent vector, the joint resource scheduling strategy, the feedback reward value, and the local observation state after execution, the network parameters of the multi-agent reinforcement learning strategy network are updated, and the updated multi-agent reinforcement learning strategy network is used to generate subsequent joint resource scheduling strategies.
[0014] Specifically, the estimated value of logical transmission resource consumption is determined based on the number of logical resource blocks allocated to each service in the current time slot, the duration of resource block occupation, and the transmission power; the estimated value of logical computing resource consumption is determined based on the logical computing power allocated to each service in the current time slot and its occupation duration.
[0015] In some embodiments of the present invention, in the step of combining the XR session quality, the large model service quality, and the total resource consumption estimate to obtain the instantaneous reward for intent guidance, the total resource consumption estimate is determined based on the logical transmission resource consumption estimate and the logical computation resource consumption estimate, and the instantaneous reward for intent guidance is determined based on the XR session quality, the large model service quality, the total resource consumption estimate, and the multidimensional intent vector.
[0016] In some embodiments of the present invention, in the step of updating the corresponding constraint penalty multiplier based on the difference between each delay default rate and the corresponding allowable default probability threshold within a sliding window, the following steps are taken: The XR service delay default rate is determined based on the number of times the motion-to-photon delay exceeds the XR service delay constraint and the statistical count of the motion-to-photon delay within the sliding window; the large model inference service delay default rate is determined based on the number of times the large model inference delay exceeds the large model inference service delay constraint and the statistical count of the large model inference delay within the sliding window; the XR service delay default rate and the large model inference service delay default rate are used as the delay constraint default data; the difference between each delay default rate and the corresponding allowable default probability threshold is determined, and the corresponding constraint penalty multiplier is updated based on this difference.
[0017] A second aspect of the present invention also provides an intent-guided hierarchical resource orchestration system, the system comprising a computer device including a processor and a memory, the memory storing computer instructions, the processor executing the computer instructions stored in the memory, and the system implementing the steps of the method described above when the computer instructions are executed by the processor.
[0018] A third aspect of the invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the aforementioned intention-guided hierarchical resource orchestration method.
[0019] Additional advantages, objects, and features of the invention will be set forth in part in the description which follows, and will also become apparent in part to those skilled in the art upon studying the text, or may be learned by practice of the invention. The objects and other advantages of the invention will become apparent from the description and the accompanying drawings.
[0020] Those skilled in the art will understand that the objectives and advantages achievable with the present invention are not limited to those specifically described above, and that the above and other objectives achievable with the present invention will become clearer from the following detailed description. Attached Figure Description
[0021] The accompanying drawings, which are provided to further illustrate the invention and form part of this application, are not intended to limit the scope of the invention.
[0022] Figure 1 This is a schematic diagram illustrating one implementation of the hierarchical resource orchestration method intended to guide this solution. Figure 2 This is a schematic diagram of the overall processing flow of this solution; Figure 3 This is the data flow diagram for the multi-agent reinforcement learning algorithm in this scheme; Figure 4 This is a schematic diagram of the hierarchical resource orchestration processing framework of this solution. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the embodiments and accompanying drawings. Here, the illustrative embodiments and descriptions of this invention are used to explain the invention, but are not intended to limit the invention.
[0024] It should also be noted that, in order to avoid obscuring the invention with unnecessary details, only the structures and / or processing steps closely related to the solution according to the invention are shown in the accompanying drawings, while other details that are not closely related to the invention are omitted.
[0025] like Figure 1 and 2 As shown, this invention proposes an intent-guided hierarchical resource orchestration method, the steps of which include: Step S100: Acquire multimodal data collected by the terminal, construct a first sensor sequence based on the multimodal data, and input the first sensor sequence into the lightweight student model to obtain low-dimensional semantic labels; In specific implementation, the terminal can be an XR head-mounted device or a terminal device that communicates with an XR head-mounted device. The multimodal data collected by the terminal includes inertial measurement data and gaze point sequences. The inertial measurement data is used to characterize the motion state of the terminal or the user's head, and the gaze point sequences are used to characterize the user's visual attention position at continuous sampling times. The terminal synchronizes the inertial measurement data and gaze point data according to the sampling time, and combines the data at each sampling time according to a preset field order to form a first sensor sequence.
[0026] During the runtime phase, the terminal only inputs the first sensor sequence into the trained lightweight student model, without inputting video frames. The student model determines a low-dimensional semantic label based on the scene category with the highest output probability and sends this low-dimensional semantic label to the orchestration center. The low-dimensional semantic label can be encoded using three-dimensional one-hot encoding, for example... , and These represent high-dynamic training scenarios, precision testing scenarios, and balanced teaching scenarios, respectively.
[0027] Step S200: At the beginning of the current macro cycle, aggregate the running statistics of the previous macro cycle, construct a global context vector based on the running statistics and low-dimensional semantic labels, and input the global context vector into a pre-trained lightweight intent prediction network to obtain a multi-dimensional intent vector. Specifically, let the first Each macro cycle includes If there are several fast-scale control time slots, then the fast-scale control time slots contained in this macrocycle satisfy the following: in, Indicates the macro-cycle number, Indicates the fast-scale control time slot number. This indicates the number of fast-scale control slots contained in a macrocycle. The orchestration center is in the [number missing]. At the start of the first macro cycle, the aggregation of the first... Operational statistics generated by each macro cycle.
[0028] Step S300: Distribute the multidimensional intent vector to a distributed decision entity, so that the distributed decision entity inputs the local observation state and the multidimensional intent vector into a multi-agent reinforcement learning policy network to obtain a joint resource scheduling policy. like Figure 4 As shown, specifically, the distributed decision-making entities include a wireless resource decision-making entity, a computational resource decision-making entity, and a session execution decision-making entity. Each distributed decision-making entity employs a reinforcement learning model and performs decentralized decisions based on its local observation state within a fast-scale control slot. During the training phase, the local observation states of each decision-making entity can be aggregated into a global state to update the policy network in a centralized manner; during the runtime phase, each decision-making entity generates actions solely based on its local observation state and the multi-dimensional intent vector of the current macro-cycle.
[0029] Step S400: Adjust network transmission resources, logical computing resources and session layer service parameters according to the joint resource scheduling strategy, and collect task execution results.
[0030] The above scheme employs a lightweight student model for semantic compression perception, extracting data into low-dimensional semantic labels, thus reducing data complexity. While preserving the semantics of the business scenario, it reduces the amount of data transmitted and processed subsequently, thereby reducing the impact of uplink semantic or contextual data on communication bandwidth and edge computing power. The low-dimensional semantic labels are combined with the operational statistics of the previous macrocycle to form a global context vector, which can uniformly represent the current business scenario, business operation status, and the usage status of shared communication and computing resources. During the online phase, the lightweight intent prediction network is used to determine the scheduling focus, resource consumption constraints, and resource adjustment range for different services. Each distributed decision entity further combines the multi-dimensional intent vector and local observation status to generate a joint resource scheduling strategy for network transmission resources, logical computing resources, and session layer service parameters. Therefore, this solution can determine the resource orchestration intention based on the business scenario and historical operating status within the slow-scale macro cycle, and coordinately adjust communication and computing resources based on the real-time status within the fast-scale control time slot. This enables resource scheduling to adapt to the low MTP latency requirements of XR services and the sudden uplink bandwidth and computing requirements of large language model inference services, thereby alleviating resource competition when the two types of services share the logical bandwidth pool and edge computing power pool, and reducing the risk of service quality degradation caused by static slicing or scheduling for a single objective.
[0031] In some embodiments of the present invention, the method further includes the following steps: determining XR session quality, large model service quality, total resource consumption estimate, and latency constraint default data based on the task execution result of a new macro cycle; combining the XR session quality, the large model service quality, and the total resource consumption estimate based on the XR service weight, the large model service weight, and the resource consumption penalty coefficient in the multidimensional intent vector to obtain an instantaneous reward for intent guidance; statistically calculating the XR service latency default rate and the large model inference service latency default rate within a sliding window based on the latency constraint default data; updating the corresponding constraint penalty multiplier based on the difference between each latency default rate and the corresponding allowed default probability threshold; correcting the instantaneous reward for intent guidance using the updated constraint penalty multiplier and the latency default result of the current fast-scale control slot to obtain a reward with a constraint penalty term; and updating the multi-agent reinforcement learning policy network based on the reward with the constraint penalty term.
[0032] In some embodiments of the present invention, the lightweight student model includes a temporal convolutional network and a gated recurrent unit. In the training step of the lightweight student model, a first sensor sequence and a second sensor sequence are constructed based on the multimodal data in the training data. The first sensor sequence is input into the lightweight student model, and the second sensor sequence is input into the teacher model. The lightweight student model is trained based on the output results of the lightweight student model and the teacher model, as well as the real labels in the training data.
[0033] In practice, the lightweight student model is trained using a teacher-student architecture. The multimodal data in the training data includes time-synchronized video streams, inertial measurement data, and gaze point sequences. The inertial measurement data and gaze point sequences constitute the aforementioned first sensor sequence; the synchronized video stream is associated with the first sensor sequence using the same time index to form a second sensor sequence for the teacher model to process. Thus, the first sensor sequence is used in the student model branch, and the second sensor sequence is used in the teacher model branch.
[0034] In one implementation, the teacher model uses a VideoMAE video mask autoencoder as its backbone network, with approximately 100 million parameters. The lightweight student model has approximately 500,000 parameters and is composed of a temporal convolutional network sequentially connected to gated recurrent units. The temporal convolutional network extracts local variation features from adjacent sampling times and different time spans in the first sensor sequence. The gated recurrent units aggregate the sequential relationships of these local variation features. The classification output layer of the student model generates probability distributions corresponding to each preset business scenario based on the aggregation results.
[0035] Let the scene classification logic value output by the student model be... The scene classification logic value output by the teacher model is The real-world labels of the training samples are Then, the lightweight student model can be trained using the knowledge distillation objective function shown in the following formula: in, This represents the knowledge distillation loss. This represents the cross-entropy loss between the student model output and the real-world scene labels. This represents the divergence between the soft probability distributions of the teacher model and the student model. This represents the Softmax function. This represents the temperature parameter used to smooth the probability distribution. This represents the weighting coefficient between the cross-entropy loss and the divergence loss. By minimizing... Update the parameters of the lightweight student model.
[0036] In one training data construction method, assembly operation videos from the Assembly101 dataset are used, and the operation segments are re-labeled according to business scenarios. Operation segments with obvious head movements or frequent perspective switching can be labeled as high-dynamic training scenarios; operation segments with long-term focus on a specific target and low motion intensity can be labeled as precision detection scenarios; and routine operation segments can be labeled as balanced teaching scenarios. For video segments lacking synchronous sensor recordings, the OpenXR toolchain can be used to generate synthetic gaze trajectories and inertial measurement trajectories corresponding to the video timeline, forming training samples that correspond to the teacher model branch and the student model branch in time.
[0037] In some embodiments of the present invention, in the step of aggregating the operational statistics of the previous macrocycle at the start of the current macrocycle and constructing a global context vector based on the operational statistics and low-dimensional semantic labels, Extract motion-to-photon latency, XR session quality, large model inference latency, GPU utilization, and bandwidth resource utilization from the runtime statistics to obtain the runtime data set; The low-dimensional semantic tags are encoded to obtain tag codes; The global context vector is obtained by combining the running data group with the tag encoding.
[0038] Specifically, the motion-to-photon latency in the runtime data set can be the average MTP latency of XR frames that have been displayed in the previous macrocycle; the XR session quality can be the average XR session quality in the previous macrocycle; the large model inference latency can be the average completion latency of LAM inference tasks that have been completed in the previous macrocycle; the GPU utilization can be the average GPU utilization in the previous macrocycle; and the bandwidth resource utilization can be the average bandwidth resource utilization in the previous macrocycle. The orchestration center processes the above five data items according to preset normalization rules to obtain a 5-dimensional runtime data set.
[0039] The low-dimensional semantic tags obtained in step S100 are encoded into three-dimensional tag codes. The data is then combined according to the field order: 5-dimensional running data group first, followed by 3-dimensional label encoding, to obtain an 8-dimensional global context vector. in, Indicates the first A global context vector with a macro-period. Indicates the first The three-dimensional label encoding corresponds to each macrocycle. Thus, although both the first sensor sequence on the terminal side and the global context vector of the orchestration center can have 8-dimensional data, they belong to different processing stages: the first sensor sequence is a time-series matrix formed by multiple sampling moments, which generates a three-dimensional label encoding after passing through a lightweight student model; the global context vector is obtained by combining the 5-dimensional macrocycle running statistics and the three-dimensional label encoding.
[0040] In some embodiments of the present invention, in the step of inputting the global context vector into a pre-trained lightweight intent prediction network to obtain a multi-dimensional intent vector, the lightweight intent prediction network includes an input layer, a first hidden layer, a second hidden layer and an output layer including a weight head and a boundary head arranged in sequence, wherein the first hidden layer and the second hidden layer both adopt the ReLU activation function. The global context vector is input into the input layer. The weight header is used to output the scheduling weights corresponding to XR services and large model inference services. The boundary header is used to output the upper limit of bandwidth resource utilization, the upper limit of GPU utilization, and the adjustment range of various resource scheduling actions. The outputs of the weight header and the boundary header are combined to obtain a multi-dimensional intent vector.
[0041] Specifically, the lightweight intent prediction network is a multilayer perceptron with two hidden layers belonging to the same shared network backbone. In one implementation, the input layer includes 8 input neurons, the first hidden layer includes 64 neurons, and the second hidden layer includes 32 neurons.
[0042] The pre-trained lightweight intent prediction network is obtained through a semi-offline meta-policy approximation method. For example... Figure 3 As shown, the semi-offline meta-policy approximation method includes offline context library construction, candidate intent generation, candidate intent verification, lightweight intent prediction network training, and online intent prediction.
[0043] The offline context library can be built from system simulation logs. The system performs simulations under different combinations of XR user load, channel signal-to-noise ratio, GPU occupancy status, and bursty LAM request states. Following a macro cycle, the system extracts the aforementioned 5-dimensional runtime data sets and their corresponding 3-dimensional label codes from the runtime logs to form offline context samples. ,in, Indicates the offline context sample number.
[0044] For each offline context sample, the system role, the context sample itself, the balance target between XR and LAM services, and intent parameter constraints are written into structured prompts and input into a large-scale intent generator to obtain multiple raw candidate intents represented in JSON format. Field names, data types, and array structures are validated on the raw candidate intents. The XR service weights, LAM service weights, and resource consumption penalty coefficients are projected using a projection operator to ensure that all three are non-negative and their sum is 1. The bandwidth resource utilization cap, GPU utilization cap, and adjustment ranges for various resource scheduling actions are projected to a value range of 0 to 1. The projected candidate intents are consistent with the output fields and field order of the lightweight intent prediction network.
[0045] For the first through checksum projection The first context sample Candidate intents Perform rollback verification of 50 fast-scale control slots under the same initial context to obtain candidate cumulative rewards: in, Indicates the cumulative reward for candidates. This indicates the rollback verification start time slot corresponding to this context sample. This indicates the relative time slot number in the rollback verification. Indicates candidate intent The intent-guided reward generated in the corresponding time slot is compared. The cumulative rewards of each candidate sample in the same context are compared, and the candidate intent with the largest cumulative reward is determined as the target intent label.
[0046] Specifically, the lightweight intent prediction network is trained using the mean squared error between the multidimensional intent vector predicted by the lightweight intent prediction network and the target intent label.
[0047] In some embodiments of the present invention, the local observation state includes the logical link quality observation matrix of the current time slot, downlink XR data queue state, uplink control and semantic information queue state, GPU utilization, XR rendering queue state, large model inference queue state, user-side estimated instantaneous latency, user-side experience quality estimate, session buffer state, and session-related semantic state; the distributed decision entity includes a wireless resource decision entity, a computing resource decision entity, and a session execution decision entity, all of which employ reinforcement learning models; the joint resource scheduling strategy includes logical resource block allocation actions, logical transmit power control actions, logical computing power allocation actions, task offloading control actions, encoding or rendering quality level adjustment actions, and service parameter adjustment actions. In the step of distributing the multidimensional intent vector to a distributed decision entity, so that the distributed decision entity inputs its local observation state and the multidimensional intent vector into a multi-agent reinforcement learning policy network to obtain a joint resource scheduling policy: The wireless resource decision entity uses the logical link quality observation matrix of the current time slot, the downlink XR data queue status, the uplink control and semantic information queue status, and the multidimensional intent vector as the first observation state. Based on the first observation state, the wireless resource decision entity outputs logical resource block allocation action and logical transmit power control action. Specifically, the first observed state of the wireless resource decision-making entity can be represented as: in, Indicates time slot The first observation state, This represents the logical link quality observation matrix. Indicates the downlink XR data queue status. This indicates the status of the uplink control and semantic information queue. Indicates time slot The multidimensional intent vector belonging to the macrocycle.
[0048] The wireless resource decision entity outputs logical resource block allocation variables. and logic transmit power ,in, Indicates the XR user serial number. Indicates the logical resource block number.
[0049] Logical resource block allocation satisfies: in, Indicates the number of XR users. Indicates in time slot The first A logical resource block is allocated to the user. , This indicates that the allocation has not been performed. The same logical resource block can be allocated to at most one user within the same time slot.
[0050] Based on logical resource block allocation actions and logical transmit power control actions, the user In the time slot The logical transmission rate can be expressed as: in, Indicates logical transmission rate, Indicates the number of logical resource blocks. This represents the logical bandwidth corresponding to a logical resource block. Indicates user Use the The logical signal-to-noise ratio (SNR) estimate for each logical resource block. This logical SNR estimate is based on logical link quality observations and the corresponding logical transmit power. This confirms that the logical transmit power control action participates in the determination of the logical transmission rate. The downlink XR data queue can be updated according to the following formula: in, Indicates user In the time slot Downlink XR data queue backlog Indicates time slot The amount of newly arrived XR data, This indicates the duration of the fast-scale control time slot. The above logical resource block allocation variables and logical transmission rates are used as examples of downlink XR data transmission. For uplink control and semantic information queues, the radio resource decision entity determines the logical resource block allocation results according to the corresponding uplink quality and queue backlog status, and ensures that the same logical resource block satisfies the preset resource mutual exclusion constraints.
[0051] The computing resource decision-making entity uses GPU utilization, XR rendering queue status, large model inference queue status and the multidimensional intent vector as the second observation status. The computing resource decision-making entity outputs logical computing power allocation actions and task unloading control actions based on the second observation status. Specifically, the second observed state of the computational resource decision-making entity can be represented as: in, Indicates time slot The second observation state, Indicates GPU utilization. This indicates the rendering queue status for each XR user. This indicates the LAM inference queue status. The computational resource decision entity outputs the logical computing power quota allocated to the XR rendering task, the logical computing power quota allocated to the LAM inference task, and the task unloading control action. Based on the logical computing power quota, time slots are determined accordingly. The amount of XR rendering computation that can be completed within the system and the computational workload of LAM inference Both conditions are met: in, This represents the total logical computation workload that the edge-side logical computing pool can provide within a fast-scale control time slot. The XR rendering queue and LAM inference queue can be updated respectively according to the following formula: in, Indicates time slot The newly arrived XR rendering computational workload, Indicates time slot The newly arrived LAM inference computation workload. The task offloading control action is used to determine whether the corresponding task remains to be executed in the current compute domain or is transferred to another compute domain for execution.
[0052] The session execution decision entity uses the instantaneous latency estimated by the user side, the estimated user experience quality, the session buffer state, the session-related semantic state, and the multidimensional intent vector as the third observation state. The session execution decision entity outputs session execution actions based on the third observation state. The session execution actions include encoding quality level adjustment actions or rendering quality level adjustment actions, as well as business parameter adjustment actions.
[0053] Specifically, for each XR user, a session execution decision entity can be set up, and its third observation state can be represented as: in, Indicates user In the time slot The third observation state, This represents the instantaneous MTP latency estimated by the user side. This represents the estimated user experience quality. Indicates the session buffer state. Indicates the session-related semantic state. The session execution decision entity outputs the encoding quality level or rendering quality level of the next XR frame. And business parameters used to adjust the business ratio between local and edge environments. ,in, Indicates time slot The frame number of the next XR frame is determined.
[0054] The upper limits of bandwidth resource utilization, GPU utilization, and resource scheduling action adjustment range in the multi-dimensional intent vector can serve as constraints on the action output of the policy network. If an action output by the policy network exceeds the corresponding adjustment range, the action is corrected to the boundary of that range; if executing the corresponding action would cause bandwidth resource utilization or GPU utilization to exceed the corresponding upper limit, the amount of resource allocation that would increase resource utilization is reduced or an executable action that meets the upper limit is selected. Thus, the output of the boundary header is actually used in the joint resource scheduling policy generation stage.
[0055] During the centralized training phase, the observed states of the wireless resource decision-making entity, the computational resource decision-making entity, and each session execution decision-making entity can be combined to form a global state: in, Indicates time slot The global state is determined by combining the local observation states, which already contain multi-dimensional intent vectors. Since each local observation state already includes a multi-dimensional intent vector, combining these local observation states allows the global state to contain multi-dimensional intent information. The centralized commentator network estimates the long-term value of the joint policy based on this global state; each actor network then uses the first... The local observation state of each decision entity serves as the input. During the operational phase, a centralized commentator network is not deployed; instead, each actor network generates scheduling actions at its corresponding decision entity. Thus, the same set of actor networks can change its scheduling emphasis based on multidimensional intent vectors with different macro-periods.
[0056] By adopting the above scheme, the same multi-dimensional intent vector is input into the wireless resource decision entity, the computational resource decision entity, and the session execution decision entity, respectively. This ensures that the three entities are guided by the same macro-periodic scheduling objective and boundary parameters even when using different local observation states. Each entity generates network, computation, and session layer actions within the fast-scale control time slot, enabling local responses to real-time changes in link quality, task queues, and user experience, and reducing resource mismatch caused by the separation of the three types of resource scheduling objectives.
[0057] In some embodiments of the present invention, in the step of adjusting network transmission resources, logical computing resources, and session layer service parameters according to the joint resource scheduling strategy... The logical resource block allocation action and the logical transmit power control action correspond to network transmission resources. Based on the logical resource block allocation action, logical resource blocks are allocated from the logical bandwidth pool for the downlink XR data queue and the uplink control and semantic information queue. Based on the logical transmit power control action, the logical transmit power of the corresponding logical link is set. The encoding or rendering quality level adjustment action and the service parameter adjustment action correspond to the session layer service parameters. The level determined based on the encoding or rendering quality level adjustment action is set as the encoding quality level or rendering quality level of the next XR frame. The service parameter adjustment action adjusts the service ratio between local and edge. The logical computing power allocation action and the task unloading control action correspond to logical computing resources. Based on the logical computing power allocation action, corresponding computing power is allocated from the logical computing power pool to the XR rendering task and the LAM inference task, respectively. Based on the task unloading control action, it is determined whether the task is retained for execution in the current computing domain or transferred to another computing domain for execution.
[0058] In some embodiments of the present invention, in the step of determining delay constraint violation data and resource utilization data based on the task execution results, and updating the multi-agent reinforcement learning policy based on the delay constraint violation data and resource utilization data: Based on the operation of a new macro cycle, a multi-dimensional intent vector is determined, and from the multi-dimensional intent vector, the scheduling weight, bandwidth resource utilization limit, and GPU utilization limit corresponding to each service are determined. Resource utilization data is determined based on the value by which the bandwidth resource utilization exceeds the upper limit of the bandwidth resource utilization and the value by which the GPU utilization exceeds the upper limit of the GPU utilization. Based on the scheduling weights corresponding to each service, the delay constraint default data and resource utilization over-limit data are combined to obtain a feedback reward value; Based on the local observation state before the execution of the joint resource scheduling strategy, the multi-dimensional intent vector, the joint resource scheduling strategy, the feedback reward value, and the local observation state after execution, the network parameters of the multi-agent reinforcement learning strategy network are updated, and the updated multi-agent reinforcement learning strategy network is used to generate subsequent joint resource scheduling strategies.
[0059] In some embodiments of the present invention, in the step of combining the XR session quality, the large model service quality, and the total resource consumption estimate to obtain the instantaneous reward for intent guidance, the total resource consumption estimate is determined based on the logical transmission resource consumption estimate and the logical computation resource consumption estimate, and the instantaneous reward for intent guidance is determined based on the XR session quality, the large model service quality, the total resource consumption estimate, and the multidimensional intent vector.
[0060] Specifically, the quality of an XR session can be determined based on the quality level of the XR frames that have been displayed and the MTP latency. For users... In the time slot Completed set of XR frames for display The quality of its XR session can be expressed as: in, Indicates the quality of the XR session. Indicates the XR frame number. This indicates the encoding quality level or rendering quality level of the XR frame. This indicates the image utility weight corresponding to this quality level. This indicates the MTP latency of the XR frame. This represents the delay sensitivity coefficient corresponding to the quality level. When... When the set is empty, the summation result is 0.
[0061] For time slots Completed LAM reasoning task set The service quality of a large model can be expressed as: in, Indicates the service quality of the large model. Indicates the LAM reasoning task number. This represents the quality score corresponding to the model's quality level in performing the task. This indicates the completion delay of the task. This represents the latency sensitivity coefficient of LAM services.
[0062] Specifically, time slots are set. The estimated logical transmission resource consumption is The estimated value of logical computation resource consumption is The estimated total resource consumption is: in, This represents the estimated total resource consumption. It can be determined based on the number of allocated logical resource blocks and the corresponding logical transmit power. The resource consumption estimate can be determined based on the logical computing power allocated to the XR rendering task and the LAM inference task. After normalizing the above resource consumption estimates to the same dimensions, the instantaneous reward for intent guidance is constructed: in, Indicates time slot In multidimensional intent vector Instantaneous rewards under guidance , and These represent the XR business weight, large model service weight, and resource consumption penalty coefficient in the multidimensional intent vector, respectively. Therefore, all three components of the lightweight intent prediction network weight head output are incorporated into the policy optimization objective.
[0063] In some embodiments of the present invention, in the step of updating the corresponding constraint penalty multiplier based on the difference between each delay default rate and the corresponding allowable default probability threshold within a sliding window, the following steps are taken: The XR service delay default rate is determined based on the number of times the motion-to-photon delay exceeds the XR service delay constraint and the statistical count of the motion-to-photon delay within the sliding window; the large model inference service delay default rate is determined based on the number of times the large model inference delay exceeds the large model inference service delay constraint and the statistical count of the large model inference delay within the sliding window; the XR service delay default rate and the large model inference service delay default rate are used as the delay constraint default data; the difference between each delay default rate and the corresponding allowable default probability threshold is determined, and the corresponding constraint penalty multiplier is updated based on this difference.
[0064] Specifically, delay constraint default data can be initially recorded as the actual completion delay of each XR frame and each LAM inference task within the sliding window, the corresponding delay constraint, and the default indication result item by item. Let the XR service delay constraint be... The latency constraint for large model inference services is When the actual delay is significantly greater than the corresponding delay constraint, the item is recorded as a breach; when the actual delay is equal to the corresponding delay constraint, it is not recorded as a breach.
[0065] Let the permissible default probability thresholds for XR services and large model inference services be respectively... and The corresponding constraint penalty multipliers are respectively and Then the constraint penalty multiplier can be updated according to the following formula: in, and These represent the first fast-scale control time slots used. Individual XR user constraint penalty multiplier and large model inference business constraint penalty multiplier; This indicates the update step size of the constraint penalty multiplier. This means correcting results less than zero to zero. Indicates the cutoff time slot The first sliding window The latency default rate of each XR user; Indicates the cutoff time slot The latency default rate of large model inference operations within the sliding window. When the latency default rate is higher than the corresponding allowable default probability threshold, the corresponding constraint penalty multiplier increases; when the latency default rate is lower than the corresponding allowable default probability threshold, the corresponding constraint penalty multiplier decreases but is not lower than zero.
[0066] Using the updated constraint penalty multiplier and the delay violation result of the current fast-scale control slot, the reward with constraint penalty term is obtained: in, This indicates a reward with a penalty clause. Indicates time slot The number of XR frames that experience MTP delay defaults within the time frame. Indicates time slot The number of LAM inference tasks that violate inference delay rules. For indicator functions, This indicates the maximum allowable latency threshold for XR services. This represents the maximum allowed latency threshold for large model inference operations, within a time slot. Inside, set Indicates the first MTP latency observations for each XR user This represents the completion delay observation value corresponding to the large model inference service.
[0067] The local observation state before the execution of the joint resource scheduling policy, the multi-dimensional intent vector of the current macro-cycle, the joint resource scheduling policy, the reward with the constraint penalty term, and the local observation state after execution are used as a policy update sample to update the multi-agent reinforcement learning policy network.
[0068] In summary, this solution first organizes continuously collected inertial measurement data and gaze point sequences into a first sensor sequence at the terminal. A lightweight student model, trained by teacher model distillation, converts this first sensor sequence into low-dimensional semantic labels. At the start of the current macrocycle, the orchestration center combines these low-dimensional semantic labels with the operational statistics from the previous macrocycle to form a global context vector. A pre-trained lightweight intent prediction network then generates a multi-dimensional intent vector. This multi-dimensional intent vector characterizes the scheduling emphasis, resource consumption constraints, and adjustment boundaries of resource scheduling actions for XR and LAM services. Within the fast-scale control time slot, each distributed decision entity inputs the multi-dimensional intent vector and its local observation state into a multi-agent reinforcement learning policy network to generate scheduling actions corresponding to network transmission resources, logical computing resources, and session layer service parameters. The service quality, resource consumption, and latency violation results generated after executing the scheduling actions are used to correct rewards and update the policy network, thus forming a closed loop where slow-scale intent generation, fast-scale resource scheduling, and execution result feedback are interconnected.
[0069] This invention also provides an intent-guided hierarchical resource orchestration system, which includes a computer device, a processor, and a memory. The memory stores computer instructions, and the processor executes the computer instructions stored in the memory. When the computer instructions are executed by the processor, the system implements the steps of the method described above.
[0070] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the aforementioned intention-guided hierarchical resource orchestration method. The computer-readable storage medium may be a tangible storage medium, such as random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, floppy disk, hard disk, removable storage disk, CD-ROM, or any other form of storage medium known in the art.
[0071] Those skilled in the art will understand that the exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Whether implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this invention. When implemented in hardware, it can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this invention are programs or code segments used to perform the desired tasks. The programs or code segments can be stored in a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried in a carrier wave.
[0072] It should be clarified that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of the present invention.
[0073] In this invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or in place of features of other embodiments.
[0074] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, various modifications and variations of the embodiments of the present invention are possible. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. An intent-guided hierarchical resource orchestration method, characterized in that, The steps of this method include: Acquire multimodal data collected by the terminal, construct a first sensor sequence based on the multimodal data, and input the first sensor sequence into a lightweight student model to obtain low-dimensional semantic labels; At the start of the current macrocycle, the running statistics of the previous macrocycle are aggregated. A global context vector is constructed based on the running statistics and low-dimensional semantic labels. The global context vector is then input into a pre-trained lightweight intent prediction network to obtain a multi-dimensional intent vector. The multidimensional intent vector is distributed to a distributed decision entity, so that the distributed decision entity inputs its local observation state and the multidimensional intent vector into a multi-agent reinforcement learning policy network to obtain a joint resource scheduling policy. The network transmission resources, logical computing resources, and session layer service parameters are adjusted according to the joint resource scheduling strategy, and the task execution results are collected.
2. The intent-guided hierarchical resource orchestration method according to claim 1, characterized in that, The method further includes the following steps: determining XR session quality, large model service quality, total resource consumption estimate, and latency constraint default data based on the task execution result of a new macro cycle; combining the XR session quality, large model service quality, and total resource consumption estimate based on the XR service weights, large model service weights, and resource consumption penalty coefficients in the multidimensional intent vector to obtain an instantaneous reward for intent guidance; statistically calculating the XR service latency default rate and large model inference service latency default rate within a sliding window based on the latency constraint default data; updating the corresponding constraint penalty multiplier based on the difference between each latency default rate and the corresponding allowed default probability threshold; correcting the instantaneous reward for intent guidance using the updated constraint penalty multiplier and the latency default result of the current fast-scale control slot to obtain a reward with a constraint penalty term; and updating the multi-agent reinforcement learning policy network based on the reward with the constraint penalty term.
3. The intent-guided hierarchical resource orchestration method according to claim 1, characterized in that, The lightweight student model includes a temporal convolutional network and a gated recurrent unit. In the training step of the lightweight student model, a first sensor sequence and a second sensor sequence are constructed based on the multimodal data in the training data. The first sensor sequence is input into the lightweight student model, and the second sensor sequence is input into the teacher model. The lightweight student model is trained based on the output results of the lightweight student model and the teacher model, as well as the real labels in the training data.
4. The intent-guided hierarchical resource orchestration method according to claim 2, characterized in that, In the step of aggregating the operational statistics of the previous macrocycle at the start of the current macrocycle, and constructing a global context vector based on the operational statistics and low-dimensional semantic labels, Extract motion-to-photon latency, XR session quality, large model inference latency, GPU utilization, and bandwidth resource utilization from the runtime statistics to obtain the runtime data set; The low-dimensional semantic tags are encoded to obtain tag codes; The global context vector is obtained by combining the running data group with the tag encoding.
5. The intent-guided hierarchical resource orchestration method according to claim 1, characterized in that, In the step of inputting the global context vector into a pre-trained lightweight intent prediction network to obtain a multi-dimensional intent vector, the lightweight intent prediction network includes an input layer, a first hidden layer, a second hidden layer, and an output layer including a weight header and a boundary header, arranged in sequence. Both the first hidden layer and the second hidden layer use the ReLU activation function. The global context vector is input into the input layer. The weight header is used to output the scheduling weights corresponding to XR services and large model inference services. The boundary header is used to output the upper limit of bandwidth resource utilization, the upper limit of GPU utilization, and the adjustment range of various resource scheduling actions. The outputs of the weight header and the boundary header are combined to obtain a multi-dimensional intent vector.
6. The intent-guided hierarchical resource orchestration method according to claim 1, characterized in that, The local observation status includes the logical link quality observation matrix of the current time slot, downlink XR data queue status, uplink control and semantic information queue status, GPU utilization, XR rendering queue status, large model inference queue status, user-side estimated instantaneous latency, user-side experience quality estimate, session buffer status, and session-related semantic status; the distributed decision-making entity includes a wireless resource decision-making entity, a computing resource decision-making entity, and a session execution decision-making entity, all of which employ reinforcement learning models; the joint resource scheduling strategy includes logical resource block allocation actions, logical transmit power control actions, logical computing power allocation actions, task offloading control actions, encoding or rendering quality level adjustment actions, and service parameter adjustment actions. In the step of distributing the multidimensional intent vector to a distributed decision entity, so that the distributed decision entity inputs its local observation state and the multidimensional intent vector into a multi-agent reinforcement learning policy network to obtain a joint resource scheduling policy: The wireless resource decision entity uses the logical link quality observation matrix of the current time slot, the downlink XR data queue status, the uplink control and semantic information queue status, and the multidimensional intent vector as the first observation state. Based on the first observation state, the wireless resource decision entity outputs logical resource block allocation action and logical transmit power control action. The computing resource decision-making entity uses GPU utilization, XR rendering queue status, large model inference queue status and the multidimensional intent vector as the second observation status. The computing resource decision-making entity outputs logical computing power allocation actions and task unloading control actions based on the second observation status. The session execution decision entity uses the instantaneous latency estimated by the user side, the estimated user experience quality, the session buffer state, the session-related semantic state, and the multidimensional intent vector as the third observation state. The session execution decision entity outputs session execution actions based on the third observation state. The session execution actions include encoding quality level adjustment actions or rendering quality level adjustment actions, as well as business parameter adjustment actions.
7. The intent-guided hierarchical resource orchestration method according to claim 6, characterized in that, In the step of adjusting network transmission resources, logical computing resources, and session layer service parameters according to the joint resource scheduling strategy, The logical resource block allocation action and the logical transmit power control action correspond to network transmission resources. Based on the logical resource block allocation action, logical resource blocks are allocated from the logical bandwidth pool for the downlink XR data queue and the uplink control and semantic information queue. Based on the logical transmit power control action, the logical transmit power of the corresponding logical link is set. The encoding or rendering quality level adjustment action and the service parameter adjustment action correspond to the session layer service parameters. The level determined based on the encoding or rendering quality level adjustment action is set as the encoding quality level or rendering quality level of the next XR frame. The service parameter adjustment action adjusts the service ratio between local and edge. The logical computing power allocation action and the task unloading control action correspond to logical computing resources. Based on the logical computing power allocation action, corresponding computing power is allocated from the logical computing power pool to the XR rendering task and the LAM inference task, respectively. Based on the task unloading control action, it is determined whether the task is retained for execution in the current computing domain or transferred to another computing domain for execution.
8. The intent-guided hierarchical resource orchestration method according to claim 2, characterized in that, In the step of combining the XR session quality, the large model service quality, and the total resource consumption estimate to obtain the instantaneous reward for intent guidance, the total resource consumption estimate is determined based on the logical transmission resource consumption estimate and the logical computation resource consumption estimate, and the instantaneous reward for intent guidance is determined based on the XR session quality, the large model service quality, the total resource consumption estimate, and the multidimensional intent vector.
9. The intent-guided hierarchical resource orchestration method according to claim 4, characterized in that, In the step of updating the corresponding constraint penalty multiplier based on the difference between each delay default rate and the corresponding allowable default probability threshold within a sliding window, the following steps are taken: First, the XR service delay default rate is determined based on the number of times the motion-to-photon delay exceeds the XR service delay constraint and the statistical count of the motion-to-photon delay within the sliding window. Second, the large model inference service delay default rate is determined based on the number of times the large model inference delay exceeds the large model inference service delay constraint and the statistical count of the large model inference delay within the sliding window. Third, the XR service delay default rate and the large model inference service delay default rate are used as the delay constraint default data. Fourth, the difference between each delay default rate and the corresponding allowable default probability threshold is determined, and the corresponding constraint penalty multiplier is updated based on this difference.
10. An intent-guided hierarchical resource orchestration system, characterized in that, The system includes a computer device, which includes a processor and a memory. The memory stores computer instructions, and the processor executes the computer instructions stored in the memory. When the computer instructions are executed by the processor, the system implements the steps of the method as described in any one of claims 1 to 9.