Resource allocation method of AI generation content service based on integrated perception and communication
By optimizing the allocation of perception and communication resources in integrated perception and communication scenarios, and combining deep reinforcement learning and ranking algorithms, the service quality problem of AI-generated content services under resource constraints was solved, achieving efficient resource allocation and improved user experience.
Patent Information
- Application Number
- CN202511053332.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-04
AI Technical Summary
Existing AI-generated content services struggle to allocate resources effectively in integrated perception and communication scenarios, leading to difficulties in guaranteeing service quality. This is especially true when perception and communication resources are limited, impacting the accuracy of generated content and user experience.
By acquiring the display capacity and channel status of user equipment within the coverage area of integrated sensing and communication functions, an iterative adjustment method for the allocation of sensing and communication resources is adopted. Combined with deep reinforcement learning algorithms and ranking algorithms, the energy allocation of sensing and communication is optimized to ensure sensing accuracy and generated image quality. Time division multiplexing and orthogonal frequency division multiplexing technologies are used for sensing and communication to construct a cross-modal sensing-understanding-generation architecture.
It achieves efficient resource allocation in integrated sensing and communication scenarios, improves the average service quality and personalized experience for users, reduces the difficulty of solving problems, and improves the accuracy of AI-generated content and communication efficiency.
Smart Images

Figure CN120897272A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of communication, in particular to the field of wireless communication technology, and more particularly to a resource allocation method for AI-generated content service based on integrated sensing and communication. BACKGROUND
[0002] With the rapid development of data analysis, hardware performance and artificial intelligence (AI), various AI-generated content (AIGC) services have become increasingly popular. For example, ChatGPT generates and understands text according to user instructions, image generation system DALL·E is used to generate high-quality images, and systems such as Sora focus on the generation and interaction of video content. The emergence of these services has promoted the development of natural language processing, computer vision and multimedia processing, making human-computer interaction more intelligent and efficient. However, these applications rely on traditional human-computer interaction methods and only generate content based on input text and voice prompts. However, certain user intentions (such as physical gestures) and environmental changes are difficult to express through text and language. Traditional methods capture this information through cameras or sensors, which can be costly and raise privacy concerns. For example, extended reality (XR) applications may collect facial features, leading to potential misuse of personal data.
[0003] Emerging integrated sensing and communication (ISAC) can perceive human gestures and environmental information by using wireless signals, thus providing a relatively private and non-contact alternative. It is expected that AIGC will be integrated with ISAC to provide various services. However, in this scenario, it is currently difficult to reasonably allocate related resources to ensure the quality of service of AI-generated content.
[0004] It should be noted that the background art is only used to introduce the related information of the present application, so as to help understand the technical solutions of the present application, but does not mean that the related information must be prior art. The related information is submitted and disclosed together with the present application scheme, and in the absence of evidence that the related information has been publicly disclosed before the filing date of the present application, the related information should not be considered as prior art. SUMMARY
[0005] Therefore, the purpose of the present application is to overcome the defects of the prior art, and to provide a resource allocation method for AI-generated content service based on integrated sensing and communication.
[0006] The purpose of the present application is achieved by the following technical solutions:
[0007] According to a first aspect of the present application, a resource allocation method is provided for allocating sensing and communication resources for an artificial intelligence based on integrated sensing and communication technology to generate content services, the method comprising: obtaining display capacity and channel related state of a plurality of user equipment participating in this time resource allocation within the coverage of a transceiver with integrated sensing and communication functions; iteratively adjusting the sensing resources and communication resources to be allocated to the plurality of user equipment this time according to the display capacity and channel related state of the plurality of user equipment, with the optimization objective of maximizing the average service quality of users, to obtain an allocation decision, wherein the service quality of each user is related to the sensing accuracy of the user posture information obtained through the sensing resources and the quality of the generated image corresponding to the digital human driven by the human posture through the user equipment. The scheme can at least achieve the following beneficial technical effects: the present application first incorporates the sensing accuracy affecting the accuracy of AI generated content into the evaluation of the service quality of users, and in addition, considers the quality of the generated image, to better improve the overall service quality from both content accuracy and user service quality.
[0008] Optionally, the service quality of each user is the product of the standard service quality and the personalized immersion coefficient required by the user. The scheme can at least achieve the following beneficial technical effects: the scheme proposes a personalized weight mechanism to support dynamic adaptation of user subjective preferences, and better matches the personalized needs of users.
[0009] Optionally, the optimization objective is set as:
[0010]
[0011] wherein, represents the sensing energy related to the user equipment, represents the communication energy related to the user equipment, represents the sensing energy related to user equipment k, represents the communication energy related to user equipment k, represents the number of user equipment participating in this time resource allocation, represents the personalized immersion coefficient of user equipment k, represents the standard service quality corresponding to user k, s.t. represents being constrained to, C1-C5 represents constraint conditions C1-C5, represents the transmission time of user equipment to transmit the generated image, and condition C1 represents that the total sensing time of all user equipment plus the transmission time of any user equipment should be less than the maximum service time, represents the number of sensing signals in the sensing resources allocated to user k, represents the duration of each sensing signal, represents the maximum service time; condition C2 represents that the total energy consumption should be less than the maximum energy of the transceiver ; condition C3 means that the perceived energy allocated to each user equipment should not exceed the maximum perceived energy , condition C4 requires that the standard quality of service of each user equipment should meet the personalized minimum requirement , condition C5 means that the resolution of the generated image should be less than the display capacity of the user equipment k . The scheme can at least achieve the following beneficial technical effects: the scheme defines the optimization target and its multiple constraint conditions, which can better guide the allocation of perceived resources and communication resources, obtain a solution that matches different user personalized needs, and improve the average service quality of users.
[0012] Optionally, the maximum perceived accuracy of each user does not exceed the upper limit of the highest perceived accuracy, and is positively correlated with the number of rounds of the perceived signal allocated thereto.
[0013] Optionally, the perceived accuracy of each user is calculated in the following manner:
[0014] , or
[0015] , or
[0016] ,
[0017] wherein, represents the upper limit of the highest perceived accuracy corresponding to the user k, represents the gain sparsity of the number of rounds of the perceived signal corresponding to the user k, represents the number of rounds of the perceived signal corresponding to the user k, represents the sensitivity coefficient of the number of rounds of the perceived signal corresponding to the user k, is a positive constant.
[0018] Optionally, the quality of the generated image corresponding to each user is equal to the resolution of the generated image divided by the display capacity of the user equipment of the user, and when the quality of the generated image exceeds 1, it is still considered as 1. The scheme can at least achieve the following beneficial technical effects: the quality of the generated image is intuitively set as the resolution of the generated image divided by the display capacity of the user equipment, which can better rely on intuitive and parameterized data to represent the quality of the generated image during optimization, and finally achieve the effect of improving user experience.
[0019] Optionally, in the allocation of sensing resources and communication resources, a preset double-layer resource allocation mechanism is adopted, which is configured to: first, in the top layer, use a deep reinforcement learning algorithm with an action filter to allocate sensing resources, wherein the deep reinforcement learning algorithm is used to calculate the action of each user according to the state related to the channel of all users, and the action filter is used to map the action of each user to the effective sensing energy range by setting the minimum sensing energy and the maximum sensing energy, so as to obtain the sensing energy allocated to each user, which can be converted into the number of sensing signals; then, in the bottom layer, use a sorting-based communication energy allocation algorithm to allocate communication resources, wherein the users are sorted in descending order according to the service quality of the users that can be improved per unit of communication energy; when the remaining energy is less than the sum of the minimum communication energy required by all users, the minimum communication energy required by the users is satisfied in order from the user at the front of the order; when the remaining energy is greater than or equal to the sum of the minimum communication energy required by all users, each user is allocated at least the minimum communication energy, and then the maximum communication energy is allocated to each user in order from the user at the front of the order until the remaining energy is allocated. This scheme can at least achieve the following beneficial technical effects: through the double-layer resource allocation mechanism, the solution space is reduced by several times, and the solution difficulty is effectively reduced.
[0020] According to the second aspect of the present application, a method for controlling the motion of a digital person based on the posture of a sensing user is provided, which comprises: obtaining the display capacity and the state related to the channel of a plurality of user devices participating in resource allocation this time, and using the allocation decision obtained by the resource allocation method according to the first aspect; generating images based on the allocation decision and a communication-sensing integrated frame structure, wherein the communication-sensing integrated frame structure is configured to: in the sensing stage, use a transceiver to send a frequency-modulated continuous wave to perform posture sensing on each user respectively in a time-division multiplexing manner using the entire bandwidth, and obtain posture sensing data of each user, wherein each user performs posture sensing in a sensing time determined by the number of sensing signals allocated to the user and the duration of each sensing signal, and a guard interval is provided after the sensing time of each user ends; in the communication stage, obtain generated images obtained after a server determines control actions of a digital person based on the posture sensing data of each user, and control the transceiver to use orthogonal frequency division multiplexing technology or non-orthogonal multiple access technology to transmit the generated images to the user devices corresponding to each user for display. This scheme can at least achieve the following beneficial technical effects: in the sensing stage, each user is sensed in a time-division multiplexing manner using the entire bandwidth, and the sensing signals of the users are provided with a guard interval, which can reduce the mutual interference of the sensing signals between users and improve the sensing accuracy; in the communication stage, the generated images are transmitted to the user devices corresponding to each user using orthogonal frequency division multiplexing technology or non-orthogonal multiple access technology, so as to improve the spectrum efficiency.
[0021] According to a third aspect of the present application, there is provided an electronic device comprising: one or more processors; and a memory, wherein the memory is configured to store executable instructions; the one or more processors are configured to implement the steps of the method of the first and / or second aspects via execution of the executable instructions. BRIEF DESCRIPTION OF DRAWINGS
[0022] The embodiments of the present application will be further described with reference to the drawings.
[0023] Figure 1 Flowchart of the resource allocation method according to the embodiments of the present application;
[0024] Figure 2 Schematic diagram of the application scenario of the resource allocation according to the embodiments of the present application;
[0025] Figure 3 Schematic diagram of the time-frequency resource frame structure of the ISAC according to the embodiments of the present application;
[0026] Figure 4 Flowchart of the AIGC service process driven by the ISAC according to the embodiments of the present application;
[0027] Figure 5 Schematic diagram of the double-layer resource allocation mechanism according to the embodiments of the present application;
[0028] Figure 6 Schematic diagram of the average CAQA change curve with the maximum total energy according to the embodiments of the present application;
[0029] Figure 7 Schematic diagram of the average CAQA change trend with the number of users according to the embodiments of the present application;
[0030] Figure 8 Schematic diagram of the average CAQA change trend with the maximum total energy under different maximum service times according to the embodiments of the present application. DETAILED DESCRIPTION
[0031] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application.
[0032] As mentioned in the background section, AIGC is expected to be integrated with ISAC to provide various services. But in this scenario, it is still difficult to allocate resources reasonably to guarantee the quality of service. The services that can be provided are, for example: digital human live streaming, virtual fitting, game and movie making, etc. Illustratively, the ISAC technology can be used to perceive the user's body posture and other information (such as the overall descendant or gesture), and then the AIGC service generates the content required by the user (i.e. the generated image / video composed of multiple images and other multimedia information corresponding to the posture) based on the perceived user posture information (for example, the user's body skeleton posture), and finally the communication capability of ISAC can be used to transmit the generated content to the user equipment (display equipment) on the user side for display.
[0033] Therefore, the AIGC service based on ISAC will bring a better experience to the user. For example, in an interactive virtual environment, real-time data such as body posture and movement that is difficult to convey through language can be captured through wireless sensing. This enables AI-generated content services to generate more accurate virtual content based on these actions, making the experience more responsive and immersive. However, there are still great challenges in AI-generated content services based on ISAC, especially in the case of limited ISAC resources. Because allocating more resources to the perception process will bring more accurate input information to the input end of the AI-generated content service, but since communication and perception share ISAC resources, the resources occupied by communication will decrease, which will lead to a decrease in communication capability, and then the transmission quality of communication will decrease, which will affect the quality of AI-generated content service. Conversely, the same is true. Therefore, how to balance the resources of perception and communication to maximize the quality of AI-generated content service is crucial.
[0034] In the scenario of ISAC-driven AI-generated content services, existing resource allocation schemes cannot be directly used. Because existing AI-generated content service resource allocation methods usually do not consider the impact of perception and communication on the quality of AI-generated content services. For example, existing AI-generated content service quality evaluation usually evaluates content generation quality (CGQ) without considering content accuracy, because they assume that the input information and command prompts are accurate. However, this method is not suitable for AIGC networks based on ISAC, because in such networks, although the command prompt is still accurate, the content generation depends on the not completely accurate perception data.
[0035] According to an embodiment of the present application, see Figure 1The application provides a resource allocation method, comprising the following steps: S1, obtaining the display capacity of a plurality of user equipment participating in resource allocation this time and the channel-related state within the coverage range of an integrated sensing and communication function transceiver; S2, iteratively adjusting the sensing resources and communication resources to be allocated to the plurality of user equipment this time according to the display capacity of the plurality of user equipment and the channel-related state, so as to obtain an allocation decision, wherein the service quality of each user is related to the sensing accuracy of user posture information obtained through sensing resources and the quality of generated images corresponding to digital people driven by human body posture obtained through user equipment. In order to facilitate comprehensive understanding of the application scheme, the following aspects are introduced: scene description, ISAC frame structure scheme, performance evaluation scheme of service quality of a single user, joint sensing and communication resource allocation scheme, etc.
[0036] I. Scene description
[0037] The existing AI generated content service usually executes a generative AI model on a cloud server and accesses a cloud-based AIGC service on a core network. The application can adopt this mode, sense user posture information of a user locally with a transceiver, transmit the user posture information to a cloud server, generate generated images corresponding to digital people according to the user posture information by the cloud server, and then transmit the generated images to a user equipment for display.
[0038] However, due to the remote nature of users, cloud services exhibit high latency. Therefore, it is also recommended to deploy generative AI services on the edge side to directly produce, process and distribute content on the edge side using generative AI technology. Alternatively, services can be deployed on the edge side and in the cloud at the same time, such as Figure 2In this paper, we propose an ISAC-driven AIGC network architecture to effectively solve the problems of high latency, insufficient interaction mode and resource allocation in traditional AIGC. In terms of network architecture, we construct a four-layer heterogeneous architecture including user layer, perception layer, edge layer and cloud layer. The perception layer realizes non-contact precise perception of multi-user posture and environmental information with the help of the transceiver of ISAC; the edge layer deploys lightweight AIGC models that can extract semantic features and generate content in real time for posture information; the cloud layer is responsible for executing complex AIGC tasks and global resource optimization. The transceiver of ISAC is connected to the AIGC service provider through wired / wireless links. Among them, the AIGC service provider deploys AIGC services on both the cloud and the edge side. The cloud relies on high-performance computing clusters to undertake AIGC model pre-training, fine-tuning and version iteration tasks (such as image generation based on diffusion models, multi-modal semantic alignment training). The trained model generates a lightweight version through knowledge distillation and dynamic pruning techniques and is delivered to the edge node (such as the server on the base station side) through wired / wireless links. The lightweight AIGC model deployed on the edge side computing resources (such as edge servers, other edge computing devices, etc.) supports low-latency inference, for example, real-time generation of character image content in VR environments. At the same time, the cloud continuously monitors the generation quality feedback of the edge node (such as user interaction satisfaction), dynamically triggers model incremental update or reconstruction distillation strategy to adapt to the dynamic changes of the edge scene (such as different VR applications, different users).
[0039] The transceiver of ISAC can capture the user's posture (which can be a global posture (such as the overall joint pose) or a local posture (such as a hand gesture)) through wireless sensing and upload it to the AIGC service provider on the edge side through wireless communication / wired communication, then the AIGC model on the edge side generates the corresponding content based on the user posture information captured by wireless sensing and returns it to the ISAC device, finally the transceiver of the ISAC device forwards the generated content (for example, a digital human image with corresponding posture for a virtual avatar in a VR game) received from the AIGC server to the user's user device (such as a VR device). Assuming that there are K users within the coverage range of the transceiver, the bandwidth of the system is B. Due to the different needs of perception tasks and communication tasks, the power of perception and communication is generally different.
[0040] Compared with traditional human-computer interaction, since the sensed posture is used as the input of AIGC, the perception accuracy of the transceiver will affect the accuracy of the content generated by the AIGC server, and further affect the user experience quality (QoE) in the virtual game process. On the other hand, the resolution of the digital human image generated by AIGC also has a positive impact on user experience quality, which depends on communication capability. Therefore, both perception and communication affect the overall QoE of ISAC-based AIGC services, and both need more resources to improve QoE.
[0041] II. ISAC Frame Structure Scheme
[0042] like Figure 3 As shown, the ISAC transceiver transmits a frequency-modulated continuous wave (FMCW) using time-division multiplexing (TDM) to perform attitude sensing on K users across the entire bandwidth B, with a guard interval after each user's sensing is completed. The sensed attitudes are sent to the AIGC provider, which generates corresponding digital human images and returns them to the ISAC transceiver via a backhaul link. Finally, the ISAC transceiver uses orthogonal frequency-division multiplexing (OFDM) technology to transmit the digital human images to the users' XR / VR devices, where bandwidth B is evenly divided into K orthogonal sub-channels for communication. In terms of frame structure design, a communication-sensing integrated frame structure with time slot-level resource allocation is adopted. The sensing time slot utilizes ISAC to achieve centimeter-level positioning and channel state information acquisition; the communication time slot employs orthogonal frequency division multiplexing (or alternatively, non-orthogonal multiple access) technology to transmit AIGC-generated content; and the control time slot is used to provide feedback on CAQA indicators and user requirements to form a closed-loop optimization. The innovation of this scheme lies in its first-ever combination of ISAC's dynamic sensing capabilities with AIGC's semantic understanding and generation capabilities, constructing a cross-modal sensing-understanding-generation architecture. Furthermore, the proposed frame structure supports microsecond-level time slot switching, enabling fine-grained resource reuse for communication sensing tasks.
[0043] like Figure 4 As shown, the ISAC-driven AIGC network operates in three phases. Illustratively, in the first phase (corresponding to the sensing phase), the ISAC transceiver transmits a frequency-modulated continuous wave (FMCW) to perform sensing and acquire the user's posture. Assume the number of cycles of the sensing signal... Let represent the number of rounds of the sensing signal for the k-th user. Each sensing signal cycle includes a period of duration . The sensing waveform. The accuracy of user posture (i.e., perception accuracy) is positively influenced by the number of rounds of assigned sensing signals. Perception accuracy is related to the number of signal rounds assigned; the higher the number of signal rounds, the higher the perception accuracy. The number of signal rounds is related to the allocated sensing energy; the greater the sensing energy, the more signal rounds are required. Illustratively, the perception accuracy for each user is calculated as follows:
[0044] ,or
[0045] , or
[0046] ,
[0047] where, represents the upper limit of the highest perceived accuracy corresponding to user k, represents the gain sparsity of the round of the perceived signal corresponding to user k, represents the round of the perceived signal corresponding to user k, represents the sensitivity coefficient of the round of the perceived signal corresponding to user k, is a positive constant.
[0048] Because the duration of each perceived signal is , the perceived energy required by user k is , where is the perceived power to user k.
[0049] In the second stage, AIGC generates the corresponding digital human body image (corresponding to the generated image) on the edge side according to the user posture and prompts perceived in the first stage, and returns it to the transceiver of ISAC through the backhaul link. Assuming that the backhaul capacity is large, the uploading and returning time can be ignored, and due to the sufficient computing resources on the edge side, the inference time of lightweight AIGC can be ignored. Therefore, the accuracy and quality of the generated content are not limited by the backhaul capacity. The focus of the present application is the resource allocation of perception and communication on the transceiver of ISAC.
[0050] In the third stage, the transceiver of ISAC transmits the generated image to the user equipment (such as VR / AR equipment) of K users through OFDM. It needs to be noted that transmitting higher quality images requires more communication capacity. According to Shannon theory, let the transmission rate of the downlink communication corresponding to user k be , where, is the bandwidth, K is the total number of users, is the channel gain of user k, including large-scale path loss and small-scale Rayleigh fading, is the power of additive white Gaussian noise, is the communication power of the user. Given the communication time of user k is , the communication capacity depends on the channel gain and communication time, and the communication energy consumed by user k is .
[0051] III. Performance evaluation scheme of service quality of a single user
[0052] In the AIGC based on ISAC, the quality of service depends on the accuracy (precision) and quality of the generated content received by the user. The precision of the received content is determined by the sensing precision of the first stage, while the quality of the received content is limited by the communication capability of the third stage. As shown in FIG. 6, in ISAC, the sensing signal and the communication signal share the same radio resource. Therefore, in order to provide good quality of service, the resources should be allocated carefully for sensing and communication. Figure 2
[0053] According to one embodiment of the present application, a performance evaluation scheme for ISAC-driven AIGC service is designed to evaluate the quality of service of users, to solve the problem that the quality evaluation of the existing AIGC service does not consider the accuracy of sensing data and lacks personalization. The scheme constructs a CAQA index system and proposes a two-dimensional quality evaluation model related to content accuracy and content generation quality. The content accuracy is measured by the matching degree of the sensing data and the real environment; the content quality is based on the resolution (image) or frame rate (video) of the generated content, and the higher the resolution / frame rate, the higher the content generation quality. Considering that if the content accuracy is not high, even if the content quality is high, it has no effect on the improvement of the quality of ISAC-driven AIGC service, therefore CAQA is defined as the product of the two. At the same time, the personalized CAQA demand based on the user's immersive experience is proposed, wherein, is the standard user CAQA demand, is the personalized immersion coefficient based on the user k's experience, is the actual CAQA demand of user k. The innovation of the scheme is that it first includes the influence of sensing data on the accuracy of AI-generated content into the AIGC service quality evaluation system, and proposes a personalized weight mechanism to support dynamic adaptation of user subjective preferences.
[0054] Since QoE depends on the subjective feelings of users, and the subjective feelings of users are different due to individual needs, resources should be allocated according to the differentiated service requirements of the task. In addition, the quality of service of AIGC based on ISAC depends on the accuracy and quality of the generated content received by the user. The present application proposes a new QoE model, namely the CAQA evaluation model, which involves both the accuracy and quality of the generated image. First, since the image generation is based on the sensed user pose, the accuracy of the generated image can be approximated by the sensing precision. Second, the image quality is measured by the resolution, and the higher the resolution, the better the quality. However, transmitting higher resolution images requires more communication capacity. However, the communication capacity is limited by the ISAC radio resource, because the ISAC radio resource is shared with sensing. If the size of the generated image exceeds the available communication capability, i.e. , the transmission will fail. Fortunately, AIGC allows the resolution of the image to be specified before the image is generated. Therefore, the resolution of the image is determined by the communication capability or the communication resource.
[0055] Assuming the image is losslessly compressed and the color information of each pixel occupies a fixed number of bits (e.g., 24-bit RGB representation, 24 bits are needed for each pixel), the resolution of the image generated by AIGC (expressed in the total number of pixels) is , where is the number of bits per pixel. Considering that if the sensing data accuracy input to AIGC is not high, even if the image quality is high, it is invalid, the user's quality of service (CAQA) can be modeled as the product of the perception accuracy and the quality of the generated image, i.e.,
[0056]
[0057] , where is the display capacity of user device k, and are positively correlated with the perception energy and the communication energy, respectively. Therefore, it is jointly affected by the user's perception and communication energy allocation.
[0058] It should be understood that the above formula is only illustrative, and those skilled in the art can adjust it as needed. For example, the user's quality of service (CAQA) can be replaced by being calculated in the following way:
[0059] .
[0060] Four, joint perception and communication resource allocation scheme
[0061] In the AIGC system supporting ISAC, all users share the energy and spectrum resources of the transceiver of ISAC. Therefore, the perception and communication energy of each user must be carefully allocated to provide good quality of service. Therefore, the perception and communication energy allocation of all users is optimized to maximize the average quality of service (AvgCAQA) of users, i.e.,
[0062]
[0063] , where represents the perception energy related to the user device, represents the communication energy related to the user device, represents the perception energy related to user device k, represents the communication energy related to user device k, indicates the number of user devices participating in this resource allocation, indicates the personalized immersion coefficient of user device k, denotes the service quality of the standard corresponding to user k, s.t. denotes subject to, C1-C5 denote constraint conditions C1-C5, denotes the transmission time of the user equipment to generate an image, condition C1 denotes that the total perceived time of all user equipment plus the transmission time of any user equipment should be less than the maximum service time, denotes the number of rounds of perception signals in the perception resource allocated to user k, denotes the duration of each perception signal, denotes the maximum service time; condition C2 denotes that the total energy consumption should be less than the maximum energy of the transceiver ; condition C3 means that the perception energy allocated to each user equipment should not exceed the maximum perception energy , condition C4 requires that the service quality of the standard corresponding to each user should meet the personalized minimum requirement , condition C5 means that the resolution of the generated image should be less than the display capacity of user equipment k .
[0064] For the above optimization objectives, due to the nonlinear relationship between the perception energy and the accuracy, the solution space of the communication and perception resource variables increases exponentially with the number of users, so the solution space for resource allocation is very large, and it is a non-convex and non-deterministic polynomial difficulty (NP-Hard) problem. The existing resource allocation algorithm is easy to fall into a local optimal solution, and the effect is not good. If the perception energy is given, the allocation of communication energy can be transformed into a standard linear programming (Linear Programming, LP) problem. However, the perception energy allocation problem is still a non-convex and NP-Hard problem. Considering these characteristics of the SenComE-AIGC problem, combined with the fact that the deep reinforcement learning algorithm (Deep reinforcement learning, DRL) can handle non-convex problems, a novel perception and communication energy allocation algorithm is proposed, which adopts a hierarchical optimization framework and adopts a preset double-layer resource allocation mechanism. The top layer is the perception resource allocation layer, which uses the deep reinforcement learning with action filter (DRL-F) algorithm to optimize the perception energy allocation, and the bottom layer is the communication resource allocation layer, which uses a ranking-based heuristic algorithm to solve the communication energy allocation problem.
[0065] According to one embodiment of the present application, the preset double-layer resource allocation mechanism is configured to: first perform sensing resource allocation using a deep reinforcement learning algorithm with an action filter in the top layer, wherein the deep reinforcement learning algorithm is used to calculate the action of each user according to the state of all users and channels, the action filter is used to map the action of each user to the effective sensing energy range by setting the minimum sensing energy and the maximum sensing energy, and the sensing energy of each user is obtained, which can be converted into the number of sensing signals; the residual energy is calculated according to the total energy and the sensing energy of each user, and communication resource allocation is performed using a ranking-based communication energy allocation algorithm in the bottom layer, wherein the users are ranked in order from high to low according to the service quality of the users that can be improved per unit of communication energy; when the residual energy is less than the sum of the minimum communication energy required by all users, the minimum communication energy required by the users is satisfied in order from the user ranked first; and when the residual energy is greater than or equal to the sum of the minimum communication energy required by all users, each user is allocated at least the minimum communication energy, and then the maximum communication energy is allocated to each user in order from the user ranked first until the residual energy is allocated completely. The hierarchical optimization framework reduces the solution space by an exponential factor, and the complexity of the algorithm is linearly related to the total number of users K. This scheme innovatively decomposes the original NP-Hard problem into two polynomial-time solvable sub-problems, reduces the solution space of the original optimization problem by using a heuristic embedding method, and the deep reinforcement learning algorithm with an action filter mechanism can further reduce the solution space of the sensing resource allocation. Due to the reduction of the solution space, the double-layer resource allocation scheme can improve AvgCAQA by more than 50% compared with existing resource allocation schemes such as the CGQ algorithm.
[0066] According to one embodiment of the present application, the schematic structure of the LPDRL-F algorithm is as shown in Figure 5As shown, the method decomposes the optimization problem into two sub-problems, i.e., the sensing energy allocation and the communication energy allocation. First, a deep reinforcement learning (DRL) algorithm, i.e., soft actor-critic (SAC) algorithm with action filter (SAC-F), is proposed to allocate the sensing energy, and then the sensing accuracy is obtained through user sensing. With the sensing energy allocation scheme, the communication energy allocation problem can be proved to be a traditional linear programming problem. Then, the users are ranked according to their sensing accuracy and channel condition, and a ranking-based communication energy allocation algorithm (RCE) is proposed to obtain the optimal communication energy allocation scheme under the current sensing energy allocation. From a mathematical point of view, the complexity of RCE is linear with respect to the number of users, and it can be proved that it can derive the globally optimal solution. Subsequently, according to the sensing energy resource solution derived by the DRL-F algorithm and the communication energy resource solution derived by the RCE algorithm, the resolution of the AIGC service generated image is specified, and the AvgCAQA is obtained, which can be used as a reward to guide the training of the proposed LPSAC-F algorithm.
[0067] The LPSAC-F framework includes five key elements: agent, environment, state, action, and LP-guided reward. For the sensing and communication energy allocation problem, the definitions of these elements are as follows:
[0068] (1) Agent and environment: the transceiver of ISAC as the agent, and all components interacting with the transceiver of ISAC, including users, channels, energy, and AIGC, are considered as the environment.
[0069] (2) State: the state s consists of factors such as channel gain, which are very important for sensing and communication energy allocation. However, the original channel gain usually has a large span, making it difficult for deep reinforcement learning algorithms to handle directly. To solve this problem, a logarithmic transformation is adopted, which compresses the variation of channel gain into a more easily handled range. Illustratively, the state is represented as where is the normalized logarithmic channel gain corresponding to user k, defined as . Here, L is a normalization parameter based on the coverage radius of the transceiver of ISAC and the path loss model.
[0070] (3) Action: the action represents the sensing energy allocation proportion of K users, where is the output of a sigmoid activation function in the behavior network, ensuring its range is between 0 and 1. According to this operation, the perceived energy of user k is given by However, directly applying without additional processing can lead to invalid operations, thus violating the minimum CAQA constraint, i.e., Here, refers to the minimum perceived energy required to ensure that the minimum CAQA constraint condition is met when the quality of the generated image reaches its maximum value (i.e., 1). To address this issue, an action filter is introduced to convert into a valid range of perceived energy Illustratively, this conversion guarantees that when ; and when . After filtering out invalid operations, the perceived energy of user k is finally given by By mapping the original actions to a valid range, the decoding layer can effectively filter out invalid actions, thus reducing the action space and improving training efficiency.
[0071] (4) LP-guided reward: According to the perceived energy allocation scheme derived from the DRL-F algorithm , the communication energy allocation is determined by the RCE algorithm (the RCE algorithm will be introduced in detail later). Since the constraint condition C3 is satisfied in the action design, the constraint conditions C1, C2, and C5 are satisfied in the RCE algorithm. Therefore, the LP-guided reward is defined to ensure compliance with the C4 constraint condition. Illustratively, if all users meet the individualized minimum CAQA requirement, i.e., , the reward is the average quality of service (AvgCAQA) of the users. Otherwise, the reward will be penalized according to the number of users violating C4, thus reducing the reward. Preferably, the reward function is as follows:
[0072]
[0073] where denotes the number of users violating the constraint condition C4. This reward function encourages the LPDRL-F algorithm to prioritize solutions that maximize AvgCAQA while ensuring compliance with the minimum CAQA requirement.
[0074] Based on the perceived energy allocation solution obtained from the DRL-F algorithm, it is found that the original problem, i.e., the SenComE-AIGC problem, is transformed into a communication energy allocation problem, i.e.:
[0075]
[0076] Among them, constraint M1 originates from the original constraints C1, C4, and C5 in the SenComE-AIGC problem, and M2 originates from the original constraint C2. M1 specifies the minimum communication energy required to satisfy all users. This ensures that each user meets CAQA requirements (C4). Furthermore, it ensures that each user's communication energy does not exceed their maximum communication energy limit. This is because exceeding this limit will cause the communication time plus the total sensing time to exceed the service time (C1), or exceed the device display capacity (C5). M2 ensures that the total energy used for communication does not exceed the remaining energy. (Total energy minus total induced energy).
[0077] Since both the objective and constraints are linear, the ComE-AIGC problem is a standard linear programming problem. A sorting-based communication energy allocation algorithm, namely the RCE algorithm, is proposed to find the optimal solution to this problem. Furthermore, the algorithm's complexity is linearly related to the number of users, significantly reducing the solution space of the original problem. The pseudocode of the RCE algorithm is shown below:
[0078]
[0079] Communication energy allocation is a standard linear programming problem. First, allocate the remaining energy. This is to meet the minimum communication energy requirements of each user. If there is any extra energy remaining... According to Distribute in descending order. This represents the improvement in the quality of service for user k per unit of communication energy, i.e. This represents the CAQA achievable by user k units of energy. The maximum communication energy allocated to each user is... Once this limit is reached, the next user in the sequence will be considered until all energy is allocated or all users have reached their maximum communication energy.
[0080] To verify the effectiveness of this invention, the proposed joint sensing and communication energy allocation scheme was also simulated. The simulation parameters are shown in Table 1, and the simulation results are as follows: Figures 6-8 As shown.
[0081] Table 1. Simulation Parameter Description of Joint Sensing and Communication Energy Allocation Scheme
[0082]
[0083] To evaluate the performance of the proposed LPSAC-F, it was compared with the following four algorithms:
[0084] (1) MSHC: Assign maximum sensing energy, then use RCE algorithm to obtain optimal communication energy allocation.
[0085] (2) CGQ: Randomly assign sensing energy, then use RCE algorithm to obtain communication energy allocation only oriented to content generation quality.
[0086] (3) HDRL: Use DRL without action filter to solve sensing energy allocation problem, use RCE algorithm to solve communication energy allocation problem.
[0087] (4) UDRL-F: Use DRL with action filter to solve sensing energy allocation problem, after meeting the minimum SA-AQE requirement, evenly allocate the remaining energy to communication.
[0088] Figure 6 The relationship between average CAQA and maximum total energy is shown. As the total energy increases, both sensing energy and communication energy become more, thus improving the average CAQA of all algorithms. The LPHDRL-F algorithm is always superior to other algorithms. LPDRL-F is superior to UDRL-F because it adopts a linear programming guided method, optimizes the allocation of communication resources according to the current sensing energy allocation, and uses the solution to guide the training of the algorithm. One reason why LPDRL-F surpasses LPDRL is the action filtering mechanism, which narrows down the solution space of sensing energy allocation, making it easier to find the optimal solution. In addition, as the total energy increases, the performance gap between LPDRL-F and MSHC is also narrowing, because in the case of limited resources, effective resource management can bring greater benefits. Specifically, in the case of energy constraints, allocating maximum sensing energy may lead to insufficient communication energy, thus reducing performance. Similarly, when energy is limited, the performance of CGQ is superior to MSHC.
[0089] Figure 7 It is shown that as the number of users increases, the average CAQA of all algorithms will decrease due to the reduction of resources (such as energy and communication bandwidth) available to each user. LPDRL-F is always superior to other algorithms. For example, in the case of 10 users, the average CAQA of LPDRL-F is improved by 77.76% and 55.54% compared with MSHC and CGQ, respectively. These advantages are due to the ability of LPDRL-F to dynamically allocate sensing and communication energy according to channel conditions and user ranking, optimizing limited resources to achieve higher CAQA. In contrast, the fixed sensing energy allocation of MSHC and the random sensing energy allocation of CGQ (only oriented to content quality) will result in subpar performance.
[0090] Figure 8The average CAQA performance of LPDRL-F is shown as a function of the maximum total energy under different maximum service times. As the energy increases, the performance gap between different service times widens. This is because when the energy is low, the CAQA is mainly limited by the availability of energy, and increasing the maximum service time has little effect on the average CAQA due to insufficient energy. However, as the energy increases, the available energy also increases, and the extension of the service time will make the average CAQA improve more obviously. For example, when the maximum total energy is , , the average CAQA of increases by 29.93%. This is because in the case of sufficient energy, the average CAQA will be limited by the service time.
[0091] In summary, some embodiments of the present application realize the system-level design of the ISAC-driven AIGC service technology system through the collaborative innovation of the ISAC-driven AIGC network architecture, the integrated frame structure design, the CAQA performance evaluation system, and the hierarchical resource allocation scheme. The use of the sensing capability of ISAC reduces the risk of privacy leakage, and the rationality evaluation of the quality of service of ISAC-driven AIGC is realized based on the new indicators related to the sensing accuracy and the quality of the generated image. Through the double-layer resource optimization framework, the solution space of the original NP-Hard problem is compressed from exponential complexity to linear complexity, and the solving efficiency of AvgCAQA is improved by more than 50% compared with traditional algorithms (such as CGQ algorithm). This technology system provides a more efficient, intelligent, and secure wireless communication solution for digital human live streaming, VR / AR interaction, and other scenarios, significantly improving the real-time performance, accuracy, and user experience of AIGC services.
[0092] It should be noted that although the above describes the steps in a specific order, it does not mean that the steps must be performed in the above specific order. In fact, some of these steps can be executed concurrently, or even in a different order, as long as the desired function can be achieved.
[0093] The present application can be a system, a method, and / or a computer program product. The computer program product can include a computer readable storage medium having computer readable program instructions loaded thereon for causing a processor to implement various aspects of the present application.
[0094] A computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves.
[0095] Embodiments of the application have been described above, with the understanding that these embodiments are exemplary only, and are not restrictive, in terms of the scope of the embodiments disclosed. Many modifications and variations of the described embodiments are possible, in light of the above teachings, without departing from the scope and spirit of the described embodiments. The choice of words in this document is intended to best explain the principles of the embodiments, practical application, or technical improvement in the art, or to enable others skilled in the art to utilize the embodiments disclosed herein.
Claims
1. A resource allocation method for allocating sensing and communication resources to an AI-generated content service based on integrated sensing and communication technologies, the method comprising: Within the coverage area of the transceiver integrating sensing and communication functions, obtain the display capacity and channel-related status of multiple user devices participating in this resource allocation; Based on the display capacity of multiple user devices and the channel-related status, with the optimization objective of maximizing the average service quality for users, the sensing resources and communication resources to be allocated to multiple user devices are iteratively adjusted to obtain the allocation decision. Among them, the service quality of each user is related to the sensing accuracy of the user posture information obtained through sensing resources and the quality of the generated image corresponding to the digital human driven by human posture obtained through the user device.
2. The method according to claim 1, characterized in that, The service quality for each user is the product of the standard service quality and the personalized immersion coefficient that the user requires for their desired experience.
3. The method according to claim 1, characterized in that, The optimization objective is set as follows: in, Represents the sensing energy related to user equipment. Represents the communication energy related to user equipment. Represents the sensing energy related to user equipment k. Represents the communication energy associated with user equipment k. This indicates the number of user devices participating in this resource allocation. This represents the personalized immersion coefficient of user device k. This represents the standard service quality corresponding to user k, st indicates that it is subject to constraints, and C1-C5 represent constraints C1-C5. This represents the transmission time of the generated image by the user equipment. Condition C1 states that the total sensing time of all user equipment plus the transmission time of any user equipment should be less than the maximum service time. This represents the number of rounds of sensing signals allocated to user k from the sensing resources. This indicates the duration of each sensed signal. Indicates the maximum service time; condition C2 indicates that the total power consumption should be less than the transceiver's maximum power. Condition C3 means that the sensing energy allocated to each user device should not exceed the maximum sensing energy. Condition C4 requires that the standard quality of service for each user device should meet the minimum personalized requirements. Condition C5 refers to the resolution of the generated image. It should be less than the display capacity of user equipment k. .
4. The method according to claim 1, characterized in that, The maximum perception accuracy for each user does not exceed the upper limit of the highest perception accuracy, and is positively correlated with the number of rounds of perception signals assigned to them.
5. The method according to claim 4, characterized in that, The perception accuracy for each user is calculated as follows: ,or ,or , in, This represents the maximum perception accuracy limit for user k. The gain sparsity indicates the number of rounds of the perceived signal corresponding to user k. This represents the number of rounds of the sensing signal corresponding to user k. The sensitivity coefficient represents the number of rounds of the perceived signal corresponding to user k. It is a positive constant.
6. The method according to claim 1, characterized in that, The quality of the generated image for each user is equal to the resolution of the generated image divided by the display capacity of the user's device, and the quality of the generated image is still considered to be 1 even if it exceeds 1.
7. The method according to any one of claims 1-6, characterized in that, When allocating sensing and communication resources, a preset two-layer resource allocation mechanism is adopted, which is configured as follows: First, a deep reinforcement learning algorithm with an action filter is used at the top layer to allocate perception resources. The deep reinforcement learning algorithm calculates the action of each user based on the state of all users related to the channel. The minimum and maximum perception energy set by the action filter are used to map the action of each user to the effective perception energy range, so as to obtain the perception energy allocated to each user. The perception energy can be converted into the number of rounds of perception signal. The remaining energy is calculated based on the total energy and the perceived energy allocated to each user. At the underlying level, a sorting-based communication energy allocation algorithm is used to allocate communication resources. Users are sorted in descending order of the service quality that can be improved by obtaining each unit of communication energy. When the remaining energy is less than the sum of the minimum communication energy required by all users, the minimum communication energy required by users is satisfied sequentially, starting from the user at the top of the sort. When the remaining energy is greater than or equal to the sum of the minimum communication energy required by all users, each user is allocated at least the minimum communication energy, and then, starting from the user at the top of the sort, each user is allocated the maximum communication energy, until the remaining energy is completely allocated.
8. A method for controlling the movement of a digital human based on perceived user posture, characterized in that, include: The display capacity and channel-related status of multiple user devices participating in this resource allocation are obtained, and an allocation decision is obtained using the resource allocation method as described in any one of claims 1-7; Image generation and downlink are performed based on an integrated frame structure for allocation decision and communication sensing. The integrated frame structure for communication sensing is configured as follows: During the sensing phase, a transceiver is used to transmit a frequency-modulated continuous wave, and the entire bandwidth is used in a time-division multiplexing manner to perform attitude sensing for each user separately, thereby obtaining attitude sensing data for each user. Each user performs attitude sensing by using the number of rounds of the assigned sensing signal and the duration of each sensing signal to determine the sensing time. After the sensing time for each user ends, a protection interval is set. During the communication phase, the server obtains the generated image after determining the control digitizer's actions based on each user's posture sensing data. The control transceiver then uses orthogonal frequency division multiplexing (OFDM) or non-orthogonal multiple access (NOAMA) technology to transmit the generated image to the corresponding user equipment for display.
9. A computer-readable storage medium, characterized in that, It stores a computer program that can be executed by a processor to implement the steps of the method according to any one of claims 1-8.
10. An electronic device, characterized in that, include: One or more processors; as well as Memory, wherein the memory is used to store executable instructions; The one or more processors are configured to implement the steps of the method according to any one of claims 1-8 by executing the executable instructions.