A Synesthesia-Integrated Waveform Adaptive Resource Allocation Method and System

By introducing a scene perception weight adjuster and a deep reinforcement learning agent into the integrated sensing system, the weights of communication and perception are dynamically adjusted, solving the problem that resource allocation schemes cannot track nonlinear demand changes, and improving communication throughput and the reliability of perception detection.

CN122496910APending Publication Date: 2026-07-31NORTHEASTERN UNIV AT QINHUANGDAO
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NORTHEASTERN UNIV AT QINHUANGDAO
Filing Date
2026-06-15
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In existing integrated sensing systems, the communication and sensing resource allocation schemes cannot track the nonlinear demand changes driven by the kinematic state of the sensing target in real time, resulting in sensing and detection failures or resource waste in dynamic scenes.

Method used

A scene-aware weight regulator is introduced to improve the weights of communication and perception in the reward function from fixed hyperparameters to state-dependent variables driven by the kinematic state of the perceived target in real time. By optimizing resource allocation through a deep reinforcement learning agent and combining the channel quality on the communication side and the kinematic state parameters on the perception side, dynamic weight adjustment is achieved.

Benefits of technology

While ensuring a perception and detection probability of over 95%, the communication throughput was increased by approximately 32%, and cross-domain resource optimization was achieved, adaptively tracking physical layer requirements and improving system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122496910A_ABST
    Figure CN122496910A_ABST
Patent Text Reader

Abstract

This invention discloses a waveform adaptive resource allocation method and system for integrated sensing and communication, relating to the field of wireless communication and radar sensing fusion technology. The method constructs a joint state vector that integrates the communication service state and the kinematic state of the sensed target. It outputs the communication and sensing function allocation scheme for each subcarrier and time slot through a deep reinforcement learning agent. The method also dynamically adjusts the weights of communication and sensing in the reward function according to the kinematic state parameters of the sensed target through a scene perception weight adjuster. This allows the optimization objective of resource allocation to automatically track the nonlinear changes in physical layer requirements and maximize communication throughput while ensuring the probability of sensing detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of deep fusion technology of wireless communication and radar sensing, and in particular to a waveform adaptive resource allocation method and system for integrated sensing and communication. Background Technology

[0002] Integrated sensing technology is one of the key enabling technologies for sixth-generation mobile communication systems. This technology achieves spectrum resource and hardware platform sharing by simultaneously carrying both communication data transmission and radar sensing detection functions on the same orthogonal frequency division multiplexing waveform. In application scenarios such as vehicle-to-everything (V2X), low-altitude economy, and smart cities, base stations need to simultaneously provide data services to communication users and detect and locate surrounding targets on the same set of time-frequency resources. However, communication and sensing have inherently contradictory requirements for waveform parameters: communication prioritizes high data rates, tending to use narrow subcarrier spacing to accommodate more subcarriers and improve spectral efficiency, and dynamically selecting modulation and coding schemes based on each user's channel conditions to approximate channel capacity; while sensing prioritizes high range resolution and long-range detection capabilities, requiring wide signal bandwidth and high power concentration to ensure sufficient echo signal-to-noise ratio for detecting weak targets at long distances, and requiring sensing subcarriers to be as continuous as possible in the frequency domain to reduce range sidelobes. These conflicting requirements cannot be fully satisfied simultaneously under limited time-frequency resource constraints. How to rationally allocate the resource ratio between communication and sensing, and maximize system-level utility while ensuring the performance constraints of both, has become the core bottleneck restricting the performance of integrated sensing systems.

[0003] Existing integrated sensing resource allocation schemes mainly fall into two categories. The first category is the static allocation scheme, in which the system designer presets the resource ratio between communication and sensing based on the characteristics of the scenario and keeps it unchanged during operation. For example, 70% of the subcarriers are fixedly allocated to communication and 30% to sensing. This type of scheme is simple to implement but cannot adapt to dynamic changes in the scenario.

[0004] International patent application WO2023230757A1 discloses an autonomous sensing resource allocation method in a sensing-integrated system. This method employs a sensing signal resource allocation mechanism based on control channel signaling. After detecting a sensing target, the base station autonomously allocates broadband sensing signal resources on demand to avoid resource waste and inter-cell interference caused by static allocation. However, the resource allocation decision of this scheme is driven by signaling rather than intelligent learning, lacking real-time adaptive capability to changes in communication service load and dynamic sensing targets. More importantly, this scheme treats the ratio of communication to sensing resources as a fixed parameter preset by the network configuration, failing to recognize the highly nonlinear physical coupling relationship between the urgency of sensing needs and the kinematic state of the target in dynamic scenarios such as vehicle-to-everything (V2X) networks.

[0005] The second category is dynamic allocation schemes based on intelligent learning, which use methods such as deep reinforcement learning to replace static rules to achieve adaptive resource allocation. This type of scheme models the resource allocation problem as a Markov decision process, where the agent interacts with the integrated sensory environment and learns the optimal allocation strategy. However, existing schemes in this category still set the weight coefficients of communication performance and perception performance as fixed hyperparameters in the reward function design, preset by the system designer based on experience. In real dynamic scenarios, the kinematic states of the perceived target, such as distance and relative velocity, are constantly changing, while the radar echo signal-to-noise ratio decreases inversely proportional to the target distance (fourth power). This means that when the target distance is halved, the resources required for perception may increase by more than sixteen times. Fixed-weight schemes cannot track the nonlinear jumps in perception requirements when the target's kinematic state changes rapidly, leading to a two-way mismatch: wasted perception resources in low-threat scenarios and failed perception detection in high-threat scenarios. This decouples the agent's optimization direction from actual physical needs.

[0006] In summary, the core technical challenge of existing integrated sensing resource allocation technologies can be summarized as: how to ensure that the allocation ratio of time-frequency resources between communication and sensing can track in real time the nonlinear demand changes driven by the kinematic state of the sensed target and the physical laws of radar equations, thereby maximizing communication throughput under the constraint of sensing detection probability. Solving this problem requires breaking away from the traditional understanding of treating weights as fixed hyperparameters and establishing a dynamic coupling relationship between weights and the target's kinematic state based on physical mechanisms. Summary of the Invention

[0007] To address the core bottleneck of existing integrated sensing systems' resource allocation schemes, which cannot track the nonlinear resource demand changes driven by the kinematic state of the sensing target in real time, this invention provides an integrated sensing waveform adaptive resource allocation method and system. By introducing a scene perception weight adjuster, the weights of communication and perception in the reward function are increased from fixed hyperparameters to state-dependent variables driven by the kinematic state of the sensing target in real time. Under the constraint of ensuring the probability of perception detection, the method maximizes communication throughput from the principle that the optimization objective of resource allocation can adaptively track the actual needs of the physical layer.

[0008] The technical solution of this invention is: a sensing-integrated waveform adaptive resource allocation method, applied to a sensing-integrated system using orthogonal frequency division multiplexing waveforms, comprising the following steps: constructing a joint state vector, the joint state vector including the channel quality index and service load index of the current communication service, and the kinematic state parameters of the sensing target; inputting the joint state vector into a deep reinforcement learning agent, the deep reinforcement learning agent outputting an allocation scheme for the communication function and sensing function of each subcarrier in each time slot; determining the dynamic weights of the communication performance index and the sensing performance index in the reward function through a scene perception weight adjuster based on the kinematic state parameters of the sensing target, wherein the dynamic weights increase the sensing weight as the threat level score of the sensing target increases and increase the communication weight as the threat level score of the sensing target decreases; calculating the reward value of the deep reinforcement learning agent based on the dynamic weights, and updating the policy parameters of the deep reinforcement learning agent based on the reward value.

[0009] This invention also provides a sensing-integrated waveform adaptive resource allocation system, applied to a sensing-integrated system employing orthogonal frequency division multiplexing waveforms, comprising: a state construction module configured to construct a joint state vector containing channel quality indicators and service load indicators of communication services, as well as kinematic state parameters of the sensing target; a resource allocation module configured to output a communication and sensing function allocation scheme for each subcarrier in each time slot based on the joint state vector using a deep reinforcement learning agent; a scene perception weight adjustment module configured to determine the dynamic weights of communication performance indicators and perception performance indicators in the reward function based on the kinematic state parameters of the sensing target; and a policy update module configured to calculate the reward value based on the dynamic weights and update the policy parameters of the deep reinforcement learning agent.

[0010] The beneficial effects of this invention include:

[0011] First, this invention transforms the weight coefficients of communication and perception in the reward function from fixed hyperparameters to state-dependent variables driven in real-time by the kinematic state of the perceived target through a scene-aware weight adjuster. This allows the optimization direction of the deep reinforcement learning agent to automatically track the nonlinear changes in physical layer requirements. Compared to existing fixed-weight schemes, communication throughput is increased by approximately 32% while maintaining a perception detection probability greater than 95%. The mechanism lies in the fact that the radar echo signal-to-noise ratio decreases inversely proportional to the target distance (fourth power). Fixed weights cannot reflect this nonlinear physical characteristic. However, the weight adjuster of this invention directly injects the physical constraints of the radar equations into the optimization target, eliminating the decoupling between the optimization direction and physical requirements—a feature not considered in existing schemes.

[0012] Second, the joint state vector simultaneously integrates communication-side channel quality indicators, service load indicators, and perception-side target kinematic state parameters. This, combined with the fine-grained allocation of action space at the subcarrier multiplication time slot level, enables physical layer information-driven cross-domain resource optimization. The mechanism lies in the fact that the communication state provides the agent with an instantaneous snapshot of resource requirements, while the perception state provides a physical measure of threat level scoring. The fusion of these two allows the agent to simultaneously consider communication service quality and perception detection reliability in a single decision, generating a synergistic gain far exceeding the sum of the effects of independent optimization in each dimension.

[0013] Third, the weight regulator and the deep reinforcement learning agent form a closed-loop collaboration: the weight regulator injects physical layer requirements into the optimization target to drive the agent to finely allocate resources in the correct direction. The perception echo feedback after allocation updates the perception state and then readjusts the weights. This closed-loop mechanism enables the system to continuously and adaptively track dynamic environmental changes without human intervention. Attached Figure Description

[0014] Figure 1 This is a flowchart illustrating the integrated sensing waveform adaptive resource allocation method provided in this embodiment of the invention.

[0015] Figure 2 This is a schematic diagram of the structure of the integrated sensing waveform adaptive resource allocation system provided in an embodiment of the present invention. Detailed Implementation

[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] See Figure 1 This invention provides an integrated sensing waveform adaptive resource allocation method, applied to an integrated sensing system employing orthogonal frequency division multiplexing (OFDM) waveforms. The system includes a base station, several communication users, and several sensing targets. The base station employs a waveform adaptive resource allocation method with... The orthogonal frequency division multiplexing waveform of each subcarrier, Communication data transmission and radar sensing / detection tasks are performed simultaneously within a consecutive time slot scheduling period. Each subcarrier is allocated for either communication or sensing functions in each time slot, and the two functions are mutually exclusive. The core method of this invention includes the following steps.

[0018] Step S1: Construct the joint state vector. In each scheduling slot... At the start of the process, the base station collects current state information from both the communication and sensing sides to construct a joint state vector. The channel quality indicators on the communication side specifically include the channel quality indication value for each subcarrier. This value is periodically reported by communication users through the uplink feedback channel, reflecting the quality of channel conditions for each subcarrier under the current propagation environment. The service load indicators on the communication side include the current length of the communication data buffer queue, i.e., the number of data packets waiting to be sent but not yet scheduled. This value is maintained and updated in real time by the base station's media access control layer. The kinematic state parameters on the sensing side include the following three components: the estimated distance between the sensing target and the base station. ,in Index for the sensed target; estimated radial velocity of the sensed target relative to this base station. Estimated radar cross-section of the target. These three parameters are obtained by the base station's sensing processing module through range-Doppler spectrum analysis of the sensed echo signal during the previous scheduling cycle. When the environment simultaneously contains... When sensing multiple targets, the kinematic state parameters of each target are incorporated into the joint state vector. In the specific implementation, the channel quality indicator value on the communication side is obtained through channel state information feedback messages periodically reported by the communication users, with the reporting period aligned with the scheduling time slot interval. For each communication user, the channel quality indicator value is one... The dimensional vector reflects the signal-to-interference-plus-noise ratio (SINR) level of the user on each subcarrier. When the system serves multiple communication users, the channel quality indicator values ​​of all users are concatenated in user index order. The service load indicator is obtained as follows: the base station's packet data aggregation protocol layer reads the data buffer queue depth of each communication user at the beginning of each scheduling time slot, quantizes it in units of packets or bytes, and encodes it as a scalar value.

[0019] The acquisition of kinematic state parameters on the sensing side depends on the sensing processing results of the previous scheduling cycle. Distance parameters The velocity parameters are determined by the peak position of the range spectrum after pulse compression of the sensed echo. The radar cross-section is obtained by converting the frequency shift after Doppler processing. The radar equations are derived by inverting the echo amplitude and known system parameters. The estimation accuracy of these three parameters directly affects the reliability of subsequent threat level scoring. In multi-target scenarios, the perception processing module distinguishes different targets and estimates their kinematic parameters separately through constant false alarm rate detection and clustering algorithms.

[0020] The state information from the communication side and the sensing side is concatenated into a joint state vector according to a predefined order. The total dimension of this vector dynamically changes with the number of communication users, subcarriers, and sensing targets. To adapt to the dimensionality change, the input layer of the deep reinforcement learning agent employs a padding and masking mechanism to pad the actual state vector to a fixed maximum dimension and mark the padded positions as invalid to avoid affecting network computation. Joint State Vector As a deep reinforcement learning agent in time slots Input observation.

[0021] Step S2: The scene-aware weight adjuster determines the dynamic weights. This step is the core innovation of the method of this invention. The scene-aware weight adjuster receives the kinematic state parameters of the perceived target output in step S1 and determines the dynamic weights of the communication performance index and the perception performance index in the current time slot reward function.

[0022] First, regarding the first For each perceived target, a threat level score is calculated based on its kinematic state parameters. In an integrated sensing system, the radar echo signal-to-noise ratio is inversely proportional to the fourth power of the target distance and directly proportional to the target's radar cross-section. When a target approaches the base station at high speed, the increased Doppler spread compresses the coherent accumulation time window required for sensing processing, further exacerbating the urgency of sensing resources. Based on the above physical mechanism, the threat level score is calculated as follows:

[0023] ,

[0024] in: For the first The threat level score for each perceived target is a dimensionless scalar, with a value range of [value range missing]. The value is calculated by this formula and represents the degree of urgency of the target's need for sensing resources. The larger the value, the higher the urgency of the perception. For the first The estimated radar cross-section of each perceived target is a scalar, with a value ranging from 1 to 2. The unit is The echo amplitude is estimated by the sensing and processing module, reflecting the electromagnetic scattering characteristics of the target. The larger the cross-sectional area, the stronger the echo energy. For the first The absolute value of the estimated radial velocity of each sensed target relative to the base station is a scalar, and its value ranges from [value missing]. The unit is The Doppler frequency shift is estimated; the higher the speed, the greater the Doppler spread, and the more urgent the sensing processing window becomes. For the first The estimated distance between each perceived target and the base station is a scalar, with a value range of [value missing]. The unit is The distance is estimated from the peak position of the distance spectrum. The closer the distance, the higher the echo signal-to-noise ratio, but the greater the risk of collision. express The fourth power of embodies the inverse fourth power attenuation law of echo signal-to-noise ratio and range in the radar equation; The reference normalization factor is denoted as , which is a scalar and takes the value . ,in , , This factor ensures that the threat level score is exactly 1 under the reference conditions, which facilitates the parameter setting of the subsequent nonlinear mapping function.

[0025] Secondly, the threat level score is mapped to a perception weight value using a monotonically increasing non-linear activation function. This mapping needs to satisfy the following physical constraints: when the threat level score approaches zero, the perception weight should be close to its minimum value to prioritize resources for communication; when the threat level score increases significantly, the perception weight should saturate to its maximum value to ensure the reliability of perception detection. Based on the above constraints, the perception weight value is calculated as follows:

[0026] ,

[0027] in: For time slots The perceived weight value is a scalar, and its value range is... , dimensionless, is calculated by this formula and represents the weight coefficient of the perception performance index in the current time slot reward function; The lower bound of the perception weight is a scalar with a value of 0.1. It is dimensionless and determined by system design constraints to ensure that perception obtains at least the minimum resources in any scenario. If the value is too small, perception will completely lose its continuous monitoring capability in low-threat scenarios. The upper bound of the perception weight is a scalar with a value of 0.9. It is dimensionless and determined by system design constraints to ensure that communication obtains at least the minimum resources in any scenario. If the value is too large, the communication service will be completely interrupted in high-threat scenarios. Let be the slope factor of the activation function, and let be a scalar with a range of values. The preferred value is 5, dimensionless, controlling the steepness of the transition from low to high weight values. The larger the value, the steeper the transition, meaning the more sensitive the weight is to changes in threat level. If the value is too small, the weight adjustment will be sluggish. If the value is too large, it can easily cause weight oscillations; For time slots The comprehensive threat level score is calculated by taking the maximum score of each target when there are multiple perceived targets. Let be the offset factor of the activation function, and let be a scalar with a range of values. The preferred value is 1, dimensionless, and the center position of the control weight transition, i.e., the threat level score reaches... The time-aware weight is exactly at the midpoint between the upper and lower bounds; It is a natural exponential function. Communication weight value. Depend on Calculated.

[0028] When multiple targets exist simultaneously in the environment, the overall threat level score is the maximum value of the threat level scores of each target.

[0029] ,

[0030] in: For time slots The overall threat level score is a scalar, with a value range of [value range missing]. , dimensionless, is calculated by this formula, and represents the threat level of the most urgent of all perceived targets; The total number of sensed targets in the current environment is a positive integer, with a value range of 1. It is maintained by the perception and detection module; This is the maximum value operator. The design consideration for using the maximum value instead of the mean or summation is that in security-critical scenarios, the failure to detect a high-threat target can lead to serious consequences, so resource allocation must be based on the needs of the highest threat target.

[0031] Furthermore, to suppress weight oscillations caused by fluctuations in the target's kinematic state in the threat level score, a smoothing regularization process based on the target's kinematic continuity constraint is applied to the perception weight values ​​between consecutive time slots. In the physical world, the target's motion state is constrained by its maximum acceleration; therefore, there is a physical upper bound on the reasonable variation range of the threat level score between adjacent time slots. The constraints of the smoothing regularization are as follows:

[0032] ,

[0033] ,

[0034] in: It represents the absolute value of the change in sensing weight between adjacent time slots; The upper bound of the physically reachable weight change is a scalar, the value of which is determined by subsequent formulas. It is dimensionless and is calculated by this formula, representing the maximum allowable change in the sensing weight value within a time slot interval. Let be the maximum expected acceleration of the perceived target, and let be a scalar with a range of values ​​of . Preferred The unit is In the context of vehicle networking, the typical acceleration metric for emergency braking of a vehicle is determined by the application scenario. If the value is too small, it will over-constrain the weight changes and lead to a slow response. If the value is too large, it will weaken the smoothing effect. Let be the scheduling time slot interval, and be a scalar with a value range of . Preferred The unit is The frame structure of the orthogonal frequency division multiplexing system is determined by the frame structure of the orthogonal frequency division multiplexing system. Same as previously defined; factor This reflects the amplification effect of the fourth power of distance term on velocity changes in the threat level score. When the calculated... When the above constraints are violated, it will be truncated to .

[0035] The aforementioned smoothing regularization process, while suppressing weight oscillations, also introduces an additional effect: when false targets or clutter interference occur in the perception estimation, the resulting abrupt change in the threat level score will exceed the upper bound of the physical reachability variation. The weight truncation mechanism automatically prevents false threats from impacting resource allocation, indirectly achieving credibility screening of perception results.

[0036] Step S3: The deep reinforcement learning agent outputs a resource allocation scheme. The deep reinforcement learning agent receives the joint state vector output in step S1. The system outputs the allocation scheme for communication and sensing functions of each subcarrier in each time slot. The agent employs an actor-critic network architecture, where the actor network outputs the allocation scheme based on the joint state vector, and the critic network evaluates the state-action value function based on the joint state vector and the allocation scheme. A scene-aware weight adjuster operates as a third network module independent of the actor and critic networks. The three network modules jointly optimize during training but each maintains its own independent parameter set.

[0037] The actor network's specific structure consists of one input layer, three hidden layers, and one output layer. The input layer receives the joint state vector. Each hidden layer contains 256 neurons and uses a modified linear activation function. The output layer contains... Several neurons are used, employing the sigmoid activation function to ensure that each output value lies within the interval [missing information]. Inside. The input to the critic network consists of the joint state vector. The probability vector of the actor network output The network is constructed by concatenating these components and passing them through three hidden layers (with the same structure as the actor network) to output a scalar state-action value estimate. During training, the parameters of the actor network are updated using the policy gradient method, while the parameters of the critic network are updated by minimizing the temporal difference error.

[0038] Specifically, the actor's online output has one dimension. The probability vector of allocation , of which element Indicates the first Each subcarrier is assigned a probability for sensing function in the current time slot. When If the decision threshold of 0.5 is exceeded, the subcarrier is assigned to the sensing function; otherwise, it is assigned to the communication function. During the training phase, to facilitate policy exploration, [further details are needed]. Random sampling of the parameters from a Bernoulli distribution is used to generate the actual allocation decision; during the inference phase, a threshold decision is directly applied to obtain a deterministic allocation scheme. The allocation scheme is represented by a binary matrix. It means that among them Indicates the first Subcarriers in time slots Perform perception functions. This indicates the execution of communication functions. Subcarriers assigned to communication functions are proportionally and fairly scheduled according to the channel quality indication values ​​of each user, while subcarriers assigned to sensing functions transmit linear frequency modulated pulse compressed waveforms to achieve distance and Doppler estimation.

[0039] To achieve forward-looking resource pre-allocation, step S3 also includes the following sub-steps: based on the current kinematic state parameters of the perceived target and the threat level score after smoothing and regularization in step S2, the future kinematic prediction model is used to infer the target's kinematics state. The predicted target position and velocity are calculated for each time slot. A constant velocity model is used in the kinematic prediction model, assuming the target maintains its current velocity throughout the prediction time domain. Based on the predicted position and velocity, the predicted echo signal-to-noise ratio (SNR) for each future time slot is calculated using radar equations. The calculation method for the predicted echo SNR is as follows:

[0040] ,

[0041] in: For time slots The predicted echo signal-to-noise ratio is a scalar with a value range of . , dimensionless (linear value), is calculated by this formula and characterizes the echo detection capability at the predicted distance; This is the predicted time slot index, which is a positive integer with a value range of 1. , To predict the time domain length, it is preferable to ; The total transmit power allocated to the sensing function is a scalar, with units of . ; denoted as antenna gain, which is a scalar, dimensionless (linear value). Where is the carrier wavelength, and is a scalar, with units of . ; Defined as the target radar cross-section as before; For the first The target is in the time slot The predicted distance, in units of ; This is the spatial attenuation constant in the radar equations; Let be the noise power spectral density, and be a scalar quantity in units of . ; The bandwidth of the sensed signal is a scalar, and the unit is . .

[0042] The predicted echo signal-to-noise ratio is lower than the detection threshold. The future time slots are marked as time slots with scarce sensing resources. Detection threshold. Based on the constant false alarm rate (CFRR) detection criterion, when the false alarm probability is... Typical values ​​under the given conditions The tagging information of resource scarcity time slots is encoded into a... A two-dimensional vector is appended to the joint state vector. At the end, it enables deep reinforcement learning agents to perceive future resource demand trends when making decisions.

[0043] When an agent observes in the current time slot that sensing resources will become scarce in future time slots, its policy network can learn a behavioral pattern of reserving additional sensing subcarrier resources in the current time slot in advance, thereby upgrading from reactive allocation to predictive allocation. This look-ahead mechanism effectively avoids the lag effect of reactive policies under the fourth power inverse proportionality law of the radar equation: in scenarios where the target approaches at high speed, it is too late to adjust resource allocation when the signal-to-noise ratio of the echo in the current time slot is insufficient, while the look-ahead mechanism allows the agent to gradually increase the reservation of sensing resources while the signal-to-noise ratio is still sufficient.

[0044] Step S4: Perception-Assisted Communication Beam Prediction. The target kinematics prediction model in Step S3 outputs the predicted target position for future time slots. In a sensor-communication integrated system, the communication user and the perceived target are often the same physical entity or located in the same physical area. Taking a vehicle-to-everything (V2X) scenario as an example, the base station simultaneously provides communication data services to moving vehicles and performs radar perception detection on them; in this case, the communication user is the perceived target. Therefore, the position prediction information from the perception side can be directly reused by the communication side, without the communication side needing to independently maintain the target motion model. Specifically, the predicted target position is converted into the corresponding predicted communication beam angle value:

[0045] ,

[0046] in: For time slots The predicted value of the communication beam angle is a scalar, and its value range is [value range missing]. The unit is The azimuth angle that the communication beam should point to, calculated by this formula, represents the azimuth angle that the beam should point to. and The first The target is in the time slot The Cartesian components of the predicted horizontal and vertical positions, in units of The distance is obtained by converting the predicted distance and current azimuth angle in polar coordinates using a kinematic prediction model. It is the arctangent function.

[0047] In the time slot preceding a time slot where sensing resources are scarce, the communication side pre-aligns the beam towards the predicted angle direction, thereby reducing beam alignment delay. The smoothing regularization process introduced in step S2 makes the input of the kinematic prediction model (the smoothed target motion state) more stable, indirectly improving the accuracy of beam angle prediction. This sensing-assisted communication beam prediction mechanism is a natural emergent effect of the preceding innovative steps: the kinematic prediction model was originally introduced to solve the problem of forward allocation of sensing resources, but its output can be reused by the communication side, and the smoothing regularization also brings additional gains to the communication side in terms of improving prediction stability.

[0048] Step S5: Reward Calculation and Policy Update Based on Dynamic Weights. Based on the dynamic weights output in Step S2 and the execution results of the allocation scheme output in Step S3, calculate the reward value for the deep reinforcement learning agent. The reward function is defined as a dynamic weighted sum of normalized communication throughput and perception detection probability. The communication throughput is normalized by dividing by the system's maximum achievable throughput to convert it into a dimensionless throughput achievement rate, which is then weighted and summed with the dimensionless perception detection probability on the same dimension. Communication throughput is measured by the total data rate carried by the subcarriers allocated for communication functions, specifically the sum of the instantaneous data rates corresponding to the modulation and coding schemes determined by each communication subcarrier based on its channel quality indicator value. Perception detection probability is measured by the target detection probability calculated using the radar equation and the constant false alarm rate criterion by the total sensing signal bandwidth carried by the subcarriers allocated for sensing functions. When the perception detection probability is lower than the constraint threshold (95%), a penalty term is introduced into the reward function to encourage the agent to increase the allocation of sensing resources in subsequent decisions. The specific calculation of the reward value is: multiply the communication throughput by the communication weight value. Add the detection probability multiplied by the perception weight value Then subtract the penalty term when the perception detection probability is lower than the constraint threshold.

[0049] During training, the agent maintains an approximate set of Pareto fronts for communication throughput and perception detection probabilities. A Pareto front is a set of solutions where any solution improving communication throughput necessarily leads to a decrease in perception detection probability, and vice versa. Each Pareto solution records its corresponding weight values ​​and policy parameters, forming a weight-policy mapping. This mapping allows the system to perform interpolation and transfer using existing Pareto solutions instead of retraining from scratch when weights change.

[0050] The Pareto front update and utilization mechanism is as follows: After each training round, the accumulated communication throughput and perception detection probability of the current round are combined into a two-dimensional performance point. If this point dominates some points in the Pareto front, it is replaced; if this point is dominated by some points in the front, it is discarded; otherwise, it is added to the front. The Pareto front policy transfer mechanism is as follows:

[0051] ,

[0052] in: For the new weight values The corresponding policy parameter estimates are vectors with dimensions equal to the total number of parameters in the actor network. They are calculated using this formula and represent the initial values ​​of the policy parameters that the agent should adopt after the weight changes. old weight values The corresponding converged policy parameters are extracted from the nearest neighbor operating point in the Pareto front; Let be the migration step size factor, and let be a scalar with a range of values. Preferred , dimensionless, controls the aggressiveness of policy transfer. Too large a value may lead to policy instability after transfer, while too small a value will result in insignificant transfer effect; Let be the gradient of the policy parameters with respect to the weights, and be a vector with the same dimensions. The optimal adjustment direction of the strategy parameters corresponding to a unit change in weight is obtained by finite difference approximation of the strategy parameters of adjacent operating points on the Pareto front. For gradient operators; and These are the perceived weight values ​​before and after the change, respectively, both of which are scalars and dimensionless.

[0053] When the dynamic weights output by the scene-aware weight adjuster change, the agent uses the above formula to slide along the gradient direction from the policy parameters at the current operating point on the Pareto front to the estimated policy parameters corresponding to the new operating point. Then, only a few fine-tuning iterations are needed to converge to the optimal policy under the new weights. Compared to retraining from randomly initialized parameters, Pareto front sliding transfer reduces the policy switching latency from thousands of training steps to dozens of fine-tuning steps, enabling the system to maintain real-time responsiveness in vehicle-to-everything (V2X) scenarios where dynamic weights change frequently.

[0054] After step S5 is completed, the allocation scheme is executed, and the base station follows the allocation matrix. Communication data transmission or sensing signal transmission is performed on each subcarrier. Subcarriers assigned to communication functions carry downlink data packets for each communication user. The base station selects the corresponding modulation and coding scheme based on the channel quality indicator value of each user to maximize spectral efficiency. Subcarriers assigned to sensing functions transmit radar detection waveforms. These waveforms are reflected by the target to generate echo signals, which are pulse-compressed and Doppler-processed by the base station receiver to extract the target's range and velocity information. The echo of the sensing signal is processed and updated to update the kinematic state parameters of the sensed target, including estimated range, estimated radial velocity, and estimated radar cross-section. The updated kinematic state parameters are acquired by the state construction module in the next time slot and incorporated into a new joint state vector, proceeding to step S1 in the next time slot.

[0055] The steps S1 to S5 described above constitute a complete closed-loop control cycle. Within each scheduling time slot, five steps are executed sequentially: state construction, weight adjustment, resource allocation, beam prediction, and policy update. These steps have strict data dependencies. This closed-loop mechanism enables the system to continuously and adaptively track the dynamic changes in communication service load and the kinematic state of the sensed target without human intervention. It makes near-globally optimal resource allocation decisions in real time within each time slot, meeting the stringent requirements of continuous perception and real-time communication for safety-critical applications such as vehicle-to-everything (V2X) applications. From an information theory perspective, the scene perception weight adjuster transforms the physical layer information (distance, velocity, cross-sectional area) of the sensed target into optimization-level control signals (weight values), realizing cross-layer transmission of physical layer perception information to resource management layer decision-making information. This cross-layer information flow is the essential advantage of the integrated sensing system compared to independent communication and sensing systems.

[0056] See Figure 2 This invention also provides a sensing-integrated waveform adaptive resource allocation system, applied to sensing-integrated systems employing orthogonal frequency division multiplexing waveforms, corresponding one-to-one with the steps in the above method embodiments. This system includes the following four modules.

[0057] State construction module 1 is configured to collect current state information from both the communication and sensing sides at the start of each scheduling time slot, and construct a joint state vector. The communication side's state information includes the channel quality indicator (CMI) value for each subcarrier and the current length of the communication data buffer queue. The CMI value is decoded by the base station's physical layer processing unit from the channel state information feedback message reported by the communication user. The buffer queue length is read by the base station's media access control layer from the data buffer of the packet data aggregation protocol layer. The sensing side's state information includes the estimated distance, estimated radial velocity, and estimated radar cross-section between each sensed target and the base station. These three parameters are extracted by the base station's sensing processing unit during the sensing echo processing of the previous scheduling cycle. This module concatenates the above information into a joint state vector according to a predefined format and adapts to changes in vector dimension under different target numbers through padding and masking mechanisms. This module corresponds to step S1 in the method embodiment.

[0058] Scene perception weight adjustment module 2 is configured to receive the kinematic state parameters of the perceived targets output by state construction module 1, calculate the threat level score, and map the threat level score to perception weight values ​​and communication weight values ​​through a nonlinear activation function. This module operates as a third network module independent of the deep reinforcement learning agent in the resource allocation module, and the parameters of its weight mapping function are jointly optimized with the agent during training. When multiple perceived targets exist, this module takes the maximum value of the threat level scores of each target as the comprehensive threat level score. This module also implements smoothing regularization to limit the change in perception weight values ​​between adjacent time slots to no more than the upper bound of the physically reachable weight change. This module corresponds to step S2 in the method embodiment.

[0059] Resource allocation module 3 is configured to receive the joint state vector output by state construction module 1 and output the allocation scheme of communication and sensing functions for each subcarrier in each time slot through an embedded deep reinforcement learning agent. The agent's actor network takes the joint state vector as input, processes it through three hidden layers (each containing 256 neurons and a modified linear activation function), and the output layer generates an allocation probability vector through a sigmoid activation function. After decision thresholding, a binary allocation matrix is ​​generated. The critic network concatenates the joint state vector and the allocation probability vector as input and outputs a scalar state-action value estimate. This module also includes a look-ahead pre-allocation submodule, which uses a target kinematics prediction model to calculate the predicted echo signal-to-noise ratio (SNR) of future time slots. Future time slots with a predicted SNR lower than the detection threshold are marked as time slots with scarce sensing resources, and this marking information is added as an extra dimension to the joint state vector so that the agent can reserve sensing resources in advance. This look-ahead mechanism upgrades the resource allocation strategy from reactive to predictive, effectively avoiding the inherent lag effect of reactive strategies under the fourth power inverse proportionality law of the radar equation. This module also includes a perception-assisted communication beam prediction submodule. This submodule converts the target prediction position output by the kinematic prediction model into a predicted communication beam azimuth angle value, and pre-adjusts the communication beam pointing in the time slot preceding the time slot where perception resources are scarce. This module corresponds to steps S3 and S4 in the method embodiment.

[0060] The policy update module 4 is configured to receive dynamic weights output by the scene perception weight adjustment module 2 and allocation execution feedback from the resource allocation module 3, and calculate the reward value of the deep reinforcement learning agent. The reward value is a weighted sum of normalized communication throughput and perception detection probability according to the dynamic weights, where the communication throughput is normalized to a dimensionless throughput achievement rate by dividing by the maximum achievable throughput of the system and then weighted with the perception detection probability under a unified dimension. This module maintains an approximate set of Pareto fronts during the training phase and quickly updates the policy parameters using the Pareto front sliding migration mechanism when the weights change. The updated policy parameters are fed back to the actor network and critic network of the resource allocation module 3. This module corresponds to step S5 in the method embodiment.

[0061] The data flow between the four modules forms a closed loop: the output of the state construction module 1 is simultaneously fed into the resource allocation module 3 and the scene perception weight adjustment module 2, realizing the parallel distribution and processing of communication-side information and perception-side information; the output of the weight adjustment module 2 is fed into the policy update module 4, which transforms the physical layer perception requirements into control signals at the optimization level; the output of the resource allocation module 3, after execution, generates communication and perception performance feedback information, which is fed into the policy update module 4; the policy parameter update result of the policy update module 4 is fed back to the agent network of the resource allocation module 3, while the perception execution result is fed back to the state construction module 1 to update the kinematic state parameters of the perception target in the next time slot.

[0062] To verify the effectiveness of the method of this invention, a sensor-integrated simulation platform for vehicle-to-everything (V2X) scenarios was built. The simulation scenario was set as an urban intersection environment, with the base station deployed on a tower in the center of the intersection, and the antenna height was [missing information]. The coverage radius is The vehicle density within the intersection area is dynamically generated according to a Poisson process, with an average vehicle arrival rate of 12 vehicles per minute. Vehicle trajectories are generated based on an intelligent driver model, including behavioral patterns such as going straight, turning, and stopping. The simulation parameters are set as follows: The orthogonal frequency division multiplexing system adopts... There are subcarriers, with a subcarrier spacing of . The corresponding system bandwidth is Scheduling time slot interval The base station's transmission power is The carrier frequency is Corresponding wavelength The number of communication users is 8, the number of sensed targets is 3, and the target distance range is [missing information]. to The target speed range is to (correspond to The target radar cross-section range is to The channel model employs a three-dimensional spatial channel model, incorporating path loss, shadowing fading, and multipath small-scale fading components. The actor and critic networks of the deep reinforcement learning agent both utilize 3 fully connected layers with a hidden layer dimension of 256 and a learning rate of [missing information]. The experience replay buffer capacity is The batch size is 256, and the discount factor is 0.99. The parameters of the scene-aware weight adjuster are set to... , , , , Simulation run Each time slot, of which the first One time slot is used for training, then... One time slot is used for testing.

[0063] The comparative schemes include: Scheme 1 is a fixed-ratio allocation scheme, which allocates 70% of the subcarriers to communication and 30% to sensing. This scheme is the most commonly used baseline scheme in engineering practice. Scheme 2 is a fixed-weight deep reinforcement learning scheme, which uses a deep reinforcement learning agent but the weights of communication and sensing in the reward function are fixed at 0.5 and 0.5 respectively. This scheme represents the typical level of existing dynamic allocation schemes based on intelligent learning. Scheme 3 is the scheme of this invention, which uses a scene-aware weight adjuster to dynamically adjust the weights and includes all mechanisms of smoothing regularization, look-ahead pre-allocation, and Pareto front sliding transfer.

[0064] The simulation results are analyzed from the following four dimensions.

[0065] In terms of communication throughput, with the constraint that the detection probability is greater than 95%, the average communication throughput of the present invention is: Compared to Option 1 It improved by approximately 32%, compared to Option 2. The throughput increased by approximately 17%. The main reason for the improved throughput is that when the perceived threat level score in the scenario is low, the weight adjuster of this invention automatically allocates more resources to communication, while Scheme 1 and Scheme 2 still maintain a high proportion of perception resources, resulting in insufficient communication resources.

[0066] In the dimension of perception and detection probability, the scenario of a target approaching at high speed is defined as the target moving from... by Approaching The detection probability of this invention remains above 97% throughout the entire process, even when the target distance is reduced to... Within a certain range, the detection probability reaches 99.5%. Scheme 1 applies when the target distance exceeds... The probability of time-sensing detection drops below 85% because the fixed 30% of sensing resources cannot provide sufficient echo signal-to-noise ratio at long distances. Scheme 2 addresses this issue when the target distance exceeds [a certain threshold]. The fluctuation in the probability of time-sensing detection is due to the fact that fixed weights prevent the agent from accurately recognizing the urgency of sensing distant targets.

[0067] In terms of policy switching response, the scheme of this invention, after adopting Pareto front sliding migration, achieves an average of 35 policy convergence steps when weights change, while the zero-search scheme without Pareto migration requires an average of 2400 steps, reducing policy switching latency by approximately 98%. This means that in Under time-slot interval conditions, the policy switching time of the Pareto migration scheme is approximately It fully meets the real-time response requirements for changes in the kinematic state of targets in vehicle networking scenarios.

[0068] In terms of the effectiveness of smoothing regularization, the convergence curves of the agent training process with and without smoothing regularization were compared. Introducing smoothing regularization reduced the variance of the reward value during training by approximately 60%, and the number of training steps required for convergence decreased from... Step down to Furthermore, smoothing regularization successfully intercepted four false target threat events caused by clutter interference in the simulation, avoiding unnecessary surges in sensing resources.

[0069] The experimental results demonstrate that this invention effectively solves the problem of decoupling the optimization direction from physical requirements in dynamic scenarios by transforming the reward function weights from fixed hyperparameters to state-dependent variables through a scene-aware weight regulator. The four mechanisms—weight regulator, smoothing regularization, look-ahead pre-allocation, and Pareto front transfer—work together to ensure a dynamic optimal balance between communication throughput and sensing detection probability in the integrated sensing system.

[0070] The embodiments of the present invention are not limited to the specific embodiments described above. Those skilled in the art can make various equivalent changes or substitutions based on the technical solutions of the present invention, and all such changes or substitutions should be included within the protection scope of the present invention.

Claims

1. A transducer-integrated waveform adaptive resource allocation method, applied to a transducer-integrated system employing orthogonal frequency division multiplexing waveforms, characterized in that, Includes the following steps: Construct a joint state vector, which includes the channel quality index and service load index of the current communication service, as well as the kinematic state parameters of the sensed target; The joint state vector is input into the deep reinforcement learning agent, and the deep reinforcement learning agent outputs the allocation scheme of communication and sensing functions of each subcarrier in each time slot. The scene perception weight adjuster determines the dynamic weights of the communication performance index and the perception performance index in the reward function based on the kinematic state parameters of the perceived target. The dynamic weights increase the perception weight as the threat level score of the perceived target increases and increase the communication weight as the threat level score of the perceived target decreases. The reward value of the deep reinforcement learning agent is calculated based on the dynamic weights, and the policy parameters of the deep reinforcement learning agent are updated based on the reward value.

2. The integrated inductive waveform adaptive resource allocation method as described in claim 1, characterized in that, The steps for determining dynamic weights by the scene perception weight adjuster include: calculating a threat level score based on the kinematic state parameters of the perceived target. The threat level score is based on a nonlinear mapping established by the inverse fourth power relationship between the echo signal-to-noise ratio and the target distance in the radar equation and the Doppler spread caused by the target's relative velocity. The closer the distance and the higher the relative velocity, the larger the threat level score. The threat level score is then mapped to a perception weight value through a monotonically increasing nonlinear activation function, and a communication weight value is obtained by subtracting the perception weight value from one.

3. The integrated sensing waveform adaptive resource allocation method as described in claim 2, characterized in that, The step of mapping the threat level score to the perception weight value further includes: applying a smoothing regularization process based on the target kinematic continuity constraint to the perception weight values ​​between consecutive time slots. The smoothing regularization process limits the change in perception weight values ​​between adjacent time slots to no more than the upper bound of the physical reachable weight change determined by the target's maximum acceleration and the time slot interval. When multiple perception targets exist simultaneously, the maximum value of the threat level scores of each target is used as the comprehensive threat level score, so that the resource allocation scheme prioritizes the detection reliability of the highest threat target.

4. The integrated sensing waveform adaptive resource allocation method as described in claim 3, characterized in that, The steps of the deep reinforcement learning agent output allocation scheme further include: based on the current kinematic state parameters of the perceived target and the threat level score after smoothing and regularization, using the target kinematic prediction model to estimate the predicted position and predicted velocity of the target in a preset number of time slots in the future; calculating the predicted echo signal-to-noise ratio of each future time slot based on the predicted position and predicted velocity using the radar equation; marking future time slots with predicted echo signal-to-noise ratios lower than the detection threshold as time slots with scarce perception resources; and adding the marking information of the time slots with scarce perception resources to the state space of the deep reinforcement learning agent, so that the deep reinforcement learning agent can reserve additional perception subcarrier resources in advance for the upcoming time slots with scarce perception resources.

5. The integrated sensing waveform adaptive resource allocation method as described in claim 4, characterized in that, The target predicted position output by the target kinematics prediction model is also used to assist in the pre-adjustment of the communication beam direction. Specifically, this includes: converting the target predicted position into the corresponding communication beam angle prediction value, and pointing the communication beam to the predicted angle direction in the time slot before the time slot where sensing resources are scarce, so as to reduce the communication beam alignment delay. The angle prediction accuracy of the sensing-assisted communication beam prediction is improved with the introduction of the target kinematic continuity constraint of the smoothing regularization process, so that the communication beam switching delay is lower than the beam switching delay without the sensing-assisted scheme.

6. The integrated inductive waveform adaptive resource allocation method as described in claim 5, characterized in that, The policy update of the deep reinforcement learning agent also includes: maintaining an approximate set of Pareto fronts for communication throughput and perception detection probability during training. The dynamic weights output by the scene perception weight adjuster correspond to an operating point on the Pareto front. When the dynamic weights change, the deep reinforcement learning agent slides along the Pareto front to the new operating point instead of searching again in the entire policy space, so that the policy switching latency when the weights change is lower than the policy convergence latency when searching from zero.

7. The integrated inductive waveform adaptive resource allocation method as described in claim 1, characterized in that, The channel quality metrics include the channel quality indication value for each subcarrier, and the service load metrics include the current length of the communication data buffer queue.

8. The integrated inductive waveform adaptive resource allocation method as described in claim 1, characterized in that, The kinematic state parameters of the sensed target include the estimated distance between the sensed target and the base station, the estimated radial velocity of the sensed target relative to the base station, and the estimated radar cross-section of the sensed target.

9. The integrated inductive waveform adaptive resource allocation method as described in claim 1, characterized in that, The deep reinforcement learning agent adopts an actor-critic network architecture, in which the actor network outputs the allocation scheme based on the joint state vector, the critic network evaluates the state-action value function based on the joint state vector and the allocation scheme, and the scene-aware weight adjuster operates as a third network module independent of the actor network and the critic network.

10. A sensor-integrated waveform adaptive resource allocation system, applied to a sensor-integrated system employing orthogonal frequency division multiplexing waveforms, used to implement the sensor-integrated waveform adaptive resource allocation method according to any one of claims 1-9, characterized in that, include: The state construction module is configured to construct a joint state vector, which includes the channel quality index and service load index of the current communication service, as well as the kinematic state parameters of the sensed target. The resource allocation module is configured to receive the joint state vector and output the allocation scheme of communication and sensing functions of each subcarrier in each time slot through a deep reinforcement learning agent. The scene perception weight adjustment module is configured to determine the dynamic weights of the communication performance index and the perception performance index in the reward function based on the kinematic state parameters of the perceived target, wherein the dynamic weights increase the perception weight as the threat level score of the perceived target increases and increase the communication weight as the threat level score of the perceived target decreases. The policy update module is configured to calculate the reward value of the deep reinforcement learning agent based on the dynamic weights, and update the policy parameters of the deep reinforcement learning agent based on the reward value.