Node selection and beam forming design method under dynamic sensing integrated network

By constructing a dynamic synesthesia integrated network model, using Markov decision-making process and deep reinforcement learning to optimize node selection and beamforming, the problems of insufficient resource allocation of communication base stations and complex spectrum are solved, and efficient communication and perception performance trade-offs are achieved, which is suitable for high-precision perception and high-quality communication in complex environments.

CN120358545AActive Publication Date: 2025-07-22SHANGHAI UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510845844.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

When existing communication base stations realize environmental perception, insufficient resource allocation and complex spectrum environment, resulting in insufficient perceptual accuracy and reliability, making it difficult to meet the application needs of high precision and high reliability.

Method used

Build a dynamic synesthesia integrated network model, optimize node selection and beamforming design through Markov decision-making process and deep reinforcement learning, and realize the trade-off between communication and perception performance.

Benefits of technology

Learn the optimal strategy independently in a dynamic environment, flexibly allocate resources, and meet the multiple needs of communication, perception and delay. It is suitable for high-precision perception and high-quality communication in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358545A_ABST
    Figure CN120358545A_ABST
Patent Text Reader

Abstract

The invention discloses a node selection and beamforming design method under a node dynamic communication and sensing integrated network. The method comprises the following steps: S1, constructing a communication and sensing model of a dynamic communication and sensing integrated network system; 2, constructing communication performance, sensing performance and time delay performance representations of the dynamic communication and sensing integrated network according to the communication and sensing model of the dynamic communication and sensing integrated network system; s3, constructing a node selection and beam forming optimization model of the dynamic communication and sensing integrated network according to the representation of the communication performance, the sensing performance and the time delay performance of the dynamic communication and sensing integrated network; and S4, expressing the optimization model as a Markov decision process, and solving by using deep reinforcement learning to obtain a node selection scheme and a beamforming design scheme. By adopting the technical scheme of the invention, the trade-off of communication and sensing performance is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of wireless communication, and particularly relates to a method for node selection and beamforming design in a dynamic communication-sensing integrated network. Background Technique

[0002] With the rapid development of intelligent applications such as smart cities, intelligent transportation, and unmanned chemical plants, the demand for wireless infrastructure with environmental perception capabilities is increasing day by day. In this trend, the breakthrough development of 5G-A and 6G technologies provides new opportunities for the intelligent upgrade of communication base stations. Currently, communication base stations are evolving from a single signal transmission function to a communication-sensing integrated direction, and the environmental perception function is realized by using the transmitted communication signals. This technical path not only fully exploits the potential of existing infrastructure but also significantly reduces the cost of deploying additional dedicated sensing devices. In complex application scenarios such as smart cities and intelligent transportation, communication base stations need to achieve high-precision and high-reliability environmental perception capabilities, which pose strict requirements on the robustness of the sensing system. However, compared with dedicated radar systems, sensing based on communication signals faces significant challenges: firstly, due to the resource allocation mechanism that prioritizes communication services, the time-frequency resources available for sensing tasks are relatively limited; secondly, communication signals generally operate in lower frequency bands, and the spectrum environment is complex, resulting in a low signal-to-interference-plus-noise ratio. These factors severely restrict the accuracy and reliability of the system in key sensing tasks such as target parameter estimation. In summary, single-base-station sensing solutions are difficult to meet the actual application requirements. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a method for node selection and beamforming design in a dynamic communication-sensing integrated network.

[0004] To achieve the above object, the present invention adopts the following technical solutions: A method for node selection and beamforming design in a dynamic communication-sensing integrated network, including: Step S1, constructing a communication and sensing model of the dynamic communication-sensing integrated network system; Step 2, constructing expressions of the communication performance, sensing performance, and delay performance of the dynamic communication-sensing integrated network according to the communication and sensing model of the dynamic communication-sensing integrated network system; Step S3, constructing an optimization model for node selection and beamforming of the dynamic communication-sensing integrated network according to the expressions of the communication performance, sensing performance, and delay performance of the dynamic communication-sensing integrated network; Step S4, formulating the optimization model as a Markov decision process and solving it using deep reinforcement learning to obtain a node selection scheme and a beamforming design scheme.

[0005] The optimization model for node selection and beamforming in the preferred dynamic communication and sensing integrated network is as follows: ; Among them, is the weight of sensing performance, is the weight of delay, is the maximum transmit power threshold of each node, is the threshold of the minimum communication speed of each communication user, is the maximum delay threshold of each sensing request; is the action taken by the system at time slot ; is at time slot the beamforming matrix of base station ; is its conjugate transpose matrix, is the average parameter estimation of the sensing request, is the average sensing delay, , and respectively represent the indices of the base station, target, and time slot, represents the communication rate of user of base station , represents the delay suffered by the sensing request generated by target at time slot .

[0006] The present invention first establishes a communication and sensing model for a dynamic communication and sensing integrated network, which includes dynamically arriving sensing requests; then defines the representations of communication, sensing, and delay performance, and establishes corresponding optimization problems. Finally, the optimization problems are solved through a Markov decision process and a deep reinforcement learning method to achieve a trade-off between communication and sensing performance. The present invention has the following technical effects: 1. Through dynamic node selection and dynamic beamforming design, the present invention can flexibly allocate resources according to real-time sensing requests and channel conditions, thereby achieving an efficient trade-off between communication and sensing tasks.

[0007] 2. By adopting the model-free Q-learning reinforcement learning method, the present invention can autonomously learn the optimal strategy in a dynamic environment, adapt to unknown transition probabilities and time-varying channel conditions, and achieve long-term performance optimization.

[0008] 3. Through a unified optimization framework, the present invention simultaneously considers communication performance, sensing performance, and delay performance. This multi-objective collaborative optimization ability enables it to simultaneously meet multiple requirements of communication, sensing, and delay in a complex dynamic environment, and is applicable to scenarios with high requirements for real-time performance, sensing accuracy, and communication quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on the provided drawings.

[0010] Figure 1 It is a flowchart of a node selection and beamforming design method under the dynamic communication and sensing integrated network of the embodiments of the present invention; Figure 2 It is a schematic structural diagram of a dynamic communication and sensing integrated network system. Specific embodiments

[0011] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0012] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0013] Embodiment 1: As Figure 1 shown, the embodiments of the present invention provide a node selection and beamforming design method, including: Step S1, constructing a communication and sensing model of a dynamic communication and sensing integrated network system; Step 2, constructing representations of the communication performance, sensing performance, and delay performance of a dynamic communication and sensing integrated network according to the communication and sensing model of the dynamic communication and sensing integrated network system; Step S3, constructing an optimization model for node selection and beamforming of a dynamic communication and sensing integrated network according to the representations of the communication performance, sensing performance, and delay performance of the dynamic communication and sensing integrated network; Step S4, formulating the optimization model as a Markov decision process and solving it using deep reinforcement learning to obtain a node selection scheme and a beamforming design scheme.

[0014] As an implementation manner of the embodiments of the present invention, in step S1, as Figure 2 shown, constructing a dynamic communication and sensing integrated network system controlled by a central processing unit; among them, there are communication and sensing integrated nodes equipped with MIMO antennas, and each node is Providing communication services for multiple communication users, and there are perceived targets. Multiple nodes use beamforming to cooperate in perceiving these targets while providing communication services for their own communication users. Each node can only perceive one target at the same time, but can use the signals of self-transmission and reception as well as reception of other nodes' transmissions for sensing. Discretize time into time slots, and each time slot is represented by . Represent the communication channel between node and its user in time slot as , and represent the sensing channel between node and target in time slot as . The perceived targets in the scenario do not always require sensing services. The sensing requests for each perceived target arrive randomly. When , it means that target needs to be sensed in time slot . The sensing request is responded to by the central processor in the next time slot at the earliest, and corresponding node selection and beamforming design are carried out. Define the set of node selection binary variables , is the action taken by the system in time slot . means that node n is responsible for sensing target m in time slot . Define the set of multi-node beamforming matrix , where is the beamforming matrix of base station in time slot .

[0015] As an implementation manner of an embodiment of the present invention, in step S2, define , indicating the number of all sensing requests. Since there may be more sensing requests arriving than the number of nodes in the same time slot, there will be sensing requests that cannot be immediately responded to in the next time slot. Define the delay borne by the sensing request generated by target in time slot , where is the response time slot of the central processor to this sensing request. Then the average delay of all sensing requests can be expressed as . The communication performance is represented by the communication rate of user of base station , and the sensing performance is represented by the average parameter estimation Representation. Among them, the node selection design is responsible for associating nodes with sensing targets, and the multi-node beamforming design distributes power at different angles. The two work together to meet the communication performance requirements and sensing performance requirements and achieve a trade-off between communication and sensing performance.

[0016] As an implementation manner of an embodiment of the present invention, in step S3, the optimization model is used to optimize the sensing performance and delay of the dynamic communication and sensing integrated network system, while satisfying the rate constraint of communication users, the transmit power constraint of the dynamic communication and sensing integrated network system, and the delay constraint of each sensing request. The optimization model is written as: ;

[0017] Wherein, is the weight of the sensing performance, is the weight of the delay, is the maximum transmit power threshold of each node, is the threshold of the minimum communication rate of each communication user, is the maximum delay threshold of each sensing request; is the action taken by the system at time slot ; is at time slot the beamforming matrix of base station ; is its conjugate transpose matrix, is the average parameter estimation of the sensing request, is the average sensing delay, , and respectively represent the indices of the base station, the target, and the time slot, represents the user of base station 's communication rate, represents the target at time slot the delay suffered by the generated sensing request.

[0018] As an implementation manner of an embodiment of the present invention, in step S4, a Markov decision process is used for modeling and policy learning. The Markov decision process can be defined by the quadruple , where: S represents the state space, depicting information such as sensing requests, channel states, and delay constraints; A represents the action space, reflecting the current node selection and beamforming decisions of the system; P is the state transition probability; R is the immediate reward function, used to quantify the impact of each step of decision-making on the overall optimization goal. At each time slot , the processor observes the current system state , and selects an action After performing the action, the system obtains an immediate reward and transfers to the next state according to the state transition probability and continuously updates the strategy. To satisfy the terminal state constraint and achieve the optimal cumulative reward, the system continuously iteratively learns the optimal strategy so as to complete the approximation of the original optimization problem, specifically including: Step S41: At each time slot, the state of the dynamic communication and sensing integrated network system includes sensing requests, time-varying channel conditions, the node selection scheme and beamforming scheme adopted by multiple nodes in the previous time slot, that is ;

[0019] Step S42: The actions of the dynamic communication and sensing integrated network system include the node selection scheme and beamforming scheme, that is ;

[0020] Step S43: The optimization objective is to minimize the average CRB and average delay of sensing requests while satisfying the rate constraint of communication users, the transmit power constraint of the system, and the delay constraint of each sensing request. The immediate reward function is used to measure the decision-making effect of the system at each time slot, and it is a key evaluation index in the reinforcement learning process. The optimization objective is transformed from the original constraint model into maximizing the expected value of this reward function. This function comprehensively considers the sensing task delay, transmit power overhead, and rate constraint violation situation, and is the main basis for guiding policy learning. Reinforcement learning gradually approaches the optimal policy by continuously interacting with the environment and adjusting the policy according to the immediate reward. The immediate reward at time slot is calculated as follows: ;

[0021] where is the optimization objective, is the normalized communication rate threshold constraint, is the normalized response delay threshold constraint, is the normalized transmit power constraint.

[0022] Step S44: Since the state changes of the sensing network are complex and difficult to model, the traditional method of solving the Markov decision process that requires knowing the state transition probability is no longer applicable. Therefore, a model-free Q-learning reinforcement learning method is adopted to learn the resource allocation and scheduling strategy in a dynamic environment with unknown state transition probabilities. This method iteratively updates the action value function Q-value by interacting with the environment and uses the immediate reward function as the evaluation basis for the optimization objective to achieve a model-free solution to the original optimization problem. The finally converged optimal strategy can effectively minimize the sensing task delay and the system transmission power while satisfying the rate and response constraints.

[0023] The embodiments described above are only descriptions of the preferred embodiments of the present invention and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A node selection and beamforming design method under a dynamic communication-sensing integrated network, characterized in that Including: Step S1: Construct the communication and sensing models of the dynamic communication-perception integrated network system; Step 2: Construct the representations of the communication performance, sensing performance, and delay performance of the dynamic communication-perception integrated network according to the communication and sensing models of the dynamic communication-perception integrated network system; Step S3: Construct the optimization model for node selection and beamforming of the dynamic communication-perception integrated network according to the representations of the communication performance, sensing performance, and delay performance of the dynamic communication-perception integrated network; Step S4: Express the optimization model as a Markov decision process and solve it using deep reinforcement learning to obtain the node selection scheme and the beamforming design scheme.

2. The method for node selection and beamforming design in a dynamic communication and sensing integrated network according to claim 1, wherein The optimization model for node selection and beamforming of the dynamic communication-perception integrated network is: ; Among them, is the weight of the sensing performance, is the weight of the time delay, is the maximum transmission power threshold of each node, is the threshold of the minimum communication speed of each communication user, is the maximum time delay threshold of each sensing request; is the action taken by the system at time slot ; is at time slot the beamforming matrix of base station ; is its conjugate transpose matrix, is the average parameter estimation of the sensing request, is the average sensing delay, , and respectively represent the indexes of the base station, the target and the time slot, represents the communication rate of user of base station , represents the time delay suffered by the sensing request generated by target at time slot .

Citation Information

Patent Citations

  • Node mode selection and beam forming method in cooperative sensing integrated scene

    CN119483668A

  • Auxiliary and inductive integrated beam forming design method for active intelligent metasurface

    CN119727806A

  • Communication and perception integration method and system for base station and user cooperation

    CN119907099A

  • Super-large scale MIMO communication perception integration method

    CN120150765A