A device and a method for triggering sensing in an integrated communication and sensing capable communication network
The reinforcement learning-based method for triggering sensing in integrated communication networks addresses inflexibility and high latency by dynamically selecting policies based on rewards, improving sensing efficiency and change detection.
Patent Information
- Application Number
- PCT/EP2025/072912
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-15
- Filing Date
- 2025-08-08
- Publication Date
- 2026-02-19
AI Technical Summary
Existing sensing mechanisms in integrated communication and sensing capable communication networks are inflexible, leading to high overhead and latency, particularly in dynamic environments.
A method and device for triggering sensing using reinforcement learning-based policies, where policies are selected based on associated rewards, allowing for flexible and efficient resource allocation and result-based learning to improve sensing capabilities.
This approach reduces overhead and latency while enhancing the ability to detect environmental changes by rewarding policies that improve sensing performance.
Smart Images

Figure EP2025072912_19022026_PF_FP_ABST
Abstract
Description
[0001] R.411681
[0002] - 1 -
[0003] Description
[0004] Title
[0005] A device and a method for triggering sensing in an integrated communication and sensing capable communication network
[0006] Background
[0007] The invention relates to a device and a method for triggering sensing in an integrated communication and sensing capable communication network.
[0008] In communication networks that are capable of integrated communication and sensing, the sensing may affect the communication.
[0009] Disclosure of the invention
[0010] A method for triggering sensing in an integrated communication and sensing capable communication network comprises providing a set of policies, wherein a respective policy of the set of policies corresponds to at least one parameter that defines a configuration of at least a part of the communication network for integrated communication and sensing, wherein the policies of the set of policies are associated with a respective reward for selecting the respective policy, wherein the method comprises selecting a policy of the set of policies depending on the rewards associated with the policies of the set of policies, sending a message for triggering the sensing, wherein the message comprises the at least one parameter defined by the selected policy, wherein the method comprises receiving at least one result of the sensing, and learning the reward for the selected policy depending on the at least one result. This method triggers the sensing based on a reward that is learned in a reinforcement learning depending on the result of the sensing. This method is more flexible than triggering the sensing in a static deterministic and cyclic manner. This method requires less overhead and latency than a dynamic event based triggering. R.411681
[0011] - 2 -
[0012] The reward for example represents a probability value, wherein selecting the policy comprises sampling the policy randomly based on a distribution of the rewards. The higher the reward, the more likely the policy is sampled. This favors the policies for that the reward is higher.
[0013] The method may comprise initializing the rewards with the same value. The uniform distribution of the rewards provides the same chance to be selected to the policies.
[0014] The method may comprise initializing at least two rewards with different values. This allows to prefer a policy over another policy, e.g., based on expert knowledge.
[0015] According to an exemplary selection scheme, selecting the policy may comprise selecting the policy that is associated with the highest reward of the rewards.
[0016] The communication network may comprise a set of base stations, wherein the method comprises selecting a subset of the set of base stations, sending the message to the base stations in the subset of the base stations, receiving the result of the sensing from the base stations, and learning the reward depending on the results.
[0017] The at least one parameter for example define at least one resource of the communication network for sensing, and / or a duration of the sensing, and / or a time of triggering the sensing.
[0018] The learning comprises determining a property of the sensing depending on the at least one result, and determining the reward depending on the property. Considering the property of the sensing for determining the reward encourages the learning of policies that improve the property of the sensing.
[0019] The property for example indicates an ability of the sensing to detect a change in an environment that is subject to the sensing, wherein the learning comprises determining the reward with a function that increases the reward with increasing ability. Thus, policies that provide a better ability of the sensing to detect the R.411681
[0020] - 3 - change are rewarded higher than policies that provide comparably less ability of the sensing to detect the change.
[0021] A device for triggering sensing in an integrated communication and sensing capable communication network is configured to execute the method.
[0022] A computer program for triggering sensing in an integrated communication and sensing capable communication network comprises computer readable instructions that, when executed by the computer and / or the device, cause the computer and / or the device to execute the method. According to a further aspect, there is provided a computer-readable medium having stored thereon the computer program.
[0023] Further advantageous embodiments are derivable from the following description and the drawing. In the drawing:
[0024] Fig. 1 schematically depicts a communication network,
[0025] Fig. 2 schematically depicts a method for triggering sensing in the communication network.
[0026] Figure 1 schematically depicts a communication network 100.
[0027] The communication network 100 comprises at least one user equipment 101 , at least one base station 102, and a core network 103. The communication network 100 is a cellular network. The cellular network comprises cells.
[0028] User equipment 101 in this context may refer to a mobile phone or a sensor, e.g., a camera, or a relay module for connecting a sensor to the at least one base station 102.
[0029] Figure 1 depicts a set of base stations 102.
[0030] Objects may be present in an environment of a base station 102. Examples for an object are a car 104, a bicycle 105 or a pedestrian 106. R.411681
[0031] - 4 -
[0032] A respective base station 102 is associated with a respective cell to provide access to the communication network 100 for the at least one user equipment 101 or the car 104 in the respective cell.
[0033] The communication network 100 is capable of integrated communication and sensing.
[0034] The integrated communication and sensing may involve sensing with radio signals, in particular reflected, refracted, and / or diffracted radio signals. Integrated communication and sensing, or integrated sensing and communication, respectively, refers to performing, preferably radar, sensing using radio signals impacted, e.g., reflected, refracted, diffracted, by the object 104, 105, 106 or the environment. In other words, radio resources of the (wireless) communication network can be used for both communication and performing (radar) sensing of the environment.
[0035] The integrated communication and sensing may involve sensing with various types of sensors, such as the camera, a radar sensor, a lidar sensor, and / or a motion sensor.
[0036] The respective base station 102 is configured for integrated communication and sensing. This means, the base station 102 is configured for sensing a presence of an object in the environment of the base station 102 and sending a result of the sensing to the core network 103.
[0037] The at least one base station 102 and the at least one user equipment 101 use radio signals for communication between a respective base station 102 and a respective user equipment 101.
[0038] When the sensing is performed, the base station 102 uses radio signals impacted, e.g., reflected, refracted, diffracted, by a respective object in the environment of the base station 102 for determining the result of the sensing.
[0039] The respective base station 102 is configured to report a result of the sensing by the respective base station 102 to the core network 103. R.411681
[0040] - 5 -
[0041] The core network 103 may be configured to determine a presence of a new object in the environment of the base station 102 depending on the result of the sensing.
[0042] The core network 103 may be configured to determine a size or velocity of an object in the environment of the base station 102 depending on the result of the sensing.
[0043] The core network 103 may be configured to determine a context map 107 depending on the result of the sensing sensed by and / or received from one base station 102 or depending on the results of the sensing sensed by and / or received from two or more base stations 102.
[0044] The communication network 100 is configurable for integrated communication and sensing.
[0045] At least one parameter defines a configuration of at least a part of the communication network 100 for integrated communication and sensing. The at least one parameter defines for example a base station 102 or multiple base stations 102. The at least one parameter defines for example at least one resource of the communication network 100 for sensing, e.g., time slots, specific frequencies, antennas, and / or beams that are allocated either for sensing or communication. The at least one parameter defines for example a duration of the sensing, and / or a time of triggering the sensing.
[0046] The communication network 103 is configured for triggering the sensing. The communication network 103 is configured for sending a message for triggering the sensing to at least one base station 102. The communication network 103 is configured to provide the at least one parameter in the message.
[0047] The at least one base station 102 is configured to perform the sensing when the communication network triggers the sensing by the at least one base station 102. The at least one base station 102 is configured to perform the sensing upon receipt of the message. The at least one base station 102 is configured to perform the sensing according to the at least one parameter provided in the message. R.411681
[0048] - 6 -
[0049] The at least one base station 102 is configured to report the result of sensing to the at core network 103.
[0050] Figure 2 depicts a method for triggering sensing in an integrated communication and sensing capable communication network. The method is for example executed in iterations. One iteration is described below.
[0051] The core network 103 is an example for a device that is configured to execute the method.
[0052] The method comprises a step 201.
[0053] The step 201 comprises providing 201 a set of policies.
[0054] A respective policy of the set of policies defines the at least one parameter that defines the configuration of at least a part of the communication network 100 for integrated communication and sensing.
[0055] Different policies define the at least one parameter differently.
[0056] Different policies for example define different base stations 102 for sensing.
[0057] Different policies for example define different resources of the communication network 100 for sensing, e.g., different time slots, different specific frequencies, different antennas, and / or different beams that are allocated either for sensing or communication. Different policies for example define different durations of the sensing, and / or different times of triggering the sensing.
[0058] The policies of the set of policies are associated with a respective reward for selecting the respective policy.
[0059] The reward represents for example a probability value.
[0060] The method may comprise initializing the rewards with the same value, or initializing at least two rewards with different values. The different values are for example given values provided by an expert. R.411681
[0061] According to an example for selecting a policy from a set of K e N policies, the rewards are probability values pk>such that where N is the set of non-negative integer numbers and i is the iteration index.
[0062] For the first iteration i = 0 the policies may have the same probability value:
[0063] The selected policyk >is determined in the iteration i of the method wherein k* e [1: K],
[0064] The method comprises a step 202.
[0065] The step 202 comprises selecting a policy of the set of policies depending on the rewards associated with the policies of the set of policies.
[0066] Selecting the policy may comprise selecting the policy that is associated with the highest reward of the rewards.
[0067] In case the reward represents the probability value, selecting the policy may comprise sampling the policy randomly based on a distribution of the rewards.
[0068] The step 202 may comprise selecting a subset of the set of base stations 102.
[0069] The method comprises a step 203.
[0070] The step 203 comprises sending the message for triggering the sensing.
[0071] The message is for example sent to at least one base station 102. R.411681
[0072] - 8 -
[0073] The step 203 for example comprises selecting a subset of the base stations 102 depending on the at least one parameter that defines the base station 102 or the base stations 102 for sensing.
[0074] The step 203 for example comprises sending a message to the respective base station 102 in the subset of the base stations 102.
[0075] The message comprises the at least one parameter defined by the selected policy.
[0076] The method comprises a step 204.
[0077] The step 204 comprises receiving at least one result of the sensing.
[0078] The step 204 may comprise receiving the result of the sensing from at least one base station 102.
[0079] In case the subset of the base stations 102 is used for sensing, the step 204 may comprise receiving the results of the sensing from the subset of the base stations 102.
[0080] The method comprises a step 205.
[0081] The step 205 comprises learning the reward for the selected policy depending on the at least one result.
[0082] The step 205 may comprise learning the reward depending on the result of sensing performed by at least one base station 102.
[0083] In case the subset of the base stations 102 is used for sensing, the step 205 may comprise learning the reward depending on the result of sensing performed by the subset of the base stations 102.
[0084] The learning may comprise determining a property of the sensing depending on the at least one result, and determining the reward depending on the property. R.411681
[0085] - 9 -
[0086] The property for example indicates an ability of the sensing to detect a change in an environment that is subject to the sensing. The learning may comprise determining the reward with a function that increases the reward with increasing ability.
[0087] According to an example for learning, a function where y is a real number representing the reward.
[0088] Using a strictly positive y results in a function that increases the value of the reward for the policy nfk>in case the policy nfk>leads to a better result than another policy. Using a strictly negative ytresults in a function that decreases the value of the reward for the policy nfk>in case the policy nfk>leads to a worse result than another policy.
[0089] A possible numerical value of ytis a real number within
[0090] The probabilities associated with the possible policies are for example updated as follows: under the conditions
[0091] 0 < p(<k*> +Yi< 1
[0092] Afterwards the step 202 is executed.
Claims
R.411681- 10 -Claims1. A method for triggering sensing in an integrated communication and sensing capable communication network (100), characterized in that the method comprises providing (201) a set of policies, wherein a respective policy of the set of policies corresponds to at least one parameter that defines a configuration of at least a part of the communication network (100) for integrated communication and sensing, wherein the policies of the set of policies are associated with a respective reward for selecting the respective policy, selecting (202) a policy of the set of policies depending on the rewards associated with the policies of the set of policies, sending (203) a message for triggering the sensing, wherein the message comprises the at least one parameter defined by the selected policy, receiving (204) at least one result of the sensing, and learning (205) the reward for the selected policy depending on the at least one result.
2. The method according to claim 1, characterized in that the reward represents a probability value, wherein selecting (202) the policy comprises sampling the policy randomly based on a distribution of the rewards.
3. The method according to claim 2, characterized in that the method comprises initializing (201) the rewards with the same value.
4. The method according to claim 2, characterized in that the method comprises initializing (201) at least two rewards with different values.
5. The method according to claim 1 , characterized in that selecting (202) the policy comprises selecting the policy that is associated with the highest reward of the rewards.R.411681- 11 -6. The method according to one of the preceding claims, characterized in that the communication network (100) comprises a set of base stations (102), wherein the method comprises selecting (202) a subset of the set of base stations, sending (203) the message to the base stations in the subset of the base stations (102), receiving (204) the result of the sensing from the base stations (102), and learning (205) the reward depending on the results.
7. The method according to one of the preceding claims, characterized in that the at least one parameter defines at least one resource of the communication network (100) for sensing and / or a duration of the sensing, and / or a time of triggering the sensing.
8. The method according to one of the preceding claims, characterized in that the learning (205) comprises determining a property of the sensing depending on the at least one result, and determining the reward depending on the property.
9. The method according to claim 8, characterized in that the property indicates an ability of the sensing to detect a change in an environment that is subject to the sensing, wherein the learning (205) comprises determining the reward with a function that increases the reward with increasing ability.
10. A device (103) for triggering sensing in an integrated communication and sensing capable communication network (100), characterized in that the device (103) is configured to execute the method according to one of the preceding claims.11 . A computer program for triggering sensing in an integrated communication and sensing capable communication network, characterized in that the computer program comprises computer readable instructions that, when executed by the computer and / or the device (103) according to claim 10, cause the computer and / or the device (103) according to claim 10 to execute the method according to one of the claims 1 to 9.R.411681- 12 -12. A computer-readable medium having stored thereon the computer program of claim 11.
Citation Information
Patent Citations
Network functions and a sensing node for sensing operations in a communication system
WO2024114930A1