An Energy-Efficient Sensor Scheduling Method for Cooperative Sensing and Related Devices

Through a high-efficiency sensor scheduling method for collaborative perception, the current status information of the sensor is obtained and probability distribution sampling is performed, and an accurate scheduling strategy is formulated, which solves the problem of insufficient scheduling strategies and neglected energy consumption in the existing technology, and achieves efficient perceived data freshness and energy consumption optimization.

CN119233376BActive Publication Date: 2025-06-17SHENZHEN INST OF ARTIFICIAL INTELLIGENCE & ROBOTICS FOR SOC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411337577.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-24
Publication Date
2025-06-17
Estimated Expiration
2044-09-24

AI Technical Summary

Technical Problem

The existing sensor scheduling methods cannot accurately reflect dynamic channel conditions and sensor working mode, resulting in inaccurate scheduling strategies, increasing information age and reducing data freshness, while ignoring the energy consumption problem at the sensor side.

Method used

Using a high-efficiency sensor scheduling method for collaborative perception, a sleep probability distribution sampling and transmit power probability distribution sampling are carried out by obtaining the current status information of each sensor, including working mode, channel status and information age, and a precise scheduling strategy is formulated to optimize energy consumption and data freshness.

Benefits of technology

Under certain coverage constraints, the freshness of perceived data is improved, and the energy consumption of sensors is reduced, solving the problem of perceived information benefits trade-off in dynamic channel situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119233376B_ABST
    Figure CN119233376B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides an energy-efficient sensor scheduling method and related devices for collaborative perception, including: obtaining the current state information of each sensor at the current moment; sampling the current sleep duration according to the current state information; inputting the current state information into the lower-layer policy network to obtain the probability distribution of the transmission power, and sampling according to the probability distribution of the transmission power to obtain the current transmission power; sending the current sleep duration and the current transmission power to each sensor so that each sensor sends perception information to the target server; updating the current age of information of all target monitoring points according to the perception information to obtain the target age of information, and obtaining the benefit function value according to the target age of information; setting the sensor scheduling experience, and updating the current weight parameter value of the target server until a preset time threshold is met, and using the target weight parameter value as the target scheduling policy of the target server for each sensor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of sensor scheduling, and in particular, to an energy-efficient sensor scheduling method and related devices for cooperative perception. Background Art

[0002] In applications such as the Internet of Things and digital twins, the cooperative perception technology collects data from a target spatial area through widely distributed sensors and transmits it to a target server to achieve real-time monitoring of the physical environment. To ensure effective monitoring of the real-time physical environment, high-coverage perception data and timely data transmission are required, that is, high-quality perception data is received.

[0003] However, in the prior art, sensor scheduling is usually based on a simplified channel model and uses heuristic algorithms to achieve energy conservation. However, these methods may not accurately reflect the actual dynamic channel conditions and sensor working modes, resulting in inaccurate scheduling strategies. During the process of transmitting multi-source data to the cooperative perception center, due to the dynamic nature of the channel, there is a risk of transmission failure, which will increase the Age of Information (AoI) and reduce the freshness of the data.

[0004] A large part of the sensor scheduling research also focuses on minimizing the data freshness at the receiving end, while ignoring the energy consumption problem at the sensor end, and not fully considering how to optimize the energy consumption of sensors while ensuring data freshness. Summary of the Invention

[0005] The embodiments of the present application provide an energy-efficient sensor scheduling method and related devices for cooperative perception, which are used to ensure data freshness and reduce energy consumption under certain coverage constraints.

[0006] In a first aspect of the embodiments of the present application, an energy-efficient sensor scheduling method for cooperative perception is provided, which is applied to a target server. The target server is composed of an upper-layer policy network, an upper-layer value network, a lower-layer policy network, and a lower-layer value network, and includes:

[0007] Obtain the current state information of each sensor at the current moment; wherein, the current state information includes the current working mode of each sensor at the current moment, the current channel state between the target server and each sensor, and the current age of information of each sensor for any target monitoring point;

[0008] According to the current working mode, the current channel state, and the current age of information, sample the sleep probability distribution of each sensor to obtain the current sleep duration corresponding to each sensor in the upper-layer policy network;

[0009] Input the current working mode, the current channel state, the current age of information, and the current sleep duration into the lower-layer policy network to obtain the probability distribution of the transmission power corresponding to each sensor, and sample the power probability distribution according to the probability distribution of the transmission power to obtain the current transmission power corresponding to each sensor;

[0010] Send the current sleep duration and the current transmission power to each corresponding sensor, so that each sensor sends sensing information to the target server according to the current scheduling policy of the target server; wherein, there is an association relationship between the scheduling policy and the current sleep duration;

[0011] Update the current age of information of all target monitoring points according to the sensing information to obtain the target age of information, and obtain the benefit function value of each sensor at the current moment according to the target age of information; wherein, there is an association relationship between the benefit function value and the power consumption value of each sensor;

[0012] Set the current state information, the current sleep duration, the current transmission power, and the benefit function value as the sensor scheduling experience, and store the sensor scheduling experience in the experience pool according to a preset storage rule, and sample a preset number of sensor scheduling experiences from the experience pool to update the current weight parameter values of the upper-layer policy network, the upper-layer value network, the lower-layer policy network, and the lower-layer value network until a preset time threshold is met, and use the target weight parameter value corresponding to the preset time threshold as the selection basis for the target scheduling policy of the target server for each sensor.

[0013] Optionally, after obtaining the current state information of each sensor at the current moment, the method further includes:

[0014] Set the maximum sleep duration corresponding to all sensors; wherein, the current working mode includes a sleep mode and an activation mode, and the maximum sleep duration is used to represent the maximum continuous time for each sensor to be in the sleep mode;

[0015] Determine a target action set that satisfies the coverage constraint condition within the maximum sleep duration according to the current working mode of each sensor and the initial action set of each sensor; wherein, the initial action set is used to represent the action set of the current working mode of any sensor within the maximum sleep duration; the coverage constraint condition is used to represent the action set of sensors that satisfy a preset coverage rate and are in the activation mode at the current moment;

[0016] Input the current state information and the set of target actions into the upper-layer policy network to obtain an action dimension vector corresponding to the number of sensors in the set of target actions, so as to determine a masking vector according to the action dimension vector; wherein, the masking vector is used to represent the result of each sensor satisfying the coverage constraint condition at the current moment.

[0017] Optionally, the step of sampling the sleep probability distribution of each sensor according to the current working mode, the current channel state, and the current age of information to obtain the current sleep duration corresponding to each sensor in the upper-layer policy network includes:

[0018] Determine the probability distribution of the sleep probability of each sensor corresponding to different sleep durations according to the masking vector, the action dimension vector, the current working mode, the current channel state, and the current age of information; wherein, the probability distribution of the sleep probability is associated with the current weight parameter value of the upper-layer policy network;

[0019] Perform sampling of the sleep probability distribution according to the probability distribution of the sleep probability to obtain the current sleep duration.

[0020] Optionally, the step of inputting the current working mode, the current channel state, the current age of information, and the current sleep duration into the lower-layer policy network to obtain the probability distribution of the transmission power corresponding to each sensor includes:

[0021] Input the current working mode, the current channel state, the current age of information, and the current sleep duration into the lower-layer policy network to obtain a mean vector; wherein, the mean vector is associated with the current weight parameter value of the lower-layer policy network at the current moment;

[0022] Based on the mean vector, combine the variance vector corresponding to each sensor to obtain the probability distribution of the transmission power of each sensor.

[0023] Optionally, before the step of updating the current age of information of all target monitoring points according to the sensing information to obtain the target age of information, the method further includes:

[0024] If the current scheduling policy is a sleep policy, control each sensor to sleep for the corresponding current sleep duration according to the sleep policy;

[0025] If the current scheduling policy is an activation policy, control each sensor according to the activation policy so that each sensor sends the sensing information to the target server based on the corresponding current transmission power.

[0026] Optionally, updating the current information age of all target monitoring points according to the perception information to obtain a target information age includes:

[0027] If the target server receives the perception information of any target monitoring point, updating the current information age to a preset constant factor according to the perception information, and using the preset constant factor as the target information age;

[0028] If the target server does not receive the perception information of any target monitoring point, adding the current information age and the preset constant factor according to the perception information, and using the added value as the target information age.

[0029] Optionally, obtaining the benefit function value of each sensor at the current moment according to the target information age includes:

[0030] Determining the power consumption value corresponding to each sensor;

[0031] Based on the power consumption value, making an approximate estimate to obtain the sensor energy consumption value of each sensor at the current moment;

[0032] According to the sensor energy consumption value, the target information age, and the quantities of all sensors and target monitoring points, obtaining the benefit function value.

[0033] Optionally, the current weight parameter values include the current upper-layer policy weight value of the upper-layer policy network, the current upper-layer value weight value of the upper-layer value network, the lower-layer policy weight value of the lower-layer policy network, and the current lower-layer value weight value of the lower-layer value network corresponding to the current moment. Updating the current weight parameter values of the upper-layer policy network, the upper-layer value network, the lower-layer policy network, and the lower-layer value network by sampling a preset number of sensor scheduling experiences from the experience pool includes:

[0034] Analyzing the preset number of sensor scheduling experiences to obtain a first scheduling experience and a second scheduling experience; wherein, the second scheduling experience is the next scheduling experience of the first scheduling experience, and the first scheduling experience and the second scheduling experience at least include current state information, the current sleep duration, the current transmission power, and the benefit function value;

[0035] Based on the stochastic gradient descent algorithm, calculate the current sleep duration, the current upper-layer policy weight value, the current state information, and the first state value function of each sensor corresponding to the first scheduling experience to obtain the upper-layer policy weight value at the next moment; wherein, the state value function at the first moment is associated with the current state information and the current upper-layer value weight value corresponding to the first scheduling experience;

[0036] Based on the stochastic gradient descent algorithm, calculate the benefit function value, the first state value function, and the second state value function corresponding to the first scheduling experience to obtain the upper-layer value weight value at the next moment; wherein, the second state value function is associated with the current state information and the current upper-layer value weight value corresponding to the second scheduling experience;

[0037] Based on the stochastic gradient descent algorithm, calculate the current sleep duration, the current lower-layer policy weight value, the current state information, the current transmission power, and the first sleep action value function of each sensor corresponding to the first scheduling experience to obtain the lower-layer policy weight value at the next moment; wherein, the first sleep action value function is associated with the current state information, the current sleep duration, and the current lower-layer value weight value corresponding to the first scheduling experience;

[0038] Based on the stochastic gradient descent algorithm, calculate the benefit function value, the first sleep action value function, and the second sleep action value function corresponding to the first scheduling experience to obtain the lower-layer value weight value at the next moment; wherein, the second sleep action value function is associated with the current state information, the current sleep duration, and the current lower-layer value weight value corresponding to the second scheduling experience.

[0039] The second aspect of the embodiments of the present application provides an energy-efficient sensor scheduling system for cooperative sensing, which is applied to a target server. The target server is composed of an upper-layer policy network, an upper-layer value network, a lower-layer policy network, and a lower-layer value network, and includes:

[0040] An acquisition unit, configured to acquire the current state information of each sensor at the current moment; wherein, the current state information includes the current working mode of each sensor at the current moment, the current channel state between the target server and each sensor, and the current age of information of each sensor for any target monitoring point;

[0041] A sampling unit, configured to sample the sleep probability distribution of each sensor according to the current working mode, the current channel state, and the current age of information to obtain the current sleep duration corresponding to each sensor in the upper-layer policy network;

[0042] An input unit, configured to input the current working mode, the current channel state, the current information age, and the current sleep duration into the lower-layer policy network, to obtain a probability distribution of the transmission power corresponding to each sensor, and perform power probability distribution sampling according to the probability distribution of the transmission power, so as to obtain the current transmission power corresponding to each sensor;

[0043] A sending unit, configured to send the current sleep duration and the current transmission power to each corresponding sensor, so that each sensor sends sensing information to the target server according to the current scheduling policy of the target server; wherein, there is an association relationship between the scheduling policy and the current sleep duration;

[0044] An updating unit, configured to update the current information age of all target monitoring points according to the sensing information, to obtain a target information age, and obtain a benefit function value of each sensor at the current moment according to the target information age; wherein, there is an association relationship between the benefit function value and the power consumption value of each sensor;

[0045] A setting unit 606, configured to set the current state information, the current sleep duration, the current transmission power, and the benefit function value as sensor scheduling experience, and store the sensor scheduling experience into an experience pool according to a preset storage rule, to sample a preset number of sensor scheduling experiences from the experience pool, and update the current weight parameter values of the upper-layer policy network, the upper-layer value network, the lower-layer policy network, and the lower-layer value network until a preset time threshold is met, and use the target weight parameter value corresponding to the preset time threshold as a selection basis for the target scheduling policy of the target server for each sensor.

[0046] The sensor scheduling system provided in the second aspect of the embodiments of the present application is configured to execute the energy-efficient sensor scheduling method for collaborative sensing described in the first aspect.

[0047] The third aspect of the embodiments of the present application provides a sensor scheduling device, including:

[0048] A central processing unit, a memory, an input / output interface, a wired or wireless network interface, and a power supply;

[0049] The memory is a transient storage memory or a persistent storage memory;

[0050] The central processing unit is configured to communicate with the memory and execute instruction operations in the memory to execute the energy-efficient sensor scheduling method for collaborative sensing described in the first aspect.

[0051] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium includes instructions that, when run on a computer, cause the computer to execute the energy-efficient sensor scheduling method for collaborative sensing described in the first aspect.

[0052] In a fifth aspect of the embodiments of the present application, a computer program product is provided. The computer program product includes instructions that, when run on a computer, cause the computer to execute the energy-efficient sensor scheduling method for collaborative sensing described in the first aspect.

[0053] As can be seen from the above technical solutions, the embodiments of the present application have the following advantages: Through an energy-efficient sensor scheduling method for collaborative sensing disclosed in the embodiments of the present application, the sensor scheduling strategy is decomposed into two sub-strategies: the sleep cycle and the transmission power, solving the problem of benefit trade-off of sensing information in a dynamic channel situation. It can select the sleep cycle and transmission power of the sensor under certain coverage constraints to adapt to the changes in the channel environment, thereby improving the freshness of sensing data and reducing energy consumption. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments recorded in the present application, and those of ordinary skill in the art can also obtain other drawings based on these drawings.

[0055] Figure 1 It is a schematic flowchart of an energy-efficient sensor scheduling method for collaborative sensing disclosed in the embodiments of the present application;

[0056] Figure 2 It is a schematic flowchart of another energy-efficient sensor scheduling method for collaborative sensing disclosed in the embodiments of the present application;

[0057] Figure 3 It is a schematic flowchart of another energy-efficient sensor scheduling method for collaborative sensing disclosed in the embodiments of the present application;

[0058] Figure 4 It is a schematic flowchart of another energy-efficient sensor scheduling method for collaborative sensing disclosed in the embodiments of the present application;

[0059] Figure 5 It is a schematic flowchart of another energy-efficient sensor scheduling method for collaborative sensing disclosed in the embodiments of the present application;

[0060] Figure 6 It is a schematic structural diagram of an energy-efficient sensor scheduling system for collaborative sensing disclosed in the embodiments of the present application;

[0061] Figure 7 This is a schematic structural diagram of an energy - efficient sensor scheduling device for collaborative perception disclosed in the embodiments of the present application. Detailed implementation manners

[0062] The terms "first", "second", "third", "fourth", etc. (if any) in the description, claims and drawings of the present application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order different from that illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non - exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily limit to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these process, method, product or device.

[0063] It should be noted that the descriptions involving "first", "second", etc. in the present application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. Additionally, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present application.

[0064] In applications such as the Internet of Things and digital twins, collaborative sensing technology collects data from a target spatial area through widely distributed sensors and transmits it to a target server to achieve real-time monitoring of the physical environment. To ensure effective monitoring of the real-time physical environment, high-coverage sensing data and timely data transmission are required, that is, receiving high-quality sensing data. However, the battery energy of sensors is limited, resulting in their inability to continuously perform data sensing and transmission. The cyclic working mode of sensors allows sensors to switch between the working and sleeping states, which can save the energy consumption of sensors to a certain extent. However, how to schedule sensors in the network when sensors operate in a cyclic working mode to achieve a balance between energy efficiency and the quality of sensing data is a key challenge. A large number of sensors in the working mode can ensure high coverage of the target area but result in high energy consumption. There is overlap between the coverage areas of sensors, that is, redundant coverage, which provides the possibility to control the working mode of sensors, reduce redundant coverage, and thus achieve energy saving. Therefore, most existing works adopt a local coverage mechanism to schedule the working mode of sensors to achieve energy saving while ensuring the coverage rate at the target sensing points. Timely data transmission can ensure the real-time update of environmental sensing data. However, sensors in the sleeping state reduce the freshness of data to a certain extent. Moreover, sensors generally use power control to avoid wasting transmission energy consumption, but reducing the transmission power also causes transmission packet loss, thereby reducing the freshness of the sensing data at the server. At the same time, in the existing technology, sensor scheduling is usually based on a simplified channel model and uses heuristic algorithms to achieve energy saving. However, these methods may not accurately reflect the actual dynamic channel conditions and sensor working modes, resulting in inaccurate scheduling strategies. During the process of multi-source data transmission to the collaborative sensing center, due to the dynamic nature of the channel, there is a risk of transmission failure, which will increase the AoI and reduce the freshness of data. A large part of the sensor scheduling research focuses on minimizing the data freshness at the receiving end while ignoring the energy consumption problem at the sensor end and not fully considering how to optimize the energy consumption of sensors while ensuring data freshness.

[0065] Therefore, the embodiments of this application propose an intelligent high-energy efficiency sensor scheduling scheme for collaborative sensing. The following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0066] It should be noted that the embodiments of the present application are mainly directed to the collaborative perception task of N sensors monitoring M target points in a certain area within a target server (or described as a central server, i.e., the central server in the target area), where the density of the sensors is high enough to ensure that all target points can be covered. There is an overlapping coverage situation among the sensing areas of the sensors, so some target points may be monitored by more than two sensors. The sensors switch between two working modes, namely, sleep and activation, so as to reduce redundant coverage and save energy consumption. Among them, the target server has a hierarchical network architecture. Among them, the upper layer includes a policy network (upper-layer policy network) composed of four fully connected layers and an invalid action masking module with weight parameters of θ (k) and a value network (upper-layer value network) composed of four fully connected layers with weight parameters of ω (k) The lower layer includes a policy network (lower-layer policy network) composed of four fully connected layers with weight parameters of ψ (k) and a value network (lower-layer value network) composed of four fully connected layers with weight parameters of φ (k) It should also be noted that k is the moment k (or the number of iterations, that is, one iteration is one calculation moment k). Further, before executing the embodiments of the present application, that is, before starting the first iteration, it is necessary to initialize the weight parameters and related parameters. The weight parameters of the above-mentioned upper-layer policy network, upper-layer value network, lower-layer policy network, and lower-layer value network are initialized to θ (0) ω (0) ψ (0) φ (0) . The remaining parameters include the number B of subsequent sensor scheduling experiences and the learning rate β

[0067] To solve the technical problems described above, please refer to Figure 1 Figure 1 which is a schematic flowchart of an energy-efficient sensor scheduling method for collaborative perception disclosed in the embodiments of the present application. It includes steps 101-step 106

[0068] 101. Obtain the current state information of each sensor at the current moment

[0069] After constructing the hierarchical network architecture, at moment k (i.e., the current moment), the target server obtains the current state information of each sensor. It should be noted that the current state information includes the current working mode of each sensor at the current moment, the current channel state between the target server and each sensor, and the current age of information of each sensor for any target monitoring point

[0070] In one specific embodiment, the target server obtains the current working mode of each sensor ​Estimate the channel state between itself and each sensor Evaluate the age of information of the sensed data at each target point Denoted as the state It should be noted that, v (k) , h (k) and A (k) represent the working mode, channel state or age of information as vectors. Additionally, N is the number of sensors, and n represents sensor n; In [ ], m is the target point to be monitored, and the corresponding M is the number of target points to be monitored.

[0071] Furthermore, v (k) represents the working mode, which is divided into two types: sleep mode and active mode, and can be expressed as h represents the channel state, which can also be said to be the channel gain. Generally, it is a number between 0 and 1, and can be expressed as A represents the age of information, and the range is from 0 to A max , and can be expressed as

[0072] 102. According to the current working mode, current channel state, and current age of information, sample the sleep probability distribution for each sensor to obtain the current sleep duration corresponding to each sensor in the upper-level policy network.

[0073] Based on the current state information obtained in step 101, the sleep probability distribution can be sampled for each sensor according to the current working mode, current channel state, and current age of information therein, so as to obtain the current sleep duration corresponding to each sensor in the upper-level policy network.

[0074] In one specific embodiment, by determining whether the working mode of the sensor at the current moment is the sleep mode or the active mode according to the current working mode of each sensor, and then determining the channel gain between the sensor and the server according to the channel state therein. Since each sensor can cover at least 1 target monitoring point (target point) when in the active state, therefore, the working modes of different sensors within different sleep times are counted as an action set, so as to obtain the probability distribution corresponding to the sleep probability according to the action set, perform sampling of the probability distribution of the sleep probability, and finally obtain the sleep duration of each sensor at the current moment in the upper-level policy network. It should be noted that the current sleep duration at this time should be understood as a vector.

[0075] Furthermore, specifically, reference can be made to Figure 2 the embodiment shown

[0076] 103. Input the current working mode, current channel state, current information age, and current sleep duration into the lower-layer policy network to obtain the probability distribution of the transmission power corresponding to each sensor, and perform power probability distribution sampling according to the probability distribution of the transmission power to obtain the current transmission power corresponding to each sensor.

[0077] Based on step 102, the current working mode, current channel state, current information age, and current sleep duration can be input into the lower-layer policy network, thereby obtaining the probability distribution of the transmission power corresponding to each sensor, and then performing power probability distribution sampling according to the probability distribution of the transmission power to obtain the current transmission power corresponding to each sensor.

[0078] In one specific embodiment, after obtaining the current sleep duration of all sensors, the current working mode, current channel state, current information age in the current state information, and the current sleep duration selected by the upper-layer policy network can be input into the lower-layer policy network to obtain the mean vector. Then, combined with the variance vector, the probability distribution of the transmission power of each sensor can be obtained, and the transmission power of each sensor can be obtained by sampling according to the probability distribution of the transmission power. It should be noted that the current transmission power at this time is a vector.

[0079] For ease of understanding, reference can be made to Figure 3 Steps 301 - 303 of the illustrated embodiment.

[0080] 104. Send the current sleep duration and the current transmission power to each corresponding sensor so that each sensor sends sensing information to the target server according to the current scheduling policy of the target server.

[0081] Based on step 103, the current sleep duration and the current transmission power can be sent to each corresponding sensor, so that each sensor sends sensing information to the target server according to the current scheduling policy of the target server.

[0082] In one specific embodiment, the target server sends the current sleep duration and the current transmission power of each sensor to the corresponding sensor together. Thus, each sensor can sleep for a certain duration (the current sleep duration) according to the scheduling policy of the target server, or control the sensor to transmit sensing information to the target server at the current transmission power.

[0083] For ease of understanding, reference can be made to Figure 3 Steps 304 - 306 of the illustrated embodiment.

[0084] 105. Update the current information age of all target monitoring points according to the sensed information to obtain the target information age, and obtain the benefit function value of each sensor at the current moment according to the target information age.

[0085] Based on step 104, according to the received sensed information, the target server can update the current information age of all target monitoring points, thereby obtaining the target information age, and obtaining the benefit function value of each sensor at the current moment according to the target information age.

[0086] In one specific embodiment, after receiving the sensed information, the target server can update the information age of each target point. Specifically, it can be adjusted and modified based on whether the sensed information is received as a precondition. For example, if the target server receives the sensed information for a target point, the information age of that target point can be adjusted to 1. If the target server receives the sensed information for a target point, add 1 to the current information age of that target point as the target information age. For details, please refer to Figure 3 Steps 304 - 306 of the illustrated embodiment.

[0087] Furthermore, based on the target information age and the power consumption of each sensor, the energy consumption of each current sensor can be estimated, thereby obtaining the benefit function value of each sensor at the current moment. For details, please refer to Figure 4 The illustrated embodiment.

[0088] 106. Set the current state information, current sleep duration, current transmission power, and benefit function value as the sensor scheduling experience, and store the sensor scheduling experience in the experience pool according to the preset storage rule, so as to sample a preset number of sensor scheduling experiences from the experience pool to update the current weight parameter values of the upper policy network, upper value network, lower policy network, and lower value network until the preset time threshold is met, and use the target weight parameter value corresponding to the preset time threshold as the basis for the target server to select the target scheduling strategy for each sensor.

[0089] Based on step 105, the current state information, current sleep duration, current transmission power, and benefit function value can be set as the sensor scheduling experience, and the sensor scheduling experience can be stored in the experience pool according to the preset storage rule, so as to sample a preset number of sensor scheduling experiences from the experience pool to update the current weight parameter values of the upper policy network, upper value network, lower policy network, and lower value network. Then, loop through steps 101 - 106 until the number of iterations or the current time meets the preset time threshold (maximum time). Thus, the target weight parameter value corresponding to the preset time threshold can be used as the target scheduling strategy of the target server for each sensor.

[0090] In one specific embodiment, the state information, sleep duration, transmission power, and benefit function at the current moment are used as a sensor scheduling experience, and then stored in the experience pool in sequence according to the first-in, first-out rule. Then, a preset number of experiences (sensor scheduling experiences) are sampled from the experience pool, and the weight parameters of the policy network and value network in the hierarchical architecture are updated respectively using the stochastic gradient descent method with a learning rate of β. Specifically, the policy network and value network at different layers each have weight parameters. In this embodiment, the weight parameters of the value network at the current moment can affect the weight parameters of the policy network and value network at the next moment. For specific details, please refer to Figure 5 the embodiments shown.

[0091] Through an energy-efficient sensor scheduling method for cooperative sensing disclosed in this embodiment, the sensor scheduling strategy is decomposed into two sub-strategies: sleep cycle and transmission power, which solves the problem of benefit trade-off of sensing information in a dynamic channel situation. Under certain coverage constraints, it can select the sleep cycle and transmission power of the sensor to adapt to the change of the channel environment, thereby improving the freshness of sensing data and reducing energy consumption.

[0092] For further detailed description Figure 1 of step 102 shown, please refer to Figure 2 , Figure 2 which is a schematic flowchart of another energy-efficient sensor scheduling method for cooperative sensing disclosed in the embodiments of the present application. It includes steps 201 - step 205.

[0093] 201. Set the maximum sleep duration corresponding to all sensors.

[0094] Based on step 101 in the above Figure 1 shown embodiment, after obtaining the current state information s of each sensor (k) , the maximum sleep duration corresponding to all sensors can be set. It should be noted that since the current working mode includes a sleep mode and an active mode, the maximum sleep duration can be used to characterize the maximum continuous time of each sensor in the sleep mode. That is to say, the maximum sleep duration is mainly for the sleep mode of the sensor. Further, setting the maximum sleep duration of the sensor can also be performed before step 101 is executed, and specific details are not limited here.

[0095] In one specific embodiment, let J be the maximum sleep time. Among them, the maximum sleep time can be set according to the rated working parameters of the sensor, and specific details are not elaborated here. For example, in this embodiment, the setting of the maximum sleep time can be 1 s or 2 s, etc.

[0096] It should be noted that Figure 1For the character descriptions of each parameter in the illustrated embodiment, the same will be referenced later and will not be elaborated on hereafter.

[0097] 202. Determine a target action set that satisfies the coverage constraint condition within the maximum sleep duration according to the current working mode of each sensor and the initial action set of each sensor.

[0098] Based on step 201, according to the current working mode of each sensor and the initial action set of each sensor, a target action set that satisfies the coverage constraint condition within the maximum sleep duration can be determined. It should be noted that the initial action set is used to represent the action set of any sensor in its current working mode within the maximum sleep duration; the coverage constraint condition is used to represent the action set of sensors that satisfy the preset coverage rate and are in the active mode at the current moment.

[0099] In one specific embodiment, the working modes of current sensors can be used to determine an invalid action set that violates the coverage constraint condition (such as sensors with excessive sleep). Specifically, an effective action set N can be determined from the original action space (initial action set) Ω1 = [0, 1,..., J] i.e., the target action set. It can be understood that Ω1 at this time or is a kind of vector. Specifically, for example, if the current working mode is that 5 sensors are in the active state and each of them can cover 1 target point, and the coverage constraint is to ensure an 80% coverage rate, then at least 4 sensors need to be in the active state. The invalid action set includes actions that make more than one sensor sleep, such as the sleep duration of all sensors being greater than 0 (where 0 is the maximum sleep duration), the sleep duration of 4 sensors being greater than 0, the sleep duration of 3 sensors being greater than 0, and the sleep duration of 2 sensors being greater than 0, all of which are invalid action sets.

[0100] 203. Input the current state information and the target action set into the upper-layer policy network to obtain an action dimension vector corresponding to the number of sensors in the target action set, so as to determine a masking vector according to the action dimension vector.

[0101] Based on step 202, then input the current state information and the target action set into the upper-layer policy network, so as to obtain an action dimension vector corresponding to the number of sensors in the target action set, and determine a masking vector according to the action dimension vector. It should be noted that the masking vector is used to represent the result of each sensor satisfying the coverage constraint condition at the current moment.

[0102] In one specific embodiment, input the current state information s (k)Input it into the upper-layer policy network to obtain the |Ω1|-dimensional vector output by the upper-layer policy network Then, the masking vector can be determined according to the various parameters determined above It should be noted that in this embodiment, i represents the i-th sleep duration in a single sensor in the vector space. Specifically, there is

[0103] Thus, the calculation of the masking vector is completed

[0104] 204. Determine the probability distribution of the sleep probability corresponding to different sleep durations for each sensor according to the masking vector, the action dimension vector, the current working mode, the current channel state, and the current age of information

[0105] Based on step 203, probability distribution analysis can be performed on the masking vector, the action dimension vector, the current working mode, the current channel state, and the current age of information, so as to obtain the probability distribution of the sleep probability corresponding to different sleep durations for each sensor. It should be noted that the probability distribution of the sleep probability is related to the current weight parameter value of the upper-layer policy network

[0106] In one specific embodiment, the formula for the probability distribution of the sleep probability is

[0107] It is not difficult to understand that f(s (k) , i; θ (k) ) is the probability distribution of the sleep probability described above

[0108] 205. Perform sampling on the probability distribution of the sleep probability to obtain the current sleep duration

[0109] Therefore, based on formula (2) in step 204 for probability distribution sampling, the current sleep duration can be obtained

[0110] In one specific embodiment, based on formula (2), the probabilities of different sleep duration vectors can be calculated, and sampling is performed according to the probability distribution to obtain the vector a of the sleep duration (current sleep duration) (k) . The action set of the sleep duration vector is [0, 0, 0..., 0], [0, 0, 0..., 1],..., [0, 0, 0..., J],..., [J, J, J..., J], where [0, 0, 0..., 0] is an n-dimensional vector (related to the number n of sensors)

[0111] A high - energy - efficient sensor scheduling method for collaborative sensing disclosed by this embodiment calculates the vector probability of different sleep durations, samples according to the probability distribution to obtain a sleep duration vector, thereby decomposing the sensor scheduling strategy into sub - strategies of the sleep cycle, and can select a sleep cycle for each sensor while satisfying the coverage constraint, so as to adapt to the change of the channel environment.

[0112] For further detailed description Figure 1 of steps 103 - 105 shown Figure 3 , Figure 3 Refer to

[0113] 301. Input the current working mode, current channel state, current age of information, and current sleep duration into the lower - layer policy network to obtain a mean vector.

[0114] Corresponding to step 103 in the above - mentioned Figure 1 shown embodiment, in this embodiment, the state information s (k) (including the current working mode, current channel state, current age of information) and the sleep duration a (k) selected by the upper - layer policy network are input into the lower - layer policy network, thereby obtaining the mean vector [μ(s (k) , a (k) , n; ψ (k) )] 1≤n≤N . It should be noted that μ is the coefficient for obtaining the mean vector, and the mean vector is related to the current weight parameter value of the lower - layer policy network at the current moment.

[0115] 302. Based on the mean vector, combined with the variance vector of each corresponding sensor, obtain the probability distribution of the transmission power of each sensor.

[0116] Based on step 301, the probability distribution of the transmission power of each sensor can be obtained through the mean vector and combined with the variance vector of each sensor.

[0117] In one specific embodiment, the target server combines the mean vector and the given variance vector [X n 1≤n≤N , to obtain the probability distribution of the transmission power of each sensor. Among them, the probability distribution of the transmission power is

[0118] It should be noted that formula (3) is the joint probability density function of the multivariate Gaussian distribution. An n - dimensional vector p′ is sampled according to the probability density function in [-∞, +∞]. The p′ in f() refers to the probability of sampling to p′.​

[0119] 303. Sample the power probability distribution according to the probability distribution of the transmission power to obtain the current transmission power corresponding to each sensor.

[0120] Then, based on the probability distribution of the transmission power in step 302, perform sampling of the power probability distribution, so as to obtain the current transmission power corresponding to each sensor.

[0121] In one specific embodiment, the transmission power p of each sensor is obtained by sampling according to the probability distribution of the above formula (3). (k) . Specifically, truncate according to the sampling result to obtain p. (k) . The specific method of truncation is to set the elements in p' that are greater than the maximum transmission power P max to P max , and set the elements less than 0 to 0. P max is the rated maximum power of the sensor.

[0122] 304. Send the current sleep duration and the current transmission power to each corresponding sensor, so that each sensor sends sensing information to the target server according to the current scheduling policy of the target server.

[0123] Then, based on step 303, step 304 can be executed. Specifically, refer to Figure 1 step 104 in the shown embodiment, which will not be elaborated here. However, it should be noted that in this embodiment, there are the following two judgment bases: if the current scheduling policy is a sleep policy, control each sensor to sleep for the corresponding current sleep duration according to the sleep policy; or, if the current scheduling policy is an activation policy, control each sensor so that each sensor sends sensing information to the target server based on the corresponding current transmission power. Specifically, the target server sends the sleep duration a (k) and the transmission power p (k) of the sensor to the corresponding sensor together. Each sensor n sleeps for a certain duration according to the scheduling policy of the target server (if ), otherwise the sensor senses the current environment and transmits sensing information to the target server at power .

[0124] Then, execute step 305 or step 306 respectively according to whether the sensing information is received.

[0125] 305. If the target server receives the sensing information of any target monitoring point, update the current information age to a preset constant factor according to the sensing information, and use the preset constant factor as the target information age.

[0126] Based on step 304, when the target server receives the perception information of any target monitoring point, it can update the current information age to a preset constant factor according to the perception information, and use the preset constant factor as the target information age.

[0127] In one specific embodiment, after the target server receives the perception information of target point m, it can update the current information age to 1 according to the perception information, and use 1 as the target information age of the current target point m. It is not difficult to understand that 1 at this time is the preset constant factor described above.

[0128] 306. If the target server does not receive the perception information of any target monitoring point, add the current information age and the preset constant factor according to the perception information, and use the added value as the target information age.

[0129] Based on step 304, when the target server does not receive the perception information of any target monitoring point, it can add the current information age and the preset constant factor according to the perception information, so as to use the added value as the target information age.

[0130] In one specific embodiment, when the target server does not receive the perception information of target point m, it can add 1 to the previous information age of target point m according to the perception information, so as to use the added value as the target information age of target point m.

[0131] Therefore, combining step 305 and step 306, there is It should be noted that m represents the serial number of the target point, and there are a total of M target points (target monitoring points).

[0132] Through an energy-efficient sensor scheduling method for cooperative perception disclosed in this embodiment, combined with the above Figure 2 illustrated embodiments, the freshness of the perception information can be effectively improved, thereby effectively reducing energy consumption (see Figure 4 ).

[0133] Combined with Figure 3 illustrated embodiments, to further elaborate in detail Figure 1 step 105 shown, please refer to Figure 4 , Figure 4 This is a schematic flowchart of another energy-efficient sensor scheduling method for cooperative perception disclosed in the embodiments of the present application. It includes step 401-step 403.

[0134] 401. Determine the power loss value corresponding to each sensor.

[0135] In the process of obtaining the benefit function value at the current moment, it is necessary to first determine the power loss value corresponding to each sensor.

[0136] 402. Perform approximate estimation based on the power loss value to obtain the sensor energy consumption value of each sensor at the current moment.

[0137] Based on step 401, approximate estimation can be performed based on the power loss value to obtain the sensor energy consumption value of each sensor at the current moment.

[0138] In one specific embodiment, the energy consumption of each current sensor can be estimated according to the power loss. Specifically, the power loss can be used to approximate the energy consumption, which will not be elaborated here.

[0139] 403. Obtain the benefit function value according to the sensor energy consumption value, the age of the target information, and the number of all corresponding sensors and the number of target monitoring points.

[0140] Thus, the benefit function value can be calculated according to the sensor energy consumption value, the age of the target information, and the number of all corresponding sensors and the number of target monitoring points.

[0141] In one specific embodiment, through calculation with the above parameters, we have

[0142] where it should be noted that η is a weight parameter used to measure the importance of energy consumption compared to AoI.

[0143] Through a high - energy - efficient sensor scheduling method for cooperative sensing disclosed in this embodiment, combined with the above - mentioned Figure 2 and Figure 3 shown embodiments, the sensor scheduling strategy is further strengthened, and the sensor scheduling strategy is decomposed into another sub - strategy of transmission power. Through adaptive selection by reinforcement learning, the computational complexity is effectively reduced. At the same time, the trade - off problem between sensor energy conservation and the quality of sensed data in a dynamic channel situation is also solved. It can automatically select the sleep cycle and transmission power of sensors under a certain coverage constraint to adapt to the change of the channel environment, thereby improving the freshness of sensed data and reducing energy consumption.

[0144] For further detailed description of Figure 1 step 106 shown, Figure 5 is a schematic flowchart of another high - energy - efficient sensor scheduling method for cooperative sensing disclosed in the embodiments of the present application. It includes steps 501 - step 505.

[0145] 501. Analyze a preset number of sensor scheduling experiences to obtain the first scheduling experience and the second scheduling experience.

[0146] As described above, the current weight parameter values include the current upper-layer policy weight value of the upper-layer policy network corresponding to the current moment, the current upper-layer value weight value of the upper-layer value network, the lower-layer policy weight value of the lower-layer policy network, and the current lower-layer value weight value of the lower-layer value network. In this embodiment, specifically for Figure 1 the detailed description of step 106 in the illustrated embodiment. Specifically, before analyzing the sensor scheduling experience, it is also necessary to store the sensor scheduling experience (s (k) ,a (k) ,p (k) ,u (k) ) into the experience pool D according to the first-in-first-out rule, denoted as D←D∪(s (k) ,a (k) ,p (k) ,u (k) )(6).

[0147] Then, B experiences can be sampled from the experience pool, and the weight parameters of the policy network and the value network in the hierarchical architecture can be updated respectively using the stochastic gradient descent method with a learning rate of β.

[0148] Specifically, in this embodiment, at least two scheduling experiences can be retrieved, namely the first scheduling experience and the second scheduling experience. It should be noted that the second scheduling experience is the next scheduling experience of the first scheduling experience, and the first scheduling experience and the second scheduling experience at least include the current state information, the current sleep duration, the current transmission power, and the benefit function value.

[0149] Furthermore, this embodiment does not limit the number of retrieved sensor scheduling experiences either. For the sake of easy understanding, the first scheduling experience and the second scheduling experience are used as examples for illustration. However, it should be understood that at this time, the first scheduling experience and the second scheduling experience only differ in the order of storage in the experience pool D, and the number of corresponding first scheduling experiences and second scheduling experiences is not limited. Since both the first scheduling experience and the second scheduling experience are taken from the experience pool D, they both contain (s (k) ,a (k) ,p (k) ,u (k) ). It should be noted that steps 502 - 505 in this embodiment can be executed simultaneously or not simultaneously, and this embodiment does not limit the execution order of steps 502 - 505.

[0150] 502. Based on the stochastic gradient descent algorithm, calculate the current sleep duration, the current upper-layer policy weight value, the current state information, and the first state value function of each sensor corresponding to the first scheduling experience to obtain the upper-layer policy weight value at the next moment.

[0151] Based on step 501, the current sleep duration, current upper-layer policy weight value, current state information, and the first state value function of each sensor corresponding to the first scheduling experience can be calculated based on the stochastic gradient descent algorithm to obtain the upper-layer policy weight value at the next moment. It should be noted that the state value function at the first moment is related to the current state information and the current upper-layer value weight value corresponding to the first scheduling experience. Specifically, there is

[0152] where y is used to represent the y-th sensor scheduling experience in the experience pool, that is, the first scheduling experience. β is the learning rate. It can be seen from formula (7) that the upper-layer policy weight value θ (k) and the upper-layer value weight value ω (k) at the current moment can affect the upper-layer policy weight value θ (k+1) at the next moment. Further, at the first iteration, or at the initial moment k = 0, the calculation can be performed according to the initially set θ (0) , ω (0) . It should also be noted that V(s (y) ; ω (k) ) represents the state value function (described as the first state value function in this embodiment).

[0153] 503. Based on the stochastic gradient descent algorithm, calculate the benefit function value, the first state value function, and the second state value function corresponding to the first scheduling experience to obtain the upper-layer value weight value at the next moment.

[0154] Based on step 501, the benefit function value, the first state value function, and the second state value function corresponding to the first scheduling experience can be calculated based on the stochastic gradient descent algorithm, so as to obtain the upper-layer value weight value at the next moment. It should be noted that the second state value function is related to the current state information and the current upper-layer value weight value corresponding to the second scheduling experience. Specifically, there is

[0155] where y + 1 is used to represent the (y + 1)-th sensor scheduling experience in the experience pool, that is, the second scheduling experience, and γ represents the discount factor. It can be seen from formula (8) that the upper-layer value weight value ω (k) at the current moment can affect the upper-layer value weight value ω (k+1) at the next moment. Further, at the first iteration, or at the initial moment k = 0, the calculation can be performed according to the initially set θ (0) , ω (0) . It should also be noted that V(s (y+1) ; ω (k)) represents the state value function (described as the second state value function in this embodiment). Further, the discount factor γ represents the degree of importance.

[0156] 504. Based on the stochastic gradient descent algorithm, calculate the current sleep duration, current lower-layer policy weight value, current state information, current transmission power of each sensor corresponding to the first scheduling experience, and the first sleep action value function to obtain the lower-layer policy weight value at the next moment.

[0157] Based on step 501, based on the stochastic gradient descent algorithm, the current sleep duration, current lower-layer policy weight value, current state information, current transmission power of each sensor corresponding to the first scheduling experience, and the first sleep action value function can be calculated to obtain the lower-layer policy weight value at the next moment. It should be noted that the first sleep action value function is related to the current state information, current sleep duration, and current lower-layer value weight value corresponding to the first scheduling experience. Specifically, there is

[0158] Among them, as can be seen from formula (9), the lower-layer policy weight value ψ at the current moment (k) and the lower-layer value weight value φ at the current moment (k) can affect the lower-layer policy weight value ψ at the next moment (k+1) . Further, at the first iteration, or at the initial moment k = 0, it can be calculated according to the initially set ψ (0) , φ (0) . It should also be noted that V(s (y) , a (y) ; φ (k) ) represents the sleep action value function (described as the first sleep action value function in this embodiment).

[0159] 505. Based on the stochastic gradient descent algorithm, calculate the benefit function value, the first sleep action value function, and the second sleep action value function corresponding to the first scheduling experience to obtain the lower-layer value weight value at the next moment.

[0160] Based on step 501, based on the stochastic gradient descent algorithm, the benefit function value, the first sleep action value function, and the second sleep action value function corresponding to the first scheduling experience can be calculated to obtain the lower-layer value weight value at the next moment. It should be noted that the second sleep action value function is related to the current state information, current sleep duration, and current lower-layer value weight value corresponding to the second scheduling experience. Specifically, there is,

[0161] As can be seen from formula (10), the lower-layer value weight value φ at the current moment (k) can affect the lower-layer value weight value φ at the next moment(k+1) 。Furthermore, at the first iteration, or at the initial moment k = 0, it is possible to calculate based on the initially set ψ (0) , φ (0) . It should also be noted that V(s (y+1) , a (y+1) ; φ (k) ) represents the dormant action value function (described as the second dormant action value function in this embodiment).

[0162] Combined with the above steps 502 - step 505, it is also possible to iteratively execute the above process until k = K. Thus, the confirmation of the weight parameters of each network structure of the final target server is ended.

[0163] Through an energy - efficient sensor scheduling method for collaborative sensing disclosed in this embodiment, a hierarchical structure is adopted, and the sensor scheduling strategy is decomposed into two sub - strategies: the dormant period and the transmission power. Through reinforcement learning for adaptive selection, the computational complexity is effectively reduced. At the upper layer, a policy network combined with an invalid action mask module is designed to select the sleep period for each sensor on the premise of meeting the coverage constraint. At the lower layer, based on the policy - value network architecture, the transmission power is selected for each sensor in the working mode to ensure successful transmission, thereby reducing the age of information.

[0164] It should be understood that although the steps in the flowcharts involved in the above - mentioned embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above - mentioned embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0165] Please refer to Figure 6 , Figure 6 , which is a schematic structural diagram of an energy - efficient sensor scheduling system for collaborative sensing disclosed in an embodiment of the present application.

[0166] An acquisition unit 601 is configured to acquire the current state information of each sensor at the current moment; wherein, the current state information includes the current working mode of each sensor at the current moment, the current channel state between the target server and each sensor, and the current age of information of each sensor for any target monitoring point;

[0167] The sampling unit 602 is configured to sample the sleep probability distribution for each sensor according to the current working mode, the current channel state, and the current age of information, so as to obtain the current sleep duration corresponding to each sensor in the upper-layer policy network;

[0168] The input unit 603 is configured to input the current working mode, the current channel state, the current age of information, and the current sleep duration into the lower-layer policy network, so as to obtain the probability distribution of the transmission power corresponding to each sensor, and sample the power probability distribution according to the probability distribution of the transmission power, so as to obtain the current transmission power corresponding to each sensor;

[0169] The sending unit 604 is configured to send the current sleep duration and the current transmission power to each corresponding sensor, so that each sensor sends the sensed information to the target server according to the current scheduling policy of the target server; wherein, the scheduling policy has an associated relationship with the current sleep duration;

[0170] The updating unit 605 is configured to update the current age of information of all target monitoring points according to the sensed information to obtain the target age of information, and obtain the benefit function value of each sensor at the current moment according to the target age of information; wherein, the benefit function value has an associated relationship with the power consumption value of each sensor;

[0171] The setting unit 606 is configured to set the current state information, the current sleep duration, the current transmission power, and the benefit function value as the sensor scheduling experience, and store the sensor scheduling experience in the experience pool according to a preset storage rule, so as to sample a preset number of sensor scheduling experiences from the experience pool, and update the current weight parameter values of the upper-layer policy network, the upper-layer value network, the lower-layer policy network, and the lower-layer value network until a preset time threshold is met, and use the target weight parameter value corresponding to the preset time threshold as the selection basis for the target server for the target scheduling policy of each sensor.

[0172] Exemplarily, the system further includes: a determining unit 607;

[0173] The setting unit 606 is further configured to set the maximum sleep duration corresponding to all sensors; wherein, the current working mode includes a sleep mode and an activation mode, and the maximum sleep duration is used to represent the maximum continuous time for each sensor in the sleep mode;

[0174] The determining unit 607 is configured to determine, according to the current working mode of each sensor and the initial action set of each sensor, a target action set that satisfies the coverage constraint condition within the maximum sleep duration; wherein, the initial action set is used to represent the action set of the current working mode of any sensor within the maximum sleep duration; the coverage constraint condition is used to represent the action set of the sensors that satisfy the preset coverage rate and are in the activation mode at the current moment;

[0175] The input unit 603 is further configured to input the current status information and the target action set into the upper-layer policy network, obtain an action dimension vector corresponding to the number of sensors in the target action set, and determine a masking vector according to the action dimension vector; wherein, the masking vector is used to represent the result of each sensor satisfying the coverage constraint condition at the current moment.

[0176] Exemplarily, the system includes:

[0177] The determination unit 607 is specifically configured to determine a probability distribution of the sleep probability corresponding to different sleep durations for each sensor according to the masking vector, the action dimension vector, the current working mode, the current channel state, and the current age of information; wherein, the probability distribution of the sleep probability is associated with the current weight parameter value of the upper-layer policy network;

[0178] The sampling unit 602 is specifically configured to perform sampling on the sleep probability distribution according to the probability distribution of the sleep probability to obtain the current sleep duration.

[0179] Exemplarily, the system includes:

[0180] The input unit 603 is specifically configured to input the current working mode, the current channel state, the current age of information, and the current sleep duration into the lower-layer policy network to obtain a mean vector; wherein, the mean vector is associated with the current weight parameter value of the lower-layer policy network at the current moment;

[0181] The acquisition unit 601 is specifically configured to obtain a probability distribution of the transmission power of each sensor based on the mean vector and in combination with the variance vector of each corresponding sensor.

[0182] Exemplarily, the system further includes: a control unit 608;

[0183] The control unit 608 is configured to, when the current scheduling policy is a sleep policy, control each sensor to sleep for the corresponding current sleep duration according to the sleep policy;

[0184] The control unit 608 is further configured to, when the current scheduling policy is an activation policy, control each sensor so that each sensor sends sensing information to the target server based on the corresponding current transmission power.

[0185] Exemplarily, the system includes:

[0186] The update unit 605 is further configured to, when the target server receives the sensing information of any target monitoring point, update the current age of information to a preset constant factor according to the sensing information, and use the preset constant factor as the target age of information;

[0187] The obtaining unit 601 is further configured to, when the target server does not receive the perception information of any target monitoring point, add the current information age to a preset constant factor according to the perception information, and use the added value as the target information age.

[0188] Exemplarily, the system includes:

[0189] The determining unit 607 is specifically configured to determine the power loss value corresponding to each sensor;

[0190] The obtaining unit 601 is specifically configured to perform an approximate estimation based on the power loss value to obtain the sensor energy consumption value of each sensor at the current moment;

[0191] The obtaining unit 601 is further configured to obtain a benefit function value according to the sensor energy consumption value, the target information age, and the number of all sensors and the number of target monitoring points corresponding thereto.

[0192] Exemplarily, the current weight parameter value includes the current upper-layer policy weight value of the upper-layer policy network corresponding to the current moment, the current upper-layer value weight value of the upper-layer value network, the lower-layer policy weight value of the lower-layer policy network, and the current lower-layer value weight value of the lower-layer value network. The system further includes: a calculation unit 609;

[0193] The obtaining unit 601 is specifically configured to analyze a preset number of sensor scheduling experiences to obtain a first scheduling experience and a second scheduling experience; wherein, the second scheduling experience is the next scheduling experience of the first scheduling experience, and the first scheduling experience and the second scheduling experience at least include the current state information, the current sleep duration, the current transmission power, and the benefit function value;

[0194] The calculation unit 609 is configured to calculate the current sleep duration, the current upper-layer policy weight value, the current state information, and the first state value function of each sensor corresponding to the first scheduling experience based on the stochastic gradient descent algorithm to obtain the upper-layer policy weight value at the next moment; wherein, the state value function at the first moment is related to the current state information and the current upper-layer value weight value corresponding to the first scheduling experience;

[0195] The calculation unit 609 is further configured to calculate the benefit function value, the first state value function, and the second state value function corresponding to the first scheduling experience based on the stochastic gradient descent algorithm to obtain the upper-layer value weight value at the next moment; wherein, the second state value function is related to the current state information and the current upper-layer value weight value corresponding to the second scheduling experience;

[0196] The calculation unit 609 is further configured to calculate, based on the stochastic gradient descent algorithm, the current sleep duration, the current lower-layer policy weight value, the current state information, the current transmission power, and the first sleep action value function of each sensor corresponding to the first scheduling experience, to obtain the lower-layer policy weight value at the next moment; wherein, the first sleep action value function is associated with the current state information, the current sleep duration, and the current lower-layer value weight value corresponding to the first scheduling experience.

[0197] The calculation unit 609 is further configured to calculate, based on the stochastic gradient descent algorithm, the benefit function value, the first sleep action value function, and the second sleep action value function corresponding to the first scheduling experience, to obtain the lower-layer value weight value at the next moment; wherein, the second sleep action value function is associated with the current state information, the current sleep duration, and the current lower-layer value weight value corresponding to the second scheduling experience.

[0198] Please refer to Figure 7 , the structural schematic diagram of an energy-efficient sensor scheduling device for cooperative sensing disclosed in an embodiment of the present application includes:

[0199] A central processing unit 701, a memory 705, an input / output interface 704, a wired or wireless network interface 703, and a power supply 702;

[0200] The memory 705 is a transient storage memory or a persistent storage memory;

[0201] The central processing unit 701 is configured to communicate with the memory 705 and execute the instruction operations in the memory 705 to execute the energy-efficient sensor scheduling method for cooperative sensing in any of the foregoing Figures 1 to 5 illustrated embodiments.

[0202] An embodiment of the present application further provides a chip system, characterized in that the chip system includes at least one processor and a communication interface, the communication interface and the at least one processor are interconnected by a line, and the at least one processor is configured to run a computer program or instruction to execute the energy-efficient sensor scheduling method for cooperative sensing in any of the foregoing Figures 1 to 5 illustrated embodiments.

[0203] An embodiment of the present application further provides a computer-readable storage medium, the computer-readable storage medium includes instructions, when the instructions run on a computer, the computer is enabled to execute the energy-efficient sensor scheduling method for cooperative sensing in any of the foregoing Figures 1 to 5 illustrated embodiments.

[0204] An embodiment of the present application further provides a computer program product including instructions, when the computer program product runs on a computer, the computer is enabled to execute the energy-efficient sensor scheduling method for cooperative sensing in any of the foregoing Figures 1 to 5A high-energy efficiency sensor scheduling method for collaborative perception in any of the illustrated embodiments.

[0205] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0206] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces, devices, or units, and can be in electrical, mechanical, or other forms.

[0207] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0208] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0209] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

Claims

1. A sensor scheduling method for collaborative sensing, characterized in that: Applied to a target server, the target server is composed of an upper layer policy network, an upper layer value network, a lower layer policy network and a lower layer value network, the method comprising: Acquire current status information of each sensor at the current moment; wherein the current status information includes the current working mode of each sensor at the current moment, the current channel state between the target server and each sensor, and the current information age of each sensor for any target monitoring point; According to the current working mode, the current channel state and the current information age, sampling the sleep probability distribution of each sensor is performed to obtain the current sleep duration corresponding to each sensor in the upper layer strategy network; Inputting the current working mode, the current channel state, the current information age and the current sleep duration into the lower layer strategy network to obtain the probability distribution of the transmit power corresponding to each sensor, and performing power probability distribution sampling according to the probability distribution of the transmit power to obtain the current transmit power corresponding to each sensor; Sending the current sleep duration and the current transmit power to each corresponding sensor, so that each sensor sends the sensing information to the target server according to the current scheduling policy of the target server; wherein the scheduling policy is associated with the current sleep duration; The current information age of all target monitoring points is updated according to the perception information to obtain the target information age, and the benefit function value of each sensor at the current moment is obtained according to the target information age; wherein the benefit function value is associated with the power consumption value of each sensor; The current state information, the current sleep time, the current transmit power and the benefit function value are set as sensor scheduling experience, and the sensor scheduling experience is stored in an experience pool according to a preset storage rule, so as to sample a preset number of sensor scheduling experiences in the experience pool, and update the current weight parameter values ​​of the upper-layer strategy network, the upper-layer value network, the lower-layer strategy network and the lower-layer value network until a preset time threshold is met, and the target weight parameter value corresponding to the preset time threshold is used as the basis for the target server to select the target scheduling strategy for each sensor.

2. The sensor scheduling method for collaborative sensing according to claim 1, characterized in that: After obtaining the current state information of each sensor at the current moment, the method further includes: Setting a maximum sleep time corresponding to all sensors; wherein the current working mode includes a sleep mode and an active mode, and the maximum sleep time is used to represent the maximum duration of each sensor in the sleep mode; According to the current working mode of each sensor and the initial action set of each sensor, a target action set that satisfies the coverage constraint condition within the maximum sleep time is determined; wherein the initial action set is used to characterize the action set of the current working mode of any sensor within the maximum sleep time; the coverage constraint condition is used to characterize the action set of the sensor that satisfies the preset coverage rate in the activation mode at the current moment; The current state information and the target action set are input into the upper-layer strategy network to obtain an action dimension vector corresponding to the number of sensors in the target action set, so as to determine a masking vector based on the action dimension vector; wherein the masking vector is used to characterize the result of each sensor satisfying the coverage constraint at the current moment.

3. The sensor scheduling method for collaborative sensing according to claim 2 is characterized in that: The step of sampling the sleep probability distribution of each sensor according to the current working mode, the current channel state, and the current information age to obtain the current sleep duration corresponding to each sensor in the upper layer strategy network includes: Determine the probability distribution of the sleep probability of each sensor corresponding to different sleep durations according to the masking vector, the action dimension vector, the current working mode, the current channel state, and the current information age; wherein the probability distribution of the sleep probability is associated with the current weight parameter value of the upper-layer strategy network; The sleep probability distribution is sampled according to the probability distribution of the sleep probability to obtain the current sleep duration.

4. The sensor scheduling method for collaborative sensing according to claim 1, characterized in that: The inputting the current working mode, the current channel state, the current information age and the current sleep duration into the lower layer strategy network to obtain a probability distribution corresponding to the transmit power of each sensor includes: Inputting the current working mode, the current channel state, the current information age and the current sleep duration into the lower layer strategy network to obtain a mean vector; wherein the mean vector is associated with the current weight parameter value of the lower layer strategy network at the current moment; Based on the mean vector and in combination with the corresponding variance vector of each sensor, a probability distribution of the transmit power of each sensor is obtained.

5. The sensor scheduling method for collaborative sensing according to claim 1, characterized in that: Before updating the current information age of all target monitoring points according to the perception information to obtain the target information age, the method further includes: If the current scheduling strategy is a sleep strategy, controlling the current sleep duration corresponding to each sensor being put into sleep according to the sleep strategy; If the current scheduling strategy is an activation strategy, each sensor is controlled according to the activation strategy so that each sensor sends the perception information to the target server based on the corresponding current transmission power.

6. The sensor scheduling method for collaborative sensing according to claim 1, characterized in that: The updating of the current information age of all target monitoring points according to the perception information to obtain the target information age includes: If the target server receives the perception information of any target monitoring point, the current information age is updated to a preset constant factor according to the perception information, and the preset constant factor is used as the target information age; If the target server does not receive the perception information of any target monitoring point, the current information age is added to the preset constant factor according to the perception information, and the added value is used as the target information age.

7. The sensor scheduling method for collaborative sensing according to claim 1, characterized in that: The obtaining of the benefit function value of each sensor at the current moment according to the target information age includes: Determining a power consumption value corresponding to each of the sensors; Perform an approximate estimation based on the power consumption value to obtain a sensor energy consumption value of each sensor at the current moment; The benefit function value is obtained according to the sensor energy consumption value, the target information age, and the corresponding number of all sensors and the number of target monitoring points.

8. The sensor scheduling method for collaborative sensing according to claim 1, characterized in that: The current weight parameter value includes the current upper strategy weight value of the upper strategy network, the current upper value weight value of the upper value network, the current lower strategy weight value of the lower strategy network and the current lower value weight value of the lower value network corresponding to the current moment, and the current weight parameter values ​​of the upper strategy network, the upper value network, the lower strategy network and the lower value network are updated by sampling a preset number of sensor scheduling experiences in the experience pool, including: Analyze the preset number of sensor scheduling experiences to obtain a first scheduling experience and a second scheduling experience; wherein the second scheduling experience is the next scheduling experience of the first scheduling experience, and the first scheduling experience and the second scheduling experience at least include current state information, the current sleep duration, the current transmit power and the benefit function value; Based on the stochastic gradient descent algorithm, the current sleep duration of each sensor corresponding to the first scheduling experience, the current upper-layer strategy weight value, the current state information and the first state value function are calculated to obtain the upper-layer strategy weight value at the next moment; wherein the first state value function is associated with the current state information corresponding to the first scheduling experience and the current upper-layer value weight value; Based on the stochastic gradient descent algorithm, the benefit function value corresponding to the first scheduling experience, the first state value function and the second state value function are calculated to obtain the upper value weight value at the next moment; wherein the second state value function is associated with the current state information corresponding to the second scheduling experience and the current upper value weight value; Based on the stochastic gradient descent algorithm, the current sleep duration, the current lower-layer strategy weight value, the current state information, the current transmit power and the first sleep action value function of each sensor corresponding to the first scheduling experience are calculated to obtain the lower-layer strategy weight value at the next moment; wherein the first sleep action value function is associated with the current state information, the current sleep duration and the current lower-layer value weight value corresponding to the first scheduling experience; Based on the stochastic gradient descent algorithm, the benefit function value corresponding to the first scheduling experience, the first sleep action value function and the second sleep action value function are calculated to obtain the lower-level value weight value at the next moment; wherein, the second sleep action value function is associated with the current state information corresponding to the second scheduling experience, the current sleep duration and the current lower-level value weight value.

9. A sensor scheduling system for collaborative sensing, characterized in that: Applied to a target server, the target server is composed of an upper layer strategy network, an upper layer value network, a lower layer strategy network and a lower layer value network, and the system includes: An acquisition unit, configured to acquire current state information of each sensor at a current moment; wherein the current state information includes a current working mode of each sensor at the current moment, a current channel state between the target server and each sensor, and a current information age of each sensor for any target monitoring point; A sampling unit, configured to perform sleep probability distribution sampling on each sensor according to the current working mode, the current channel state and the current information age, to obtain a current sleep duration corresponding to each sensor in the upper layer strategy network; An input unit, used to input the current working mode, the current channel state, the current information age and the current sleep duration into the lower layer strategy network to obtain the probability distribution of the transmit power corresponding to each sensor, and to perform power probability distribution sampling according to the probability distribution of the transmit power to obtain the current transmit power corresponding to each sensor; A sending unit, configured to send the current sleep duration and the current transmit power to each corresponding sensor, so that each sensor sends the sensing information to the target server according to the current scheduling policy of the target server; wherein the scheduling policy is associated with the current sleep duration; An updating unit, configured to update the current information age of all target monitoring points according to the perception information to obtain a target information age, and obtain a benefit function value of each sensor at the current moment according to the target information age; wherein the benefit function value is associated with the power consumption value of each sensor; A setting unit is used to set the current state information, the current sleep time, the current transmit power and the benefit function value as sensor scheduling experience, and store the sensor scheduling experience into an experience pool according to a preset storage rule, so as to sample a preset number of sensor scheduling experiences in the experience pool, update the current weight parameter values ​​of the upper strategy network, the upper value network, the lower strategy network and the lower value network until a preset time threshold is met, and use the target weight parameter value corresponding to the preset time threshold as the basis for the target server to select the target scheduling strategy for each sensor.

10. A sensor scheduling device for collaborative sensing, characterized in that: The device comprises: CPU, memory, input and output interfaces, wired or wireless network interfaces, and power supply; The memory is a transient storage memory or a persistent storage memory; The central processing unit is configured to communicate with the memory and execute instruction operations in the memory to perform the sensor scheduling method for cooperative sensing as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • State updating method based on related information age in energy harvesting Internet of Things

    CN116056033A

  • Sensor control method and device, optical detection equipment and storage medium

    CN118090733A