A motion recognition method based on multi-device collaboration based on sensory integration

By utilizing cellular traffic load and location prediction combined with deep reinforcement learning algorithms to optimize resource allocation in multi-device collaboration, the accuracy of human motion recognition in dynamic environments and the complexity of resource allocation are solved, achieving efficient motion recognition results.

CN119767262BActive Publication Date: 2025-09-26NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411924074.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-09-26
Estimated Expiration
2044-12-25

AI Technical Summary

Technical Problem

In the task of multi-device collaborative human action recognition, how to reasonably allocate resources to maximize the accuracy of human action recognition under dynamic perception environments and network status changes? Especially for devices with limited resources, how to allocate network resources between different devices to improve recognition performance and solve the challenges brought by the complexity and dynamic changes of resource allocation across time periods.

Method used

By obtaining the current and historical observation data of ISAC devices, using cellular traffic load prediction and location prediction, and combining deep reinforcement learning algorithms to optimize resource allocation strategies, the perception and communication resources of each ISAC device are dynamically adjusted to optimize the human motion recognition process.

Benefits of technology

It improves the accuracy and resource utilization efficiency of human motion recognition, optimizes the multi-device collaborative motion recognition method, and improves the recognition accuracy in dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119767262B_ABST
    Figure CN119767262B_ABST
Patent Text Reader

Abstract

The present invention specifically relates to a motion recognition method and system based on multi-device collaboration with integrated sensory and communication capabilities, a storage medium, a computer program product, and an electronic device. This method proposes a deep reinforcement learning online algorithm that dynamically allocates sensory and communication resources to adapt to time-varying sensory environments and network conditions. This algorithm optimizes the allocation of sensory and communication resources across multiple devices over multiple time periods, thereby maximizing the accuracy of human motion recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technology, and in particular to a motion recognition method and system based on sensory integration and multi-device collaboration, a storage medium, a computer program product, and an electronic device. Background Art

[0002] In related technologies, with the rapid development of sensing technology, human motion recognition (HMR) has attracted widespread attention in various application scenarios, such as elderly care, search and rescue operations, and human-computer interaction. Various methods have been proposed for HMR. Among them, wireless recognition offers greater privacy and robustness than visual and acoustic recognition, making it a key research direction in the field. Wireless recognition leverages the impact of human motion on the radio signal propagation path (such as reflection, diffraction, and scattering) by extracting and analyzing features in the received signal to identify specific movements. However, the collection and transmission of data from multiple devices to a fusion center often incurs significant communication overhead. Communication and perception processes compete for limited wireless resources (such as energy) on the devices. To improve resource utilization, the Integrated Sensing and Communication (ISAC) technology has been proposed. ISAC implements perception and communication functions in hardware through shared modules, improving resource utilization. Furthermore, the perception and communication processes are jointly designed from the perspectives of waveform design, beamforming, and resource allocation to balance their conflicting requirements. On this basis, some researchers have also proposed task-oriented ISAC, which maximizes the completion of downstream tasks by jointly designing perception and communication.

[0003] For multi-device collaborative human action recognition tasks, the target's echoes appear randomly in both time and space due to the target's constant movement. Furthermore, the presence of interference signals is also random. Therefore, it is necessary to combine perception data from multiple time periods to achieve temporal diversity gains. To achieve this, resource allocation across multiple devices in multiple time periods must be carefully designed. Maximizing human action recognition accuracy is particularly critical for resource-constrained devices. However, multi-device collaborative human action recognition based on integrated sensory communication faces three major challenges. First, due to the varying positions of different devices relative to the target, their contributions to improving human action recognition performance vary. Therefore, allocating network resources across multiple ISAC devices to maximize human action recognition accuracy is challenging. Second, resource allocation within one time period determines the remaining resources, which in turn affects resource allocation in subsequent time periods. This makes cross-time resource allocation for each device complex and coupled. Finally, the perception environment and network state are dynamic. The random movement of the target and interfering objects causes dynamic changes in echoes in both time and space, resulting in dynamic changes in the perception environment. As for the network status, the number of cellular users and wireless channels are also time-varying, which increases the difficulty of dynamic resource allocation and thus affects the final accuracy of human motion recognition.

[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present invention, and therefore may include information that does not constitute prior art known to ordinary technicians in this field. Summary of the Invention

[0005] The present invention provides a motion recognition method and system based on sensory integration and multi-device collaboration, a storage medium, a computer program product, and an electronic device. By sensing dynamic environmental changes and network status changes, dynamic resource allocation is performed, thereby improving the accuracy of human motion recognition, and thus being able to overcome the defects in the existing technology to a certain extent.

[0006] Other features and advantages of the present invention will become apparent from the following detailed description, or may be learned in part by practice of the present invention.

[0007] According to a first aspect of the present invention, a method for motion recognition based on sensory integration and multi-device collaboration is provided, the method comprising:

[0008] Obtaining current observation data corresponding to the current time slot of the ISAC device and historical observation data corresponding to historical time slots; wherein the observation data includes: power configuration information of the ISAC device, number of cellular users, and motion status information of the target to be analyzed;

[0009] Based on the temporal correlation of cellular traffic loads between consecutive time slots, the current observation data and the historical observation data are used to perform cellular traffic load prediction processing to obtain a predicted number of cellular users corresponding to the next time slot; and based on the predicted number of cellular users, the available communication bandwidth prediction data of all ISAC devices is determined;

[0010] Performing position prediction processing based on the current motion state of the target to be analyzed to obtain the position prediction result corresponding to the next time slot of the target to be analyzed; and determining the posterior distribution of the feature data collected by each ISAC device based on the position prediction result;

[0011] Based on the power amounts corresponding to the current time slot and the historical time slots, estimate the available power amount prediction data for the next time slot;

[0012] Determine the resource allocation strategy corresponding to each ISAC device based on the available communication bandwidth prediction data, the available power prediction data, and the position prediction result corresponding to the target to be analyzed in the next time slot;

[0013] The characteristic data transmitted by each ISAC device based on the resource allocation strategy is received, and a motion recognition result corresponding to the target to be analyzed is determined based on the characteristic data.

[0014] In some exemplary embodiments, the method further comprises:

[0015] A tuple is defined based on the number of ISAC devices, the state space, the set of joint actions, the state transition probability function, the reward function, the set of joint observations, and the observation probability function; the tuple is used to represent a decentralized partially observable Markov decision process for estimating the amount of available power in the next time slot;

[0016] The goal of defining the global reward of the tuple includes maximizing the accumulated global reward by optimizing the resource allocation strategy of each agent within M time slots.

[0017] In some exemplary embodiments, the method further comprises:

[0018] Construct a sample dataset based on the historical resource allocation policy of ISAC devices;

[0019] Build an initial policy optimization model based on the deep deterministic policy gradient algorithm;

[0020] Input the sample data set into the policy optimization initial model, use the actor network of the initial model to output the selected action, and use the critic network of the policy optimization initial model to evaluate the action;

[0021] The parameters of the actor network and critic network are updated through the preset loss function to obtain the trained strategy optimization model.

[0022] In some exemplary embodiments, performing cellular traffic load prediction processing based on the time correlation of cellular traffic loads between consecutive time slots using the current observation data and the historical observation data to obtain a predicted number of cellular users corresponding to the next time slot includes:

[0023] The observation data corresponding to multiple consecutive time slots are input into the cellular traffic load prediction model based on the RNN network to obtain the predicted number of cellular users corresponding to the next time slot output by the model.

[0024] In some exemplary embodiments, the current motion state includes position information, motion direction, and speed information of the target to be analyzed in the current time slot.

[0025] In some exemplary embodiments, the method further comprises:

[0026] The edge server receives the feature set sent by each ISAC device; wherein each ISAC device shares a channel with the current cellular communication user through orthogonal frequency division multiple access;

[0027] Feature data integration processing is performed based on the feature sets corresponding to multiple consecutive time slots to obtain the action recognition results corresponding to the target to be identified.

[0028] According to a second aspect of the present invention, a motion recognition system based on sensory integration and multi-device collaboration is provided, comprising: an edge server and an ISAC device; wherein,

[0029] The edge server is used to obtain current observation data corresponding to the current time slot of the ISAC device and historical observation data corresponding to historical time slots; wherein the observation data includes: power configuration information of the ISAC device, the number of cellular users, and motion state information of the target to be analyzed; based on the time correlation of cellular traffic load between consecutive time slots, the current observation data and historical observation data are used to perform cellular traffic load prediction processing to obtain the predicted number of cellular users corresponding to the next time slot; and based on the predicted number of cellular users, available communication bandwidth prediction data of all ISAC devices is determined; based on the current motion state of the target to be analyzed, position prediction processing is performed to obtain a position prediction result corresponding to the target to be analyzed in the next time slot; and based on the position prediction result, the posterior distribution of feature data collected by each ISAC device is determined; based on the power corresponding to the current time slot and the historical time slots, the available power prediction data for the next time slot is estimated; based on the power corresponding to the current time slot and the historical time slots, the resource allocation strategy corresponding to each ISAC device is determined based on the available communication bandwidth prediction data, the available power prediction data, and the position prediction result corresponding to the target to be analyzed in the next time slot; the feature data transmitted by each ISAC device based on the resource allocation strategy is received, and the action recognition result corresponding to the target to be analyzed is determined based on the feature data;

[0030] The ISAC device is used to send feature data sets to the edge server according to the resource allocation policy.

[0031] According to a third aspect of the present invention, there is provided a storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned motion recognition method based on sensory integration and multi-device collaboration.

[0032] According to a fourth aspect of the present invention, there is provided an electronic device, comprising:

[0033] processor; and

[0034] a memory for storing executable instructions of the processor;

[0035] Among them, the processor is configured to implement the above-mentioned action recognition method based on integrated sensory communication and multi-device collaboration by executing the executable instructions.

[0036] According to a fifth aspect of the present invention, there is provided a computer program product having a computer program stored thereon, which implements the above-mentioned action recognition method based on sensory integration and multi-device collaboration when the computer program is executed by a processor.

[0037] The embodiment of the present invention provides a method for motion recognition based on integrated sensory communication and multi-device collaboration. The method obtains current observation data corresponding to the current time slot of the ISAC device and historical observation data corresponding to the historical time slot, and uses the current observation data and historical observation data to perform cellular traffic load prediction processing to obtain the predicted number of cellular users corresponding to the next time slot; and determines the available communication bandwidth prediction data of all ISAC devices based on the predicted number of cellular users; performs position prediction processing based on the current motion state of the target to be analyzed to obtain the position prediction result corresponding to the target to be analyzed in the next time slot; and determines the posterior distribution of the feature data collected by each ISAC device based on the position prediction result; estimates the available power prediction data for the next time slot based on the power corresponding to the current time slot and the historical time slot; determines the resource allocation strategy corresponding to each ISAC device based on the available communication bandwidth prediction data, the available power prediction data, and the position prediction result corresponding to the target to be analyzed in the next time slot; receives the feature data transmitted by each ISAC device based on the resource allocation strategy, and determines the motion recognition result corresponding to the target to be analyzed based on the feature data. By optimizing the perception and communication resource allocation of multiple devices in multiple time periods, the accuracy of human motion recognition is maximized.

[0038] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The accompanying drawings are incorporated into and constitute a part of this specification, illustrate embodiments consistent with the present invention, and together with the description, serve to explain the principles of the present invention. Obviously, the drawings described below are only some embodiments of the present invention, and it is clear that those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0040] Figure 1 A schematic diagram schematically illustrates an exemplary embodiment of the present invention, a method for motion recognition based on sensory integration and multi-device collaboration;

[0041] Figure 2 A schematic diagram schematically illustrates a framework of a method for human motion recognition based on sensory integration and multi-device collaboration according to an exemplary embodiment of the present invention;

[0042] Figure 3 A schematic diagram schematically illustrates a variation trend of human motion recognition accuracy along with time domain resource budget according to an exemplary embodiment of the present invention;

[0043] Figure 4 A schematic diagram schematically illustrates a changing trend of human motion recognition accuracy along with spatial resource budget according to an exemplary embodiment of the present invention;

[0044] Figure 5 A schematic diagram schematically illustrates the accuracy of human motion recognition under different bandwidth budgets according to an exemplary embodiment of the present invention;

[0045] Figure 6 A schematic diagram schematically illustrates the accuracy of human motion recognition under different energy budgets according to an exemplary embodiment of the present invention;

[0046] Figure 7 A schematic diagram schematically illustrates an exemplary embodiment of the present invention, a motion recognition system based on sensory integration and multi-device collaboration;

[0047] Figure 8 The figure schematically shows the composition of an electronic device in an exemplary embodiment of the present invention. DETAILED DESCRIPTION

[0048] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0049] In addition, the accompanying drawings are merely schematic illustrations of the present invention and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the blocks shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0050] In related technologies, with the rapid development of sensing technology, human motion recognition (HMR) has attracted widespread attention in various application scenarios, such as elderly care, search and rescue operations, and human-computer interaction. Various methods have been proposed for HMR. Among them, wireless recognition offers improved privacy and robustness compared to visual and acoustic recognition, making it a key research area in the field. Wireless recognition leverages the impact of human motion on the radio signal propagation path (such as reflection, diffraction, and scattering) to extract and analyze features in the received signal to identify specific movements. In contrast, collecting data from multiple devices and transmitting it to a fusion center typically incurs significant communication overhead. Communication and perception processes compete for limited wireless resources (such as energy) on the devices, making improving resource utilization efficient a key issue. To address this issue, integrated sensing and communication (ISAC) technology has emerged. ISAC integrates perception and communication functions through shared hardware modules, thereby improving resource utilization. Furthermore, ISAC's joint design of waveform design, beamforming, and resource allocation effectively balances the demands of perception and communication. Building on this foundation, some research has proposed task-oriented ISACs, which maximize the completion of downstream tasks by jointly designing perception and communication. For multi-device collaborative human action recognition tasks, due to the constant movement of the target, echo signals appear randomly in both time and space. Furthermore, the presence of interference signals is also random. Therefore, to improve action recognition accuracy, it is necessary to combine perception data from multiple time periods to achieve temporal diversity gains. To achieve this goal, resource allocation across multiple devices must be carefully designed across multiple time periods. Maximizing action recognition accuracy is particularly critical for devices with limited resources. In practical applications, multi-device collaborative action recognition methods based on integrated perception and communication face three major challenges. First, different devices have different positions relative to the target, resulting in varying contributions to human action recognition performance. Therefore, how to rationally allocate resources to improve recognition accuracy is a complex problem. Second, resource allocation within a time period determines resource availability, which in turn affects resource allocation in subsequent time periods, resulting in cross-time period resource allocation coupling. The key challenge is how to allocate resources to the time period that ultimately achieves the highest human action recognition accuracy. Finally, since the perception environment and network status are dynamically changing, the random movement of targets and interfering objects, as well as the fluctuations in the number of cellular users and wireless channels, further affect the accuracy of human action recognition.

[0051] To address the shortcomings and deficiencies of existing technologies, this example embodiment provides a multi-device collaborative motion recognition method based on integrated sensory and communication. This method optimizes the human motion recognition process by rationally allocating the sensory and communication resources across multiple devices. By jointly designing resource allocation strategies for ISAC devices, devices can more efficiently sense and transmit data within each time slot, ultimately outputting accurate motion recognition results.

[0052] In this example implementation, reference Figure 1 As shown, the multi-device collaborative action recognition method based on sensory integration may specifically include the following steps:

[0053] Step S11, obtaining current observation data corresponding to the current time slot of the ISAC device and historical observation data corresponding to historical time slots; wherein the observation data includes: power configuration information of the ISAC device, the number of cellular users, and motion status information of the target to be analyzed;

[0054] Step S12, based on the time correlation of cellular traffic loads between consecutive time slots, using the current observation data and the historical observation data, performs cellular traffic load prediction processing to obtain a predicted number of cellular users corresponding to the next time slot; and determining available communication bandwidth prediction data for all ISAC devices based on the predicted number of cellular users;

[0055] Step S13: performing position prediction processing based on the current motion state of the target to be analyzed to obtain a position prediction result corresponding to the next time slot of the target to be analyzed; and determining the posterior distribution of the feature data collected by each ISAC device based on the position prediction result;

[0056] Step S14, estimating the available power prediction data for the next time slot based on the power corresponding to the current time slot and the historical time slots;

[0057] Step S15, determining a resource allocation strategy corresponding to each ISAC device based on the available communication bandwidth prediction data, the available power prediction data, and the position prediction result corresponding to the target to be analyzed in the next time slot;

[0058] Step S16: receiving the characteristic data transmitted by each ISAC device based on the resource allocation strategy, and determining the action recognition result corresponding to the target to be analyzed based on the characteristic data.

[0059] Below, each step of the multi-device collaborative action recognition method based on sensory integration in this example implementation will be described in more detail with reference to the accompanying drawings and embodiments.

[0060] In step S11, current observation data corresponding to the current time slot of the ISAC device and historical observation data corresponding to historical time slots are obtained; wherein the observation data includes: power configuration information of the ISAC device, the number of cellular users, and motion status information of the target to be analyzed.

[0061] Exemplarily, the edge server may allocate resources to the ISAC device. For the edge server, when the current time slot is time slot m, in order to allocate resources to each ISAC device in the next time slot, that is, the resource configuration of time slot m+1, the current observation data of the current time slot and the historical observation data of multiple consecutive historical time slots can be used to predict the resource allocation of the next time slot. The observation data collected for each time slot may include: power configuration information of the ISAC device, the number of cellular users, and the motion state information of the target to be analyzed. Specifically, the power configuration information of the ISAC device may be the available power of the ISAC device, which can affect the effect of the ISAC device in the perception stage and the communication stage. The number of cellular users may be the number of cellular users in the time slot, which can be used to reflect the network load. The target to be analyzed may be an object that requires motion recognition, and the motion state information may be the location information of the target; in addition, it may also include motion direction and speed information.

[0062] The observation data of ISAC device k in time slot m can be expressed as:

[0063]

[0064] Among them, m Represents the position of the target to be analyzed; l m Expressed as the number of cellular users; It is represented as the power configuration of ISAC device k.

[0065] For example, considering data inference accuracy and data processing performance, the number of historical time slots can be configured to be 1-5. One time slot can be configured to be 0.7 seconds.

[0066] In step S12, based on the time correlation of the cellular traffic load between consecutive time slots, the current observation data and the historical observation data are used to perform cellular traffic load prediction processing to obtain the predicted number of cellular users corresponding to the next time slot; and based on the predicted number of cellular users, the available communication bandwidth prediction data of all ISAC devices is determined.

[0067] For example, a deep reinforcement learning-based online optimization algorithm (DRLO) can be provided, which can maximize the accuracy of reasoning in a dynamic environment. The algorithm can predict the environmental state and adaptively adjust to cope with changes in the new environmental state to allocate resources of multiple ISAC devices.

[0068] To implement the DRLO algorithm, it is necessary to predict the sensory environment and network status information for the next time slot. To this end, a prediction model based on a recurrent neural network (RNN) is provided, which includes a cellular traffic load prediction model and a target location prediction model. The cellular traffic load prediction model outputs the predicted number of cellular users in the next time slot, thereby determining the available communication bandwidth for all ISAC devices. The target location prediction model outputs the target location for the next time slot, thereby determining the posterior distribution of the features obtained by each ISAC device.

[0069] Exemplarily, the performing cellular traffic load prediction processing based on the time correlation of cellular traffic loads between consecutive time slots using the current observation data and the historical observation data to obtain the predicted number of cellular users corresponding to the next time slot includes:

[0070] The observation data corresponding to multiple consecutive time slots are input into the cellular traffic load prediction model based on the RNN network to obtain the predicted number of cellular users corresponding to the next time slot output by the model.

[0071] Specifically, in the cellular traffic load prediction model, the RNN network is used to predict the number of cellular users in the next time slot. o Taking the observation data of time slots as input, the RNN network can capture the temporal correlation of cellular traffic loads between consecutive time slots and thus provide predictions for the next time slot. Figure 2 As shown, from the first (mM o ) time slot to the (m-1)th time slot, the observation data of each time slot is sequentially input into an RNN network, and the output of the last RNN network is input into the output layer of the entire module.

[0072] Specifically, the RNN network architecture can include an input layer, an RNN layer, a fully connected layer, and an output layer. The process of RNN network processing M0 group observation data usually includes the following steps:

[0073] 1) Data Input: Input M0 sets of observation data into the RNN network. Each set of data typically contains time series information such as target location, number of cellular users, and device power.

[0074] The input data can be expressed as [M0, T, N]; where M0 is the number of observation data groups, for example, 10-100 groups; T is the time slot length, for example, 10 time slots; and N is the number of features contained in each time slot, such as the number of cellular users, device power, etc.

[0075] 2) Time-step propagation: At each time step, the RNN network processes the current observation data together with the hidden state (or memory) of the previous time step.

[0076] The RNN network generates the current output based on the current input and previous memory (through recurrent connections) and updates the internal state of the network.

[0077] The RNN network gradually reads the input data of each time slot and passes the information to the next time step. Each time step updates the current state based on the current input and the previous hidden state (or memory). The specific update formula can include:

[0078]

[0079] Among them, x t is the current input; W and U are weight matrices; b is the bias; and f is the activation function, such as the ReLU activation function.

[0080] 3) Inter-layer transfer: If there are multiple RNN layers, the network will transfer information between layers and continuously extract features in the time series.

[0081] 4) Fully connected layer output: After being processed by the RNN layer, the data will be passed to the fully connected layer for final prediction processing.

[0082] Generate prediction results: Through the fully connected layer, the final output is the predicted value of the next time slot, that is, the number of cellular users. The formula may include:

[0083]

[0084] in, is the predicted number of cellular users, W out 、b out are the weights and biases of the output layer.

[0085] Finally, the RNN network outputs the predicted value of the number of cellular users in the next time slot

[0086] In step S13, position prediction processing is performed based on the current motion state of the target to be analyzed to obtain the position prediction result corresponding to the target to be analyzed in the next time slot; and the posterior distribution of the feature data collected by each ISAC device is determined according to the position prediction result.

[0087] For example, assuming there are no sudden turns in the target's trajectory, the target's next position can be estimated based on its current position and velocity. Although this estimation method may introduce some slight bias, due to the short duration of each time slot, the impact of this bias on the final action recognition accuracy is negligible.

[0088] Specifically, let (x m ,y m ) is the position of the target at time slot m, θ m is the current moving direction of the target, v m is the target’s speed. Then at time slot m+1, the target’s position (x m+1 ,y m+1 ) can be estimated by the following formula:

[0089]

[0090] Here, Δm is the duration of one time slot.

[0091] In step S14, based on the power amounts corresponding to the current time slot and the historical time slots, available power amount prediction data for the next time slot is estimated.

[0092] For example, the amount of power used in the current and historical time slots determines the amount of power available in subsequent time slots. Therefore, the problem can be converted into a decentralized partially observable Markov decision process.

[0093] Exemplarily, the method further includes: defining a tuple based on the number of ISAC devices, the state space, the set of joint actions, the state transition probability function, the reward function, the set of joint observations, and the observation probability function; the tuple is used to represent a decentralized partially observable Markov decision process for estimating the amount of available power in the next time slot;

[0094] The goal of defining the global reward of the tuple includes maximizing the accumulated global reward by optimizing the resource allocation strategy of each agent within M time slots.

[0095] Specifically, a tuple<K,S,A,P,R,O,Z,ψ> To represent the decentralized partially observable Markov decision process, where K is the number of agents (ISAC devices), S is the state space, A = × k∈K A k is a set of joint actions, P is the state transition probability function, R is the reward function, O = × k∈K O k is the set of joint observations, Z is the observation probability function, and ψ is the discount factor. In time slot m, based on the observation data Agent k follows the strategy Determine its action. The policy is a set of actions When all agents perform actions according to their respective strategies, a global reward R is generated. m =R(s m , a m ), and the state changes to s m+1 The goal is to maximize the cumulative global reward by optimizing the policy of each agent within M time slots.

[0096] Specifically, the state space can be defined according to the problem, including the location of the target, the number of cellular users, and the amount of power available to each ISAC device. The state changes with time slots. The observation data at time slot m is expressed as Among them m 、l m and They represent the location of the target, the number of cellular users, and the available power of ISAC device k at time slot m, respectively.

[0097] The joint action includes the resources that need to be allocated, namely, quantization gain, sensing power, transmission power and communication bandwidth. It is necessary to determine the action for each ISAC device in each time slot, that is, to allocate resources. The action of ISAC device k in time slot m is expressed as

[0098] For the above rewards, it can be expressed as For a task-oriented multi-device ISAC system, its goal is to maximize the action recognition performance, which is equivalent to maximizing the discriminant gain and is represented by the objective function. Therefore, the discriminant gain is used in the reward function. The cumulative global reward function at time slot m is:

[0099]

[0100] Where Gd,d′(·) represents the discrimination gain between category d and category d′. m Next, when the time slot is m, take action When , the nth recovery feature is obtained

[0101] Furthermore, the discount factor determines the importance of future rewards in the cumulative reward. This is a parameter used in reinforcement learning to discount future rewards, typically 0 ≤ ψ ≤ 1. It reflects the importance the agent places on future rewards. A larger discount factor (closer to 1) means the agent prioritizes long-term rewards, while a smaller discount factor means the agent prioritizes short-term rewards.

[0102] The state transition function P is used to describe the probability of transitioning from one state to another. In this problem, P represents the probability of transitioning from one state to another given the current state s. m and the current action (i.e., the combined action set of all devices), how does the system transition to the next state s? m+1 .

[0103] Specifically, the ISAC device adjusts the device resource configuration (such as power allocation, bandwidth allocation, etc.) through actions, thereby affecting the system state. m+1 |s m , a m ) gives the current state s m Next, take action Then transfer to s m +1 probability.

[0104] The observation probability function Z describes the observation probability given a hidden state s m and actions In the case of a certain observation data that the ISAC device can observe Each ISAC device has the possibility of obtaining a partial observation based on its current state and action in each time slot.

[0105] Observation probability function Z(s m+1 |s m , a m ) describes the observation data of ISAC device k in time slot m In the current state m and actions The probability of the next.

[0106] In step S15, the resource allocation strategy corresponding to each ISAC device is determined according to the available communication bandwidth prediction data, the available power prediction data, and the position prediction result corresponding to the target to be analyzed in the next time slot.

[0107] Exemplarily, the method further includes:

[0108] Step S21, constructing a sample data set based on the historical resource allocation policy of the ISAC device;

[0109] Step S22, constructing an initial policy optimization model based on a deep deterministic policy gradient algorithm;

[0110] Step S23: input the sample data set into the strategy optimization initial model, use the actor network of the initial model to output the selected action, and use the critic network of the strategy optimization initial model to evaluate the action;

[0111] In step S24, the parameters of the actor network and the critic network are updated using a preset loss function to obtain a trained policy optimization model.

[0112] Specifically, in order to maximize the cumulative global reward function, we use a deep reinforcement learning algorithm to optimize the strategy of each agent. At the same time, we use random mini-batch training, that is, using a small number of training samples in each iteration to reduce the amount of computation and speed up training. i Represents the index of each sample in the mini-batch. In particular, we use the Deep Deterministic Policy Gradient (DDPG) algorithm to obtain the optimal policy for each agent.

[0113] Using Observation Sets As input, with parameters θ actor The actor network outputs the selected action Based on the observation set and the selected action, with parameters θ critic The critic network evaluates these actions.

[0114] During the training process, the parameters of the critic network are adjusted by minimizing the loss function Update. The gradient of the loss function is expressed as:

[0115]

[0116] in, and are the weights of the critic network and the actor network, respectively. Similar to the Deep Q Network (DQN), these two target networks are updated by partially copying weights from the main network.

[0117] To train the parameters of the actor network, the loss function The gradient of is expressed as:

[0118]

[0119] In step S16, feature data transmitted by each ISAC device based on the resource allocation policy is received, and a motion recognition result corresponding to the target to be analyzed is determined based on the feature data.

[0120] For example, after determining the resource allocation policy for each ISAC device in the next time slot based on observation data corresponding to the current time slot and multiple historical time slots, the edge server can distribute the updated resource allocation policy to the corresponding ISAC device. This allows each ISAC device to determine specific actions for the next time slot based on the resource allocation policy, such as sensing power, transmission power, and communication bandwidth. Feature data corresponding to the collected sensing data is then sent to the edge server. The ISAC device can compress the sensing data and send the compressed feature data to the edge server.

[0121] In some exemplary embodiments, the method further includes: an edge server receiving a quantitative feature set sent by each ISAC device; wherein each ISAC device shares a channel with a current cellular communication user via orthogonal frequency division multiple access; and performing feature data integration processing based on feature sets corresponding to multiple consecutive time slots to obtain an action recognition result corresponding to the target to be identified.

[0122] For example, the edge server can integrate feature data from multiple consecutive time slots, use the integrated feature data to perform action recognition, and output the corresponding action recognition results. For example, the integrated feature data can be input into a trained action recognition model to obtain the action recognition results output by the model.

[0123] This paper addresses the problem of human action recognition and maximizes its accuracy by optimizing the allocation of perception and communication resources across multiple devices over multiple time periods. Specifically, the accuracy of human action recognition is measured using discriminant gain, where a greater discriminant gain indicates higher accuracy.

[0124] For example, consider the human action recognition scenario. Five ISAC devices are configured to extract the features of the target action from the received echo, and then transmit the features to the edge server for action recognition. The action candidates considered include children walking, children running, adults walking, and adults running. In the simulation, we use a wireless sensing simulator to simulate high-fidelity target actions. The target moves along a certain trajectory at a specific speed to generate action data. Assume that the height H of children and adults is uniformly distributed in the range of [0.9m, 1.2m] and [1.6m, 1.9m] respectively. The speeds of standing, walking, and running are set to 0m / s, 0.25H m / s, and 0.5H m / s, respectively. The target's movement direction is uniformly distributed in the range of [-180°, +180°]. ISAC devices are distributed in a circular area with a radius of 50 meters. The channel from the device to the edge server follows an independent and identically distributed (iid) Rayleigh distribution with zero mean and variance. Configurable The power of the additive white Gaussian noise (AWGN) is N0 = -120 dB.

[0125] In order to represent the performance gain of this scheme, the following three baseline algorithms are considered for comparison, as follows: 1) Sensing Mode Allocation (SMA): Only the parameters related to the sensing process, namely, sensing power and quantization gain, are optimized. Communication resources (including the transmission power and communication bandwidth of the ISAC device in each time slot) are evenly distributed. 2) Communication Mode Allocation (CMA): Only communication resources are optimized, including the transmission power and communication bandwidth of the ISAC device in each time slot. The sensing power is evenly distributed in all time slots, and the quantization gain is set to a specific value. 3) Single Slot Sensing Allocation (SSA): The energy budget is first evenly distributed to each time slot. Then, resources are allocated based on the amount of allocated energy and available bandwidth, thereby maximizing the performance of each time slot. The comparison results are shown in Figure 2. Figure 3-Figure 6 Extensive numerical results show that the proposed DRLO algorithm outperforms other comparison algorithms in overall performance of human action recognition.

[0126] The method provided by the embodiment of the present invention provides an online optimization algorithm based on deep reinforcement learning to maximize the accuracy of reasoning in a dynamic environment. The algorithm can predict the environmental state and adaptively adjust to cope with the changes in the new environmental state to allocate the resources of multiple ISAC devices. At the same time, the DRLO algorithm is applied to a multi-device collaborative human motion recognition system. The system has a dynamic interference and time-varying network state and is composed of multiple devices and edge servers with limited energy and bandwidth. These devices obtain perception data from the received signals in a continuous time period based on the resources allocated by the DRLO algorithm and transmit it to the edge server to complete the motion recognition task of the mobile target.

[0127] The motion recognition method based on integrated sensory communication and multi-device collaboration provided by the present invention dynamically allocates communication resources to multiple ISAC devices at different perspectives according to changes in the dynamic perception environment and changes in network status, thereby achieving reasonable allocation of resources. This enables the edge server to obtain the perception data of each ISAC device in a timely and accurate manner, thereby improving the accuracy of human motion recognition.

[0128] It should be noted that the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention and are not intended to be limiting. It is readily understood that the processes illustrated in the above figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0129] For further reference, Figure 7As shown, the embodiment of this example also provides a motion recognition system 70 based on sensory integration and multi-device collaboration, including: an edge server 702 and an ISAC device 701.

[0130] The edge server 702 is configured to obtain current observation data corresponding to the current time slot of the ISAC device and historical observation data corresponding to historical time slots; the observation data includes: power configuration information of the ISAC device, the number of cellular users, and motion state information of the target to be analyzed; based on the time correlation of cellular traffic load between consecutive time slots, the current observation data and the historical observation data are used to perform cellular traffic load prediction processing to obtain the predicted number of cellular users corresponding to the next time slot; and based on the predicted number of cellular users, available communication bandwidth prediction data for all ISAC devices is determined; based on the current motion state of the target to be analyzed, a position prediction processing is performed to obtain a position prediction result corresponding to the target to be analyzed in the next time slot; and based on the position prediction result, a posterior distribution of feature data collected by each ISAC device is determined; based on the power corresponding to the current time slot and the historical time slots, available power prediction data for the next time slot is estimated; based on the power corresponding to the current time slot and the historical time slots, a resource allocation strategy corresponding to each ISAC device is determined based on the available communication bandwidth prediction data, the available power prediction data, and the position prediction result corresponding to the target to be analyzed in the next time slot; a feature data set transmitted by each ISAC device based on the resource allocation strategy is received, and a motion recognition result corresponding to the target to be analyzed is determined based on the feature data set.

[0131] The ISAC device 701 is configured to send a feature data set to an edge server according to a resource allocation policy.

[0132] Illustratively, the system may include multiple ISAC devices 701 .

[0133] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to an embodiment of the present invention, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.

[0134] Figure 8 A schematic diagram of an electronic device suitable for implementing an embodiment of the present invention is shown.

[0135] It should be noted that Figure 8 The electronic device 1000 shown is only an example and should not limit the functions and scope of use of the embodiments of the present invention.

[0136] like Figure 8As shown, electronic device 1000 includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to the program stored in read-only memory (ROM) 1002 or the program loaded from storage portion 1008 into random access memory (RAM) 1003. Various programs and data required for system operation are also stored in RAM 1003. CPU 1001, ROM 1002 and RAM 1003 are connected to each other via bus 1004. Input / output (I / O) interface 1005 is also connected to bus 1004.

[0137] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, and the like; an output section 1007 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 1008 including a hard disk and the like; and a communication section 1009 including a network interface card such as a LAN (Local Area Network) card or a modem. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. Removable media 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 1010 as needed, so that computer programs read therefrom can be installed into the storage section 1008 as needed.

[0138] In particular, according to an embodiment of the present invention, the process described below with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product that includes a computer program carried on a storage medium, the computer program containing program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1009 and / or installed from a removable medium 1011. When the computer program is executed by the central processing unit (CPU) 1001, the various functions defined in the system of the present application are performed.

[0139] It should be noted that the storage medium shown in the embodiments of the present invention can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device or device. In the present invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any storage medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Program code contained on the storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0140] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0141] The units involved in the embodiments of the present invention may be implemented in software or hardware, and the units described may also be provided in a processor. In some cases, the names of these units do not limit the units themselves.

[0142] It should be noted that, as another aspect, the present application also provides a storage medium, which can be included in an electronic device; or it can exist independently without being installed in the electronic device. The above storage medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device implements the method described in the following embodiments. For example, the electronic device can implement the following Figure 1 The individual steps of the method are shown.

[0143] In one embodiment, the present application provides a computer program product, including a computer program, which implements the steps in the above-mentioned method embodiments when executed by a processor.

[0144] Furthermore, the above-described figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention and are not intended to be limiting. It is readily understood that the processes illustrated in the above-described figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0145] Other embodiments of the present invention will readily occur to those skilled in the art after considering the specification and practicing the invention herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the invention being indicated by the claims.

[0146] It should be understood that the present invention is not limited to the exact construction described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof, which is limited only by the appended claims.

Claims

1. A motion recognition method based on sensory integration and multi-device collaboration, characterized in that: The method comprises: Obtaining current observation data corresponding to the current time slot of the ISAC device and historical observation data corresponding to historical time slots; wherein the observation data includes: power configuration information of the ISAC device, number of cellular users, and motion status information of the target to be analyzed; Based on the temporal correlation of cellular traffic loads between consecutive time slots, the current observation data and the historical observation data are used to perform cellular traffic load prediction processing to obtain a predicted number of cellular users corresponding to the next time slot; and based on the predicted number of cellular users, the available communication bandwidth prediction data of all ISAC devices is determined; Performing position prediction processing based on the current motion state of the target to be analyzed to obtain the position prediction result corresponding to the next time slot of the target to be analyzed; and determining the posterior distribution of the feature data collected by each ISAC device based on the position prediction result; Based on the power amounts corresponding to the current time slot and the historical time slots, estimate the available power amount prediction data for the next time slot; Determine the resource allocation strategy corresponding to each ISAC device based on the available communication bandwidth prediction data, the available power prediction data, and the position prediction result corresponding to the target to be analyzed in the next time slot; The characteristic data transmitted by each ISAC device based on the resource allocation strategy is received, and a motion recognition result corresponding to the target to be analyzed is determined based on the characteristic data.

2. The method according to claim 1, characterized in that The method further comprises: A tuple is defined based on the number of ISAC devices, the state space, the set of joint actions, the state transition probability function, the reward function, the set of joint observations, and the observation probability function; the tuple is used to represent a decentralized partially observable Markov decision process for estimating the amount of available power in the next time slot; The goal of defining the global reward of the tuple includes maximizing the accumulated global reward by optimizing the resource allocation strategy of each agent within M time slots.

3. The method according to claim 2, characterized in that The method further comprises: Construct a sample dataset based on the historical resource allocation policy of ISAC devices; Build an initial policy optimization model based on the deep deterministic policy gradient algorithm; Input the sample data set into the policy optimization initial model, use the actor network of the initial model to output the selected action, and use the critic network of the policy optimization initial model to evaluate the action; The parameters of the actor network and critic network are updated through the preset loss function to obtain the trained strategy optimization model.

4. The method according to claim 1, wherein The method of performing cellular traffic load prediction processing based on the time correlation of cellular traffic loads between consecutive time slots and utilizing the current observation data and the historical observation data to obtain a predicted number of cellular users corresponding to the next time slot includes: The observation data corresponding to multiple consecutive time slots are input into the cellular traffic load prediction model based on the RNN network to obtain the predicted number of cellular users corresponding to the next time slot output by the model.

5. The method according to claim 1, wherein The current motion state includes the position information, motion direction, and speed information of the target to be analyzed in the current time slot.

6. The method according to claim 1, wherein The method further comprises: The edge server receives the quantized feature set sent by each ISAC device; wherein each ISAC device shares a channel with the current cellular communication user through orthogonal frequency division multiple access; Feature data integration processing is performed based on the feature sets corresponding to multiple consecutive time slots to obtain the action recognition results corresponding to the target to be identified.

7. A motion recognition system based on sensory integration and multi-device collaboration, characterized in that: The system includes: an edge server and an ISAC device; wherein, The edge server is used to obtain current observation data corresponding to the current time slot of the ISAC device and historical observation data corresponding to historical time slots; wherein the observation data includes: power configuration information of the ISAC device, the number of cellular users, and motion state information of the target to be analyzed; based on the time correlation of cellular traffic load between consecutive time slots, the current observation data and historical observation data are used to perform cellular traffic load prediction processing to obtain the predicted number of cellular users corresponding to the next time slot; and based on the predicted number of cellular users, the available communication bandwidth prediction data of all ISAC devices is determined; based on the current motion state of the target to be analyzed, a position prediction processing is performed to obtain a position prediction result corresponding to the target to be analyzed in the next time slot; and based on the position prediction result, the posterior distribution of the feature data collected by each ISAC device is determined; based on the power corresponding to the current time slot and the historical time slots, the available power amount prediction data for the next time slot is estimated; based on the power amount corresponding to the current time slot and the historical time slots, the resource allocation strategy corresponding to each ISAC device is determined based on the available communication bandwidth prediction data, the available power amount prediction data, and the position prediction result corresponding to the target to be analyzed in the next time slot; the feature data set transmitted by each ISAC device based on the resource allocation strategy is received, and the action recognition result corresponding to the target to be analyzed is determined based on the feature data set; The ISAC device is used to send feature data sets to the edge server according to the resource allocation policy.

8. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the motion recognition method based on sensory integration and multi-device collaboration as described in any one of claims 1 to 6 is implemented.

9. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the motion recognition method based on sensory integration and multi-device collaboration according to any one of claims 1 to 6 is implemented.

10. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; The processor is configured to execute the action recognition method based on sensory integration and multi-device collaboration as described in any one of claims 1 to 6 by executing the executable instructions.

Citation Information

Patent Citations

  • Resource allocation method of multi-cell communication perception integrated system based on DRL

    CN116546506A

  • Intelligent distribution method and system for power resources of communication and inductance integrated internet of things

    CN117769010A