BFI-based dynamic adversarial reinforcement learning cross-domain behavior recognition system and method
By using BFI data and dynamic adversarial reinforcement learning technology in the human behavior recognition system, a cross-domain behavior recognition model that can adapt to different environments is built, solving the problem of degradation of recognition performance in the cross-domain environment by the existing technology, and achieving efficient and accurate behavior recognition.
Patent Information
- Application Number
- CN202510316903.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-27
AI Technical Summary
Existing human behavior recognition technology based on WiFi signals has the problem of degradation of recognition performance in cross-domain environments, especially when the environment changes, it is difficult to effectively identify human behavior.
A dynamic adversarial reinforcement learning cross-domain behavior recognition system based on BFI is adopted. This system includes WiFi signal transceiver and receive devices, data acquisition devices and identification and processing terminals. Through technical means such as BFI data extraction, preprocessing, domain adversarial training and reinforcement learning optimization, a cross-domain behavior recognition model that can adapt to different environments is built.
It realizes efficient and accurate identification of human behavior in different recognition environments, reduces deployment costs and complexity, improves human-computer interactive experience, and is suitable for large-scale promotion and application.
Smart Images

Figure CN120217099A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent recognition, and particularly relates to a cross-domain behavior recognition system and method based on BFI dynamic adversarial reinforcement learning. Background Art
[0002] With the rapid development of information technology, emerging technologies such as the Internet of Things have deeply penetrated into people's daily lives and provided users with more opportunities to interact with devices. The human-computer interaction technology has emerged, aiming to study and design the interaction methods between humans and computer systems to improve the user experience and the usability of the system. Traditional human-computer interaction methods mainly rely on input devices such as keyboards, mice, and touchscreens. However, with the progress of technology, voice, as a semantic-rich interaction method, has gradually become the mainstream. People can execute diverse control commands on various intelligent devices through voice. However, voice interaction requires vocalization, which limits its applicability in all scenarios. In these scenarios, people's demand for a more natural and intuitive interaction method has gradually increased. Against this background, human behavior recognition has become an important part of human-computer interaction. By recognizing the rich semantic information in human actions and directly using it as a control instruction to make intelligent devices perform corresponding actions, the convenience of human-computer interaction has been greatly improved. At present, this recognition and control method has been widely applied in fields such as smart home, healthcare, and virtual reality.
[0003] Human behavior recognition solutions are classified into two categories: contact-based and non-contact-based according to whether specific devices need to be worn. Contact-based recognition relies on the sensing devices worn by users to achieve behavior recognition by monitoring physiological signals and environmental changes. For example, accelerometers, inertial measurement units, gyroscopes, magnetometers, and barometers, etc. However, in actual applications, the worn sensing devices may bring additional burdens and there are also problems of inconvenient wearing in some scenarios. Therefore, in recent years, research has tended to focus on non-contact recognition solutions, which can not only reduce users' dependence on devices but also improve the applicability in different scenarios.
[0004] Non-contact recognition solutions mainly include the following several types: (1) vision-based methods; (2) radio frequency-based methods. Vision-based methods capture human behavior images through installed cameras and then use computer vision and machine learning technologies to recognize behaviors. This method depends on the captured field of view and lighting conditions, and obstacles or dim environments will affect the recognition accuracy and there is a risk of privacy leakage. Radio frequency-based methods utilize the influence of human body and environmental activities on wireless signals to sense, recognize, and detect human behaviors by analyzing the received signals. Due to its many advantages such as non-contact, non-line-of-sight, and privacy protection, the radio frequency method has more advantages in specific scenarios.
[0005] With the popularization of WiFi technology, human behavior recognition technology based on WiFi signals has received increasing attention. Most researchers use channel state information (CSI) for behavior perception, but the required information can only be extracted from a few commercial WiFi network cards through cracking. In addition, current WiFi-based human behavior recognition research usually requires the assumption that the training data and test data have similar distributions. However, WiFi signals are extremely sensitive to environmental changes due to factors such as reflection, diffraction, or scattering during propagation, making the collected data contain not only human behavior information but also coupled environmental information. The same behavior will cause different channel changes when occurring at different locations or by different people, that is, the distributions of the source domain dataset and the target domain dataset will inevitably show differences, namely the so-called domain shift, which will directly affect the performance of the recognition model. In practical applications, it is not realistic to re-collect labeled data and train the model when facing a new environment. To effectively solve the above technical problems, there is an urgent need to provide a new type of cross-domain behavior recognition system and method. Summary of the Invention
[0006] Aiming at the problems existing in the above-mentioned prior art, the present invention provides a cross-domain behavior recognition system and method based on dynamic adversarial reinforcement learning of BFI. The system has a simple structure, low deployment cost, wide application range, and high intelligence level. It can conveniently, accurately, and efficiently recognize human behaviors in different recognition environments; the method has simple implementation steps, low implementation cost, high detection efficiency, high intelligence level, and accurate detection results. It can accurately and efficiently recognize human behaviors in different target fields by using easily obtainable BFI data, which is beneficial to improving the human-computer interaction experience and is suitable for large-scale popularization and application.
[0007] To achieve the above object, the present invention provides a cross-domain behavior recognition system based on dynamic adversarial reinforcement learning of BFI, including a WiFi signal transceiver device, a data acquisition device, and an identification and processing terminal;
[0008] The WiFi signal transceiver device is arranged inside the activity space of the object to be recognized. The WiFi signal transceiver device includes a router and a receiving terminal. The router is used to emit WiFi signals, and the receiving terminal is used to receive WiFi signals and generate a feedback signal with BFI data as identification data;
[0009] The data acquisition device is internally provided with a data packet capture module, and the data packet capture module is used to collect WiFi data packets with identification data from the WiFi signals generated by the WiFi signal transceiver device;
[0010] The recognition processing terminal is connected to the data acquisition device in a wired or wireless manner. The recognition processing terminal includes a BFI data extraction module, a data preprocessing module, a domain adversarial training module, a reinforcement learning optimization module, and a behavior recognition module. The BFI data extraction module is used to extract BFI data from the collected WiFi data packets. The data preprocessing module is used to preprocess the BFI data. The behavior recognition module is used to extract BFI features from the BFI data by using a built-in feature extraction network. At the same time, a built-in classifier is used to perform behavior recognition based on the BFI features and output a behavior recognition result. The domain adversarial training module is used to adversarially train the feature extraction network so that the feature extraction network extracts domain-invariant features. The reinforcement learning optimization module is used to optimize the parameters of the feature extraction network so that the feature extraction network dynamically adapts to domain differences and optimizes the target domain performance.
[0011] As a preference, the data acquisition device is arranged inside the space of the object to be recognized.
[0012] As a preference, the data packet capture module is a network packet capture tool.
[0013] Furthermore, to ensure a powerful processing capacity, the recognition processing terminal is a remote server.
[0014] Furthermore, to ensure accurate and reliable acquisition of BFI data, the WiFi signal transceiver device supports the IEEE802.11ac / ax protocol.
[0015] In the present invention, a router and a receiving terminal are set in the environment to be recognized, which can facilitate the generation of WiFi signals by the router and the reception of WiFi signals by the receiving terminal while generating feedback signals with BFI data. In this way, only the WiFi signal transceiver device needs to support the IEEE 802.11ac / ax protocol, and there is no need to make any hardware improvements to the router or the receiving terminal, greatly reducing the application cost and significantly improving the applicability of the system. By setting a data acquisition device configured with a data packet capture module, it is possible to facilitate the acquisition of WiFi data packets for recognition from the WiFi signals generated by the transceiver device. Enabling network communication between the data acquisition device and the recognition processing terminal can facilitate the recognition processing terminal to timely obtain the captured WiFi data packets. By setting a BFI data extraction module in the recognition processing terminal, it is possible to facilitate the extraction of BFI data from the received WiFi data packets. By setting a data preprocessing module in the recognition processing terminal, it can facilitate the preprocessing of BFI data, thereby facilitating the acquisition of more accurate BFI data and being conducive to obtaining more accurate recognition results. By setting a behavior recognition module in the recognition processing terminal, it can facilitate the use of the behavior recognition module to extract BFI features from BFI data based on the built-in feature extraction network, and can use the behavior recognition module to perform behavior recognition based on the BFI features based on the built-in classifier. By setting a domain adversarial training module in the recognition processing terminal, it can facilitate improving the ability of the feature extraction network to extract domain-invariant features through adversarial training. By setting a reinforcement learning optimization module in the recognition processing terminal, it can facilitate the optimization of the feature extraction network to improve the ability to dynamically adapt to domain differences.
[0016] The system has a simple structure, low deployment cost, wide application range, and high intelligence level. It can conveniently, accurately and efficiently recognize human behaviors in different recognition environments, greatly improving the user's human-computer interaction experience.
[0017] The present invention also provides a dynamic adversarial reinforcement learning cross-domain behavior recognition method based on BFI, which adopts a dynamic adversarial reinforcement learning cross-domain behavior recognition system based on BFI, including the following steps:
[0018] Step 1: Acquisition of sample WiFi data;
[0019] S11: In the activity space of the object to be recognized, use the router to emit WiFi signals. At the same time, use the receiving terminal to receive the WiFi signals and generate feedback signals with BFI data as recognition data. Execute this process in multiple different object activity spaces to obtain a large amount of recognition data;
[0020] S12: The data acquisition device uses the data packet capture module to capture a large number of WiFi data packets from the WiFi signals generated by the WiFi signal transceiver device, and sends the captured large number of WiFi data packets to the recognition and processing terminal;
[0021] Step 2: After receiving the WiFi data packets, the recognition and processing terminal constructs a cross-domain behavior recognition model based on BFI through the following process;
[0022] S21: First, use the BFI data extraction module to extract BFI data from the WiFi data packets, and then use the data preprocessing module to preprocess the BFI data; then, use all the preprocessed BFI data from different acquisition conditions as the data set X of N d source domains, and obtain the behavior label Y corresponding to each data, as well as the domain label d corresponding to each data;
[0023] S22: Adversarially train the feature extraction network through the domain adversarial training module, and use the reinforcement learning optimization module to optimize the feature extraction network parameters, and finally construct a cross-domain behavior recognition model based on BFI;
[0024] S22-1: Input the data set X into the shared feature extractor G θ for feature extraction to obtain temporal features, and then use the obtained temporal features to construct an input vector suitable for training to obtain the input feature G θ (X); where, θ is the parameter of the feature extractor;
[0025] S22-2: Input the input feature G θ (X) into the behavior classifier for behavior classification prediction, and calculate the classification loss according to formula (1)
[0026]
[0027] In the formula, N a is the number of behavior categories, is the true behavior label of the input data X, is the behavior classifier 's parameter;
[0028] At the same time, input the input feature G θ (X) into N d domain classifiers for domain classification, and calculate the domain classification loss according to formula (1)
[0029]
[0030] Wherein, N d is the number of domain categories, is the domain true label of the input data X, is the domain classifier parameters, where i = 1, …, N d ;
[0031] S22-3: Obtain the current environmental state vector s according to formula (3), select the action a according to the ∈-Greedy strategy, and update each λ i ←λ i +Δλ i , where λ i is the weight parameter of the loss of the i-th domain classifier, i = 1, …, N d , Δλ i respectively represent decreasing, remaining unchanged or increasing the weight, Δλ i ∈{-δ, 0, +δ}, for i = 1, …, N d ; δ is the preset step size;
[0032]
[0033] In the formula, is the current overall behavior classification accuracy rate, is the classification accuracy rate of each domain classifier;
[0034] S22-4: Obtain the total loss function according to formula (4) and update the total loss function using the newly obtained λ i
[0035] S22-5: Repeat S22-1 to S22-4 multiple times to continue training the shared feature extractor G θ , the behavior classifier and the domain classifier After reaching the set training cycle, first calculate the new environmental state vector s′ based on formula (3), and calculate the reward r according to formula (5); then store (s, a, r, s′) in the experience replay buffer, then sample data from the experience replay buffer, and calculate the loss function L(Θ) of the multi-layer perceptron network Q according to formula (6), update the parameters of the multi-layer perceptron network Q, and then synchronize the parameters of the multi-layer perceptron network Q to the target multi-layer perceptron network Q regularly;
[0036]
[0037] L(Θ) = E[(y - Q(s, a; Θ)) 2 (6);
[0038] where α is the balance factor; Θ is the parameter of the multi-layer perceptron network Q, Q(s,a;Θ) is the Q value of each possible action corresponding to the current environmental state vector s as the input data, and y = r + γmax a′ Q(s′,a′;Θ′), where γ is the discount factor and Θ′ is the parameter of the target multi-layer perceptron network Q;
[0039] S22-6: Repeat S22-5 multiple times until the set total number of training rounds is reached, then stop the training process and construct the final cross-domain behavior recognition model based on BFI;
[0040] Step 3: Use the cross-domain behavior recognition model based on BFI to recognize real-time behaviors;
[0041] S31: In the activity space of the object to be recognized, use the router to emit WiFi signals in real time. At the same time, use the receiving terminal to receive the WiFi signals in real time and generate a feedback signal with BFI data as the recognition data;
[0042] S32: The data acquisition device uses the packet capture module to capture real-time WiFi packets from the WiFi signals generated by the transceiver device and send them to the recognition processing terminal;
[0043] S33: After receiving the real-time WiFi packets, the recognition processing terminal first uses the BFI data extraction module to extract the real-time BFI data from the real-time WiFi packets, and then uses the data preprocessing module to preprocess the real-time BFI data; then use the preprocessed real-time BFI data as the input data and input it into the cross-domain behavior recognition model based on BFI, and the cross-domain behavior recognition model based on BFI performs behavior recognition and outputs the behavior recognition result.
[0044] Furthermore, in order to ensure the accuracy and efficiency of subsequent recognition, in S21 of Step 2, the preprocessing includes extracting amplitude information data, resampling, denoising, standardizing, and outlier processing of the data.
[0045] In the present invention, during the sample data collection process, a large number of WiFi data packets are generated by routers and terminals arranged in various sensing environments, which can ensure that a recognition model with higher recognition accuracy and simpler structure can be trained subsequently. During the training process, first, the BFI data extraction module extracts the BFI data containing human behavior information from the BFI data collected from multiple environments, and then preprocesses it to ensure the accuracy of the sample data. On this basis, by sending data from multiple source domains into the shared feature extractor, the core features related to behavior classification can be extracted, and at the same time, the domain differences between different source domains can be effectively suppressed, enabling the subsequent classifier to better generalize to the unseen target domain. Using the behavior classifier to perform behavior recognition based on the input features can ensure that the model can accurately recognize each action on the known source domain data, thereby providing a reliable supervision signal for domain adversarial training and greatly improving the training effect and efficiency. Using multiple domain classifiers set in parallel to perform domain classification based on the input features can effectively predict the source domain to which the current sample belongs. Then, based on the domain adversarial training module, cross-domain recognition training is carried out in the dynamic adversarial reinforcement learning framework, which can force the feature extractor to learn domain-invariant feature representations, making it impossible for each domain classifier to accurately determine the sample source, which is beneficial to improving the cross-domain recognition ability. In actual multi-source domain adversarial training, how to balance the behavior classification loss and the losses of each domain classifier is crucial. Too strong an adversarial effect may lead to a decline in behavior classification performance, while insufficient adversarial training cannot fully eliminate domain differences. On this basis, the reinforcement learning optimization module is used to optimize the parameters of the feature extraction network, and the loss weights of each domain classifier can be adjusted according to the metrics calculated on the source domain data. At the same time, the environmental state vector is designed as the accuracy rates of the behavior classifier and each domain classifier and their corresponding loss weights, the action of the agent is designed as the adjustment direction for each weight, the reward is designed as an index comprehensively considering the behavior classification performance and the domain classification confusion degree, and the state and action are used to form an experience. The DQN network is used to update the parameters to find the optimal weight parameter combination, which can maximize the behavior classification accuracy rate, and at the same time make each domain classifier close to random performance, that is, the domain adversarial effect is good, enabling the recognition model to dynamically adapt to domain differences and being beneficial to optimizing the target domain performance. Thus, through the cooperation of the domain adversarial training module and the reinforcement learning optimization module, the feature extraction network and the classifier can be efficiently trained, and the recognition performance in the cross-domain situation can be optimized. During the actual recognition application process, the BFI data collected in real time in the specific usage environment is sent into the trained cross-domain behavior recognition model, and after extracting the behavior features, they are sent into the classifier for recognition, and the recognition result can be accurately obtained.
[0046] This method makes full use of the supervision information of multiple source domains, and optimizes the weights in real time according to the training process and the performance of each module. It does not require manual search for the optimal weight combination, but completely relies on the metrics of the source domain data for optimization, which conforms to the actual situation where the target domain labels are unknown. By combining domain adversarial and reinforcement learning, cross-domain behavior detection is achieved without target domain labels. This method has simple implementation steps, low implementation cost, high detection efficiency, high intelligence level, and accurate detection results. It can accurately and efficiently identify human behaviors in different target domains using easily accessible BFI data, which is beneficial to improving the human-computer interaction experience and is suitable for large-scale promotion and application. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 is the structural schematic diagram of the system part in the present invention;
[0048] Figure 2 is the flowchart of the recognition method part of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0049] The present invention will be further described below with reference to the drawings and embodiments.
[0050] As Figure 1 shown, the present invention provides a cross-domain behavior recognition system based on BFI, including a WiFi signal transceiver device, a data acquisition device, and an identification and processing terminal;
[0051] The WiFi signal transceiver device is arranged inside the activity space of the object to be recognized. The WiFi signal transceiver device includes a router and a receiving terminal. The router is used to emit WiFi signals, and the receiving terminal is used to receive WiFi signals and generate a feedback signal with BFI data as identification data;
[0052] The data acquisition device is internally provided with a data packet capture module, and the data packet capture module is used to collect WiFi data packets with identification data from the WiFi signals generated by the WiFi signal transceiver device;
[0053] The recognition processing terminal is connected to the data acquisition device in a wired or wireless manner. The recognition processing terminal includes a BFI data extraction module, a data preprocessing module, a domain adversarial training module, a reinforcement learning optimization module, and a behavior recognition module. The BFI data extraction module is used to extract BFI data from the collected WiFi packets. The data preprocessing module is used to preprocess the BFI data. The behavior recognition module is used to extract BFI features from the BFI data by using a built-in feature extraction network. At the same time, a built-in classifier is used to perform behavior recognition based on the BFI features and output a behavior recognition result. The domain adversarial training module is used to adversarially train the feature extraction network so that the feature extraction network extracts domain-invariant features. The reinforcement learning optimization module is used to optimize the parameters of the feature extraction network so that the feature extraction network dynamically adapts to domain differences and optimizes the performance of the target domain.
[0054] As a preference, the BFI data extraction module extracts BFI data from the WiFi packets through the Wi-BFI tool library;
[0055] As a preference, the data acquisition device is arranged inside the space of the object to be recognized.
[0056] As a preference, the packet capture module is a network packet capture tool.
[0057] To ensure powerful processing capabilities, the recognition processing terminal is a remote server.
[0058] To ensure accurate and reliable acquisition of BFI data, the WiFi signal transceiver device supports the IEEE802.11ac / ax protocol.
[0059] In the present invention, a router and a receiving terminal are set in the environment to be recognized, which facilitates the generation of WiFi signals by the router and the reception of WiFi signals by the receiving terminal while generating a feedback signal with BFI data. In this way, only the WiFi signal transceiver device needs to support the IEEE 802.11ac / ax protocol, and there is no need to make any hardware improvements to the router or the receiving terminal, greatly reducing the application cost and significantly improving the applicability of the system. By setting a data acquisition device configured with a data packet capture module, it is possible to facilitate the acquisition of WiFi data packets for recognition from the WiFi signals generated by the transceiver device. Enabling network communication between the data acquisition device and the recognition processing terminal can facilitate the recognition processing terminal to timely obtain the captured WiFi data packets. By setting a BFI data extraction module in the recognition processing terminal, it is possible to facilitate the extraction of BFI data from the received WiFi data packets. By setting a data preprocessing module in the recognition processing terminal, it is possible to facilitate the preprocessing of BFI data, thereby facilitating the acquisition of more accurate BFI data and being conducive to obtaining more accurate recognition results. By setting a behavior recognition module in the recognition processing terminal, it is possible to facilitate the use of the behavior recognition module to extract BFI features from BFI data based on the built-in feature extraction network, and to use the behavior recognition module to perform behavior recognition based on the BFI features based on the built-in classifier. By setting a domain adversarial training module in the recognition processing terminal, it is possible to facilitate the improvement of the extraction ability of the feature extraction network for domain-invariant features through adversarial training. By setting a reinforcement learning optimization module in the recognition processing terminal, it is possible to facilitate the optimization of the feature extraction network to improve the ability to dynamically adapt to domain differences.
[0060] The system has a simple structure, low deployment cost, wide application range, and high degree of intelligence. It can conveniently, accurately and efficiently recognize human behaviors in different recognition environments, greatly improving the user's human-computer interaction experience.
[0061] As Figure 2 shown, the present invention also provides a dynamic adversarial reinforcement learning cross-domain behavior recognition method based on BFI, which adopts a dynamic adversarial reinforcement learning cross-domain behavior recognition system based on BFI, including the following steps:
[0062] Step 1: Acquisition of sample WiFi data;
[0063] S11: In the activity space of the object to be recognized, use the router to emit WiFi signals. At the same time, use the receiving terminal to receive the WiFi signals and generate a feedback signal with BFI data as recognition data. Execute this process in multiple different object activity spaces to obtain a large amount of recognition data;
[0064] S12: The data acquisition device captures a large number of WiFi data packets from the WiFi signals generated by the WiFi signal transceiver device using the data packet capture module, and sends the captured large number of WiFi data packets to the recognition and processing terminal;
[0065] Step 2: After receiving the WiFi data packets, the recognition and processing terminal constructs a cross-domain behavior recognition model based on BFI through the following process;
[0066] S21: First, use the BFI data extraction module to extract BFI data from the WiFi data packets, and then use the data preprocessing module to preprocess the BFI data; then, use all the preprocessed BFI data from different acquisition conditions as the data set X of N d source domains, and obtain the behavior label Y corresponding to each data, and the domain label d corresponding to each data;
[0067] S22: Adversarially train the feature extraction network through the domain adversarial training module, and use the reinforcement learning optimization module to optimize the feature extraction network parameters, and finally construct a cross-domain behavior recognition model based on BFI;
[0068] S22-1: Input the data set X into the shared feature extractor G θ for feature extraction to obtain temporal features, and then use the obtained temporal features to construct an input vector suitable for training to obtain the input feature G θ (X); where θ is the parameter of the feature extractor;
[0069] S22-2: Input the input feature G θ (X) into the behavior classifier for behavior classification prediction, and calculate the classification loss according to formula (1)
[0070]
[0071] In the formula, N a is the number of categories of behaviors, is the true behavior label of the input data X, is the parameter of the behavior classifier ;
[0072] At the same time, input the input feature G θ (X) into N d domain classifiers for domain classification, and calculate the domain classification loss according to formula (1)
[0073]
[0074] Wherein, N d is the number of domain categories, is the domain true label of the input data X, is the domain classifier parameters, where i = 1, …, N d ;
[0075] S22-3: Obtain the current environmental state vector s according to formula (3), select the action a according to the ∈-Greedy strategy, and update each λ i ←λ i +Δλ i where λ i is the weight parameter of the loss of the i-th domain classifier, i = 1, …, N d , Δλ i respectively represent decreasing, remaining unchanged or increasing the weight, Δλ i ∈{-δ, 0, +δ}, for i = 1, …, N d ; δ is the preset step size;
[0076]
[0077] Wherein, is the current overall behavior classification accuracy rate, is the classification accuracy rate of each domain classifier;
[0078] S22-4: Obtain the total loss function according to formula (4) and update the total loss function using the newly obtained λ i
[0079]
[0080] S22-5: Repeatedly execute S22-1 to S22-4 multiple times to continue training the shared feature extractor G θ , the behavior classifier and the domain classifier After reaching the set training cycle, first calculate the new environmental state vector s′ based on formula (3), and calculate the reward r according to formula (5); then store (s, a, r, s′) in the experience replay buffer, then sample data from the experience replay buffer, and calculate the loss function L(Θ) of the multi-layer perceptron network Q according to formula (6), update the parameters of the multi-layer perceptron network Q, and then synchronize the parameters of the multi-layer perceptron network Q to the target multi-layer perceptron network Q regularly;
[0081]
[0082] L(Θ) = E[(y - Q(s, a; Θ)) 2 (6);
[0083] In the formula, α is the balance factor; Θ is the parameter of the multi-layer perceptron network Q, Q(s,a;Θ) is the Q value corresponding to each possible action based on the current environmental state vector s as the input data, and y = r + γmax a′ Q(s′,a′;Θ′), where γ is the discount factor and Θ′ is the parameter of the target multi-layer perceptron network Q;
[0084] S22-6: Repeat S22-5 multiple times until the set total number of training rounds is reached, then stop the training process to construct the final cross-domain behavior recognition model based on BFI;
[0085] Step Three: Use the cross-domain behavior recognition model based on BFI to recognize real-time behaviors;
[0086] S31: In the activity space of the object to be recognized, use the router to emit WiFi signals in real time. At the same time, use the receiving terminal to receive the WiFi signals in real time and generate a feedback signal with BFI data as the recognition data;
[0087] S32: The data acquisition device uses the packet capture module to capture real-time WiFi packets from the WiFi signals generated by the transceiver device and send them to the recognition processing terminal;
[0088] S33: After receiving the real-time WiFi packets, the recognition processing terminal first uses the BFI data extraction module to extract the real-time BFI data from the real-time WiFi packets, and then uses the data preprocessing module to preprocess the real-time BFI data; then use the preprocessed real-time BFI data as the input data and input it into the cross-domain behavior recognition model based on BFI, and the cross-domain behavior recognition model based on BFI performs behavior recognition and outputs the behavior recognition result.
[0089] To ensure the accuracy and efficiency of subsequent recognition, in S21 of Step Two, the preprocessing includes extracting amplitude information data, resampling, denoising, normalizing, and outlier processing of the data.
[0090] In the present invention, during the sample data collection process, a large number of WiFi data packets are generated by routers and terminals arranged in various sensing environments, which can ensure that a recognition model with higher recognition speed and simpler structure can be trained subsequently. During the training process, first, the BFI data extraction module extracts the BFI data containing human behavior information from the BFI data collected from multiple environments, and then preprocesses it to ensure the accuracy of the sample data. On this basis, by sending data from multiple source domains into the shared feature extractor, the core features related to behavior classification can be extracted, and at the same time, the domain differences between different source domains can be effectively suppressed, enabling the subsequent classifier to better generalize to unseen target domains. Using the behavior classifier to perform behavior recognition based on the input features can ensure that the model can accurately recognize each action on the known source domain data, thereby providing a reliable supervision signal for domain adversarial training and greatly improving the training effect and efficiency. Using multiple domain classifiers set in parallel to perform domain classification based on the input features can effectively predict the source domain to which the current sample belongs. Then, based on the domain adversarial training module, cross-domain recognition training is carried out in the dynamic adversarial reinforcement learning framework, which can force the feature extractor to learn domain-invariant feature representations, making it impossible for each domain classifier to accurately judge the sample source, which is beneficial to improving the cross-domain recognition ability. In actual multi-source domain adversarial training, how to balance the behavior classification loss and the losses of each domain classifier is crucial. Too strong an adversarial effect may lead to a decline in behavior classification performance, while insufficient adversarial training cannot fully eliminate domain differences. On this basis, the reinforcement learning optimization module is used to optimize the parameters of the feature extraction network, and the loss weights of each domain classifier can be adjusted according to the metrics calculated on the source domain data. At the same time, the environmental state vector is designed as the accuracy of the behavior classifier and each domain classifier and their corresponding loss weights, the action of the agent is designed as the adjustment direction for each weight, the reward is designed as an index comprehensively considering the behavior classification performance and the domain classification confusion degree, and the state and action are used to form an experience. The DQN network is used to update the parameters to find the optimal weight parameter combination, which can maximize the behavior classification accuracy, and at the same time make each domain classifier close to random performance, that is, the domain adversarial effect is good, enabling the recognition model to dynamically adapt to domain differences and being beneficial to optimizing the target domain performance. Thus, through the cooperation of the domain adversarial training module and the reinforcement learning optimization module, the feature extraction network and the classifier can be efficiently trained, and the recognition performance in the cross-domain situation can be optimized. During the actual recognition application process, the BFI data collected in real time in the specific usage environment is sent into the trained cross-domain behavior recognition model, and after extracting the behavior features, they are sent into the classifier for recognition, and the recognition result can be accurately obtained.
[0091] This method makes full use of the supervision information of multiple source domains, and in accordance with the training process and the performance of each module, optimizes the weights in real time. It does not require manual search for the optimal weight combination, but fully relies on the metrics of the source domain data for optimization, which conforms to the actual situation where the target domain labels are unknown. By combining domain adversarial and reinforcement learning, cross-domain behavior detection is achieved without target domain labels. The implementation steps of this method are simple, the implementation cost is low, the detection efficiency is high, the degree of intelligence is high, and the detection results are accurate. It can accurately and efficiently identify human behaviors in different target domains using easily obtainable BFI data, which is conducive to improving the human-computer interaction experience and is suitable for large-scale popularization and application.
Claims
1. A dynamic adversarial reinforcement learning cross-domain behavior recognition system based on BFI, characterized in that: Including WiFi signal transceiver equipment, data collection equipment and identification processing terminal; The WiFi signal transceiver is arranged inside the activity space of the object to be identified, and the WiFi signal transceiver includes a router and a receiving terminal, the router is used to send WiFi signals, and the receiving terminal is used to receive WiFi signals and generate a feedback signal with BFI data as identification data; The data acquisition device is internally provided with a data packet capture module, and the data packet capture module is used to collect WiFi data packets with identification data from the WiFi signal generated by the WiFi signal transceiver device; The identification processing terminal is connected to the data acquisition device in a wired or wireless manner. The identification processing terminal includes a BFI data extraction module, a data preprocessing module, a domain adversarial training module, a reinforcement learning optimization module and a behavior recognition module. The BFI data extraction module is used to extract BFI data from the collected WiFi data packets, the data preprocessing module is used to preprocess the BFI data, the behavior recognition module is used to extract BFI features from the BFI data using a built-in feature extraction network, and at the same time, use a built-in classifier to perform behavior recognition based on the BFI features and output the behavior recognition results; the domain adversarial training module is used to adversarially train the feature extraction network so that the feature extraction network extracts domain-invariant features; the reinforcement learning optimization module is used to optimize the feature extraction network parameters so that the feature extraction network dynamically adapts to domain differences and optimizes the target domain performance.
2. According to claim 1, a BFI-based dynamic adversarial reinforcement learning cross-domain behavior recognition system is characterized in that: The data acquisition device is arranged inside the space of the object to be identified.
3. A cross-domain behavior recognition system based on dynamic adversarial reinforcement learning based on BFI according to claim 1 or 2, characterized in that: The data packet capture module is a network packet capture tool.
4. According to claim 3, a BFI-based dynamic adversarial reinforcement learning cross-domain behavior recognition system is characterized in that: The identification processing terminal is a remote server.
5. According to claim 4, a BFI-based dynamic adversarial reinforcement learning cross-domain behavior recognition system is characterized in that: The WiFi signal transceiver device supports the IEEE 802.11ac / ax protocol.
6. A method for cross-domain behavior recognition based on dynamic adversarial reinforcement learning based on BFI, using a cross-domain behavior recognition system based on dynamic adversarial reinforcement learning based on BFI according to any one of claims 1 to 5, characterized in that: The steps include: Step 1: Collection of sample WiFi data; S11: In the activity space of the object to be identified, a WiFi signal is sent out by a router, and at the same time, a receiving terminal is used to receive the WiFi signal and generate a feedback signal with BFI data as identification data. This process is performed in multiple different object activity spaces to obtain a large amount of identification data; S12: The data acquisition device uses the data packet capture module to capture a large number of WiFi data packets from the WiFi signal generated by the WiFi signal transceiver device, and sends the captured large number of WiFi data packets to the identification and processing terminal; Step 2: After receiving the WiFi data packet, the identification processing terminal builds a cross-domain behavior identification model based on BFI through the following process; S21: First, use the BFI data extraction module to extract BFI data from the WiFi data packet, and then use the data preprocessing module to preprocess the BFI data; then, all the preprocessed BFI data from different acquisition conditions are used as N d A data set L of the source domain is obtained, and the behavior label Y corresponding to each data and the domain label d corresponding to each data are obtained; S22: The feature extraction network is adversarially trained through the domain adversarial training module, and the parameters of the feature extraction network are optimized using the reinforcement learning optimization module, and finally a cross-domain behavior recognition model based on BFI is constructed; S22-1: Input the data set X to the shared feature extractor G θ Feature extraction is performed in , and time series features are obtained. Then, the obtained time series features are used to construct an input vector suitable for training, and the input feature G is obtained. θ (X); where θ is the parameter of the feature extractor; S22-2: Input feature G θ (X) Input to the behavior classifier Behavior classification prediction is performed in and the classification loss is calculated according to formula (1) Where N a is the number of behavior categories, is the true value label of the input data X, is a behavior classifier Parameters; At the same time, the input feature G θ (X) respectively input to N d Domain Classifier In the above formula, we perform domain classification and calculate the domain classification loss according to formula (1): Where N d is the number of categories in the domain, is the domain truth label of the input data X, is a domain classifier Parameters, where i = 1,…,N d ; S22-3: Obtain the current environment state vector s according to formula (3), select action a according to the ∈-Greedy strategy, and update each λ i ←λ i +Δλ i , where λ i is the weight parameter of the i-th domain classifier loss, i = 1, ..., N d , Δλ i Respectively represent reducing, unchanged or increasing weight, Δλ i ∈{-δ,0,+δ},fori=1,…,N d ;δ is the preset step size; In the formula, is the current overall behavior classification accuracy, is the classification accuracy of each domain classifier; S22-4: Obtain the total loss function according to formula (4) And using the newly obtained λ i Update the total loss function S22-5: Repeat S22-1 to S22-4 multiple times to continue training the shared feature extractor G θ , Behavior Classifier and domain classifier After reaching the set training cycle, the new environment state vector s′ is first calculated based on formula (3), and the reward r is calculated according to formula (5); then (s, a, r, s′) is stored in the experience replay buffer, and then data is sampled from the experience replay buffer, and the loss function L(Θ) of the multilayer perceptron network Q is calculated according to formula (6), the parameters of the multilayer perceptron network Q are updated, and then the parameters of the multilayer perceptron network Q are periodically synchronized to the target multilayer perceptron network Q; Where α is the balance factor; Θ is the parameter of the multilayer perceptron network Q, Q(s,a;Θ) is the Q value of each possible action corresponding to the current environment state vector s as input data, y = r + γmax a′ Q(s′,a′; Θ′), where γ is the discount factor and Θ′ is the parameter of the target multilayer perceptron network Q; S22-6: Repeat S22-5 for multiple times until the set total number of training rounds is reached, stop the training process, and build the final BFI-based cross-domain behavior recognition model; Step 3: Use the BFI-based cross-domain behavior recognition model to identify real-time behaviors; S31: In the activity space of the object to be identified, a router is used to send a WiFi signal in real time, and at the same time, a receiving terminal is used to receive the WiFi signal in real time and generate a feedback signal with BFI data as identification data; S32: The data acquisition device uses the data packet capture module to capture the real-time WiFi data packet from the WiFi signal generated by the transceiver device, and sends it to the identification processing terminal; S33: After receiving the real-time WiFi data packet, the identification and processing terminal first uses the BFI data extraction module to extract the real-time BFI data from the real-time WiFi data packet, and then uses the data preprocessing module to preprocess the real-time BFI data; then the preprocessed real-time BFI data is input as input data to the BFI-based cross-domain behavior recognition model, and the BFI-based cross-domain behavior recognition model performs behavior recognition and outputs the behavior recognition result.
7. According to claim 6, a method for cross-domain behavior recognition based on dynamic adversarial reinforcement learning based on BFI is characterized in that: In step 2 S21, the preprocessing includes extracting amplitude information data, resampling, denoising, standardizing and outlier processing of the data.
Citation Information
Patent Citations
Battlefield target cross-domain identification method and system based on deep reinforcement learning
CN118799559A
Cited By
User feedback classification method and electronic equipment
CN121093097A