Human body activity behavior recognition system and method, electronic equipment and storage medium

By designing a human activity behavior recognition system based on CSI signals and using a deep network for model training, the problem of low accuracy of human activity recognition in the prior art is solved, and high-precision and low-cost human activity monitoring is achieved.

CN120223565APending Publication Date: 2025-06-27BOLIU INTELLIGENT TECH (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311817433.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing techniques for detecting human activities using CSI signals are low in accuracy, and more data are used for model training.

Method used

A human activity behavior recognition system is designed to collect CSI signal data under the Wifi wireless network through the data collection module, the feature extraction module extracts CSI signal characteristics, the AI ​​training module uses the sound classification deep network for model training, and the AI ​​inference module makes predictions to identify human activity behavior.

Benefits of technology

It improves the accuracy of human activities recognition, ensures user privacy and security, and at the same time, the system cost is low.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223565A_ABST
    Figure CN120223565A_ABST
Patent Text Reader

Abstract

The invention discloses a human body activity behavior recognition system and method, electronic equipment and a storage medium. The human body activity behavior recognition system comprises a data collection module, a feature extraction module, an AI training module and an AI inference module. The data collection module is used for collecting CSI signal data under a Wifi wireless network in a manned or unmanned environment to form a data set; the feature extraction module is used for extracting CSI signal features from the data collected by the data collection module; the AI training module is used for marking data sets collected by the data collection module in a manned environment and an unmanned environment, and training a network by utilizing the characteristic that a voice classification deep network performs model training through a voice spectrogram; and the AI inference module is used for performing prediction by utilizing the characteristic variable to form an inference value, and performing final prediction by utilizing the inference value to obtain human body activity behaviors. According to the invention, human body activity monitoring can be carried out by using wireless signals, so that the privacy of a user is guaranteed; in addition, the whole system is extremely low in cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of human activity recognition, and relates to a human activity behavior recognition system, and particularly to a human activity behavior recognition system, method, electronic device and storage medium based on CSI signals. Background Art

[0002] Channel State Information (CSI for short), a term in the field of wireless communication, refers to the channel attributes of a communication link. It describes the attenuation factors of the signal on each transmission path, that is, the value of each element in the channel gain matrix H, such as signal scattering, environmental attenuation (fading, multipath fading or shadowing fading), distance attenuation (power decay of distance), etc. CSI can enable a communication system to adapt to the current channel conditions and provide guarantee for high-reliability and high-rate communication in a multi-antenna system. Generally, the receiving end evaluates CSI and quantifies and feeds it back to the transmitting end (in a time-division duplex system, reverse evaluation is required). Therefore, CSI can be divided into CSIR and CSIT.

[0003] Based on the research on the characteristics of CSI signals, technologies for detecting human activities using CSI signals have emerged; existing solutions for detecting human activities using CSI signals usually only use the amplitude or phase derived from I and Q points as the input data of the AI model, mostly with an image recognition model as the backbone, mostly use a time interval for prediction, and use less data for model training; this makes the accuracy of human activity detection solutions in the prior art low.

[0004] In view of this, there is an urgent need to design a new human activity recognition method today to overcome at least some of the above defects existing in the existing human activity recognition methods. Summary of the Invention

[0005] The present invention provides a human activity behavior recognition system, method, electronic device and storage medium, which can use wireless signals for human activity monitoring, ensuring user privacy while having a low system cost.

[0006] To solve the above technical problems, according to one aspect of the present invention, the following technical solution is adopted:

[0007] A human activity behavior recognition system, the human activity behavior recognition system includes:

[0008] A data collection module for collecting CSI signal data under a Wifi wireless network in both occupied and unoccupied environments to form a data set;

[0009] A feature extraction module for extracting CSI signal features from the data collected by the data collection module. The CSI signal is used to extract features from the points on the constellation diagram showing the relationship between the modulation signal distribution and the digital coding bits. The CSI signal features include IQ point data, and the IQ point data includes the real number I on the X-axis and the imaginary number Q on the Y-axis.

[0010] An AI training module for annotating the data sets collected by the data collection module in the presence and absence of people, and training the network by utilizing the characteristics of the sound classification deep network for model training through voice spectrograms.

[0011] An AI inference module for making predictions using feature variables to form inference values, and using the inference values to obtain the final prediction values to get human activity behaviors.

[0012] As an implementation manner of the present invention, the feature extraction module includes a CSI signal feature extraction unit; the CSI signal feature extraction unit is used to extract CSI signal features from the data collected by the data collection module.

[0013] The feature extraction module is also used to extract appropriate features in combination with the sound classification deep network under the condition of large-scale data collection; the sound classification network extracts features for the sound in the frequency domain segment. Although this system is not a sound signal, since CSI is data in the frequency domain segment, the sound classification network is used as the deep learning network architecture.

[0014] The deep learning model adopts the sound classification deep network because the input feature of the sound classification model itself is the spectrum, and the CSI signal is also presented as signals in different frequency segments. Therefore, its AI model architecture is applied to this system.

[0015] As an implementation manner of the present invention, the AI training module uses the IQ point data of 56 subcarriers of the CSI signal as the input of the model to train the model; the sum of the real number Q and the imaginary number I of the IQ point data is 112 data.

[0016] The AI training module is used to annotate the presence and absence of people environment; for the data of different time periods when a person stands between the development board and the router, it is manually marked as "occupied"; if it is data of an unoccupied time period, it is marked as "unoccupied"; through marking, the AI model can effectively distinguish the input environmental signal features into the environmental states of occupied or unoccupied through real labels.

[0017] The input data trained by the AI training module includes 200,000 pieces of occupied data and 200,000 pieces of unoccupied data. The number of occupied data and unoccupied data is equivalent to balance the number of category data, and an appropriate proportion is extracted for the environments of different scenarios, so that the data will not be biased towards a specific scenario.

[0018] The training will be carried out in batches of 400 according to the batch size and put into the model. It will continuously find the solution with the minimum error to update the weight values and repeat the execution ten times. Finally, the test data in different scenarios will be collected to confirm the effectiveness of the model accuracy.

[0019] As an implementation manner of the present invention, the AI inference module uses the weights of each layer of the trained sound classification deep network to update the input data and predict the results. Specifically, the sound classification deep network extracts various convolutional features from 112 IQ points of the input data, and finally extracts a total of 512 features. Then, through the update of the feature weights of the fully connected layer and the activation function for non-linear classification, the inference result of whether there is someone or no one will be obtained finally.

[0020] According to another aspect of the present invention, the following technical solution is adopted: A method for recognizing human activity behaviors, the method for recognizing human activity behaviors includes:

[0021] Data collection step; collecting CSI signal data under Wifi wireless network in the presence and absence of people to form a data set;

[0022] Feature extraction step; extracting CSI signal features from the data collected in the data collection step. The CSI signal features are extracted from the points on the constellation diagram of the relationship between the modulation signal distribution and the digital coding bits. The CSI signal features include IQ point data, and the IQ point data includes the real number I on the X-axis and the imaginary number Q on the Y-axis;

[0023] AI training step; annotating the data set collected in the data collection step and training the network by using the characteristics of the sound classification deep network through the voice spectrogram for model training;

[0024] AI inference step; using the feature variables for prediction to form inference values, and using the inference values as the final prediction values to obtain human activity behaviors.

[0025] As an implementation manner of the present invention, in the feature extraction step, appropriate features can be extracted by cooperating with the sound classification deep network under the condition of collecting data in a large range; the sound classification network will extract features for the sound in the frequency domain segment. Although the system is not a sound signal, because CSI is data in the frequency domain segment, the sound classification network is used as the deep learning network architecture.

[0026] The deep learning model adopts the sound classification deep network because the input feature of the sound classification model itself is the spectrum, and the CSI signal is also presented as signals in different frequency segments. Therefore, its AI model architecture is applied to this system;

[0027] As an implementation of the present invention, in the AI training step, the IQ point data of 56 subcarriers of the CSI signal is used as the input of the model to train the model; the sum of the real number Q and the imaginary number I of the IQ point data is 112 data in total.

[0028] In the AI training step, the environment with people and the environment without people are labeled; for the data of different time periods when people stand between the development board and the router, it is manually marked as having people; if it is the data of the time period without people, it is marked as without people; through marking, the AI model can effectively distinguish the input environmental signal features into the environmental states of having people or without people through real labels.

[0029] The input data for training includes data with people and data without people, and the quantities of data with people and data without people are equivalent, so as to balance the quantity of category data, and for the environments of different scenarios, an appropriate proportion is extracted, and the data will not be biased towards a specific scenario.

[0030] During training, the data is put into the model in batches, and a solution that minimizes the error is continuously found to update the weight values, and this is repeated several times. Finally, the test data of different scenarios is collected to confirm the effectiveness of the model accuracy.

[0031] As an implementation of the present invention, in the AI inference step, the weights of each layer of the trained deep network for voice classification are used to update the input data and predict the result; specifically, the deep network for voice classification extracts various convolution features from the 112 IQ points of the input data, and finally a total of 512 features are extracted. Then, after the feature weights of the fully connected layer are updated and the activation function is used for non-linear classification, the inference result, whether there are people or not, will finally be obtained.

[0032] According to another aspect of the present invention, the following technical solution is adopted: an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, the steps of the above method are implemented.

[0033] According to another aspect of the present invention, the following technical solution is adopted: a storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by the processor, the steps of the above method are implemented.

[0034] The beneficial effect of the present invention is as follows: The human activity behavior recognition system, method, electronic device and storage medium proposed by the present invention can use wireless signals to monitor human activities, ensuring the privacy of users; in addition, the construction of the entire system is simple, and identification can be directly performed through a wifi machine and a receiving chip in the space, and overall, the cost is extremely low. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1Schematic diagram of the composition of the human activity behavior recognition system in an embodiment of the present invention.

[0036] Figure 2 Flowchart of the human activity behavior recognition method in an embodiment of the present invention.

[0037] Figure 3 Schematic diagram of the IQ point amplitude derivation in an embodiment of the present invention.

[0038] Figure 4 Schematic diagram of the IQ point phase derivation in an embodiment of the present invention.

[0039] Figure 5 Working schematic diagram of the CSI stable rotation structure in an embodiment of the present invention.

[0040] Figure 6 Working schematic diagram of the large - range data collection module in an embodiment of the present invention.

[0041] Figure 7 Working schematic diagram of the CSI signal feature extraction module in an embodiment of the present invention.

[0042] Figure 8 Working schematic diagram of the AI training module in an embodiment of the present invention.

[0043] Figure 9 Working schematic diagram of the AI inference module in an embodiment of the present invention.

[0044] Figure 10 Schematic diagram of the composition of the electronic device in an embodiment of the present invention. Detailed implementation manners

[0045] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0046] In order to further understand the present invention, the preferred implementation manners of the present invention will be described below in conjunction with embodiments. However, it should be understood that these descriptions are only for further explaining the features and advantages of the present invention, rather than limiting the claims of the present invention.

[0047] The description of this part only focuses on several typical embodiments, and the present invention is not limited to the scope described in the embodiments. The mutual replacement of some technical features between the same or similar prior art means and the embodiments is also within the scope of the description and protection of the present invention.

[0048] The expression of the steps in each embodiment in the specification is only for convenience of description, and the implementation manner of the present application is not limited by the order of step implementation.

[0049] "Coupled" or "connected" in the specification includes both direct connection.

[0050] The present invention discloses a human activity behavior recognition system. Figure 1 It is a schematic diagram of the composition of the human activity behavior recognition system in an embodiment of the present invention; please refer to Figure 1 The human activity behavior recognition system includes: a data collection module 1, a feature extraction module 2, an AI training module 3, and an AI inference module 4.

[0051] The data collection module 1 is used to collect CSI signal data under Wi-Fi wireless networks in both occupied and unoccupied environments to form a data set.

[0052] The feature extraction module 2 is used to extract CSI signal features from the data collected by the data collection module. The CSI signal extracts features from the points on the constellation diagram of the relationship between the modulation signal distribution and the digital coding bits. The CSI signal features include IQ point data, and the IQ point data includes the real number I on the X-axis and the imaginary number Q on the Y-axis; the CSI signal features include IQ point data, and the IQ point data includes the real number Q and the imaginary number I (the real number I (in phase) and the imaginary number Q (quadrature phase)).

[0053] The AI training module 3 is used to label the data set collected by the data collection module in both occupied and unoccupied environments, and use the characteristics of the sound classification deep network to train the network through the voice spectrogram.

[0054] The AI inference module 4 is used to make predictions using the feature variables to form inference values, and use the inference values as the final prediction values to obtain the human activity behavior.

[0055] In an embodiment of the present invention, the feature extraction module 2 includes a CSI signal feature extraction unit; the CSI signal feature extraction unit is used to extract CSI signal features from the data collected by the data collection module;

[0056] The feature extraction module 2 can also be used to extract appropriate features in combination with the sound classification deep network under large-scale data collection; the sound classification network extracts features for the sound in the frequency domain segment. Although this system is not a sound signal, because CSI is data in the frequency domain segment, the sound classification network is used as the deep learning network architecture.

[0057] The deep learning model uses the sound classification deep network because the input feature of the sound classification model itself is the spectrum, and the CSI signal is also presented as signals in different frequency segments. Therefore, its AI model architecture is applied to this system;

[0058] The AI training module 3 uses the IQ point data of 56 subcarriers of the CSI signal as the input of the model to train the model; the sum of the real number Q and the imaginary number I of the IQ point data is 112 data in total.

[0059] The AI training module 3 is used to label the occupied environment and unoccupied environment; for the data of different time periods when a person stands between the development board and the router, it is manually marked as occupied; if it is the data of an unoccupied time period, it is marked as unoccupied; through marking, the AI model can effectively distinguish the input environmental signal features into occupied or unoccupied environmental states through real labels;

[0060] The input data trained by the AI training module 3 includes 200,000 pieces of occupied data and 200,000 pieces of unoccupied data. The number of occupied data and unoccupied data is equivalent to balance the number of category data, and for the environments of different scenarios, an appropriate proportion is extracted so that the data will not be biased towards a specific scenario;

[0061] During training, with a batch size of 400, the data is put into the model in batches, and the scheme that minimizes the error is continuously found to update the weight values, and this is repeated ten times. Finally, the test data of different scenarios is collected to confirm the effectiveness of the model accuracy.

[0062] In an embodiment of the present invention, the AI inference module 4 uses the weights of each layer of the trained sound classification deep network to update the input data and predict the result; specifically, the sound classification deep network extracts 512 features after passing through various convolution features from the 112 IQ points of the input data, and then through the feature weight update of the fully connected layer and the activation function for non-linear classification, and finally the inference result, occupied or unoccupied, will be obtained.

[0063] The present invention also discloses a method for recognizing human activity behaviors, Figure 2 is a flowchart of the method for recognizing human activity behaviors in an embodiment of the present invention; please refer to Figure 2 and the method for recognizing human activity behaviors includes:

[0064]

Step S1

[0065]

Step S2

[0066]

Step S3

[0067]

Step S4

[0068] In an embodiment of the present invention, in the feature extraction step, appropriate features can be extracted by cooperating with the sound classification deep network under the condition of collecting data on a large scale; the sound classification network will extract features for the sound in the frequency domain segment. Although the system in this case is not a sound signal, since the CSI is data in the frequency domain segment, the sound classification network is used as the deep learning network architecture.

[0069] The deep learning model adopts the sound classification deep network because the input feature of the sound classification model itself is the spectrum, and the CSI signal is also presented as signals in different frequency segments. Therefore, its AI model architecture is applied to this system;

[0070] In an embodiment of the present invention, in the AI training step, the IQ point data of 56 subcarriers of the CSI signal are used as the input of the model to train the model; the sum of the real number Q and the imaginary number I of the IQ point data is a total of 112 data.

[0071] In the AI training step, the environment with people and the environment without people are annotated; for the data of different time periods when a person stands between the development board and the router, it is manually marked as having people; if it is data of a time period without people, it is marked as without people; through marking, the AI model can effectively distinguish the input environmental signal features into the environmental states of having people or without people through real labels;

[0072] The input data for training includes 200,000 pieces of data with people and 200,000 pieces of data without people. The number of data with people is equivalent to the number of data without people, so as to balance the quantity of category data, and extract an appropriate proportion for environments of different scenarios, so that the data will not be biased towards a specific scenario;

[0073] During training, it will be put into the model in batches with a batch size of 400, and continuously find the solution that minimizes the error to update the weight values, and repeat the execution ten times. Finally, collect the test data of different scenarios to confirm the effectiveness of the model accuracy.

[0074] In an embodiment of the present invention, in the AI inference step, the weights of each layer of the trained sound classification deep network are used to update the input data and predict the result; specifically, the sound classification deep network extracts various convolutional features from 112 IQ points of the input data, and finally extracts a total of 512 features. Then, after the feature weights of the fully connected layer are updated and the activation function is used for non-linear classification, the inference result of whether there is someone or no one will be obtained.

[0075] In a usage scenario of the present invention, the IQ points are divided into real number Q and imaginary number I, and a set of complex points (I, Q) is represented by two variables; no matter how the IQ increases, it is a rotating circle on the two-dimensional plane (I, Q), which can avoid the problem of exceeding the boundary of (-3.14, 3.14) during mathematical operations such as phase angle correction. In addition, the rotation feature of the two-dimensional plane (I, Q) is used to replace the traditional method of correcting each phase angle into a phase feature without time influence. Even though the phase of the two-dimensional plane (I, Q) will shift due to time influence and cause phase angle deviation and error, the structure still has its stability. If the IQ points are viewed one by one from sub-carrier 1, sub-carrier 2... to sub-carrier 56, it will be found that the points rotate regularly according to different phases.

[0076] The AI training module is used to label the data set collected by the large-range data collection module for the environment with or without people, and use the characteristics of the sound classification deep network using the voice spectrogram to train the network. The individual IQ points of the 56 sub-carriers of CSI are a total of 56×2 = 112, which are used as the input of the model to train the model. The AI inference module is used to make predictions using 112 feature variables under each packet, and use one hundred inference values as the final prediction value.

[0077] In the large-range data collection module (Large Database Collection Module) of this system, a self-made data collector is used to collect data for a long time in the environment with people and without people, comprehensively considering the possible data changes at different times and locations.

[0078] In the CSI Feature Extraction Module, the IQ points are separated into the real part Q and the imaginary part I. A set of complex points (I, Q) is represented by two variables. No matter how the IQ values increase, they form a rotating circle on the two-dimensional plane (I, Q), which can avoid the problem of exceeding the boundary of (-3.14, 3.14) during mathematical operations such as phase angle correction. Additionally, the rotation feature of the two-dimensional plane (I, Q) is utilized to replace the traditional method of correcting each phase angle into a phase feature without the influence of time. Even though the phase of the two-dimensional plane (I, Q) may shift due to time and cause phase angle deviation errors, the structure still has its stability. If we look at the IQ points of each subcarrier from subcarrier 1, subcarrier 2... to subcarrier 56 one by one, we will find that the points rotate regularly according to different phases. With a large-scale data collection, appropriate features can be extracted by combining with a sound classification deep network.

[0079] The AI Training Module annotates the dataset of the presence or absence of people in the environment collected by the Large Database Collection Module, and uses the characteristics of the sound classification deep network using voice spectrograms to train the network. The individual IQ points of the 56 CSI subcarriers total 56X2 = 112, which are used as the input of the model to train the model. The AI Inference Module then makes predictions using 112 feature variables under each packet, and finally uses one hundred inference values to obtain the final prediction value.

[0080] This system collects 600,000 pieces of data covering both the presence and absence of people in the Large Database Collection Module for model training. In the CSI Feature Extraction Module, since the environments with people and without people are different, the environment with people has large variations and the environment without people has small variations. First, the Median Absolute Deviation (MAD) is used to remove outliers in each environment. Then, for the 56 subcarriers of each packet, by separating the IQ points into two variables, a total of 56X2 = 112 variables, namely variable I and variable Q, are used as the input data for the one-dimensional matrix deep learning model. Then, a sound classification deep network is used to extract the features on the spectrum.

[0081] After the AI Training Module extracts 512 features using the CSI Feature Extraction Module and the deep network for voice classification, it finally fits the labeled data for presence or absence of people, and after testing its effectiveness, it achieves the training goal. The AI Inference Module combines 100 values of a single packet as the final predicted probability value, making the prediction result less affected by the error value.

[0082] Figure 3 It is a schematic diagram for the Amplitude of Constellation Diagram. The left figure shows the original IQ points, and the IQ points rotate from T1, T2, T3 to T4 over time. The right figure is the schematic diagram of the amplitude calculated by the IQ points from the origin.

[0083] Figure 4 It is a schematic diagram for the Phase of Constellation Diagram. The left figure shows the original IQ points, and the right figure is the schematic diagram of the phase calculated by the angle between the IQ points and the I axis. It can be seen from the figure that the IQ points contain amplitude and phase information.

[0084] Figure 5 It is a working schematic diagram of the CSI Rotating Structure. The left figure is the original IQ point scatter plot, and the right figure is the structure diagram connecting each point over time. By observing the (I, Q) points from subcarrier 1 to subcarrier 56, it can be found that the (I, Q) points rotate according to a similar phase. Even with an estimated phase error, its stability can still be found in the structure.

[0085] Figure 6 It is a working schematic diagram of the Large Database Collection Module. It collects CSI signal data in parallel and stores it in the SD card memory, collecting CSI signals of people present and absent in the environment for a long time.

[0086] Figure 7It is a schematic diagram of the working process of the CSI Feature Extraction Module. Using the data of human presence and absence environment changes collected by the Large Database Collection Module, the Median Absolute Deviation (MAD) is used to filter out outliers in the human presence environment and outliers in the absence environment respectively. MAD is a statistical tool for calculating outliers. The median of the overall data is added and subtracted by three times MAD, that is, the average deviation of the value from the median, which is regarded as the reasonable value range, and the rest are outliers. Finally, the I and Q values of 56 subcarriers are split into a one-dimensional matrix variable of 1X112 as the input data size of the final model.

[0087] Figure 8 It is a schematic diagram of the working process of the AI Training Module. The one-dimensional matrix variable of 1X112 from the CSI Feature Extraction Module is input into the sound classification deep network, 1X512 features on the channel are extracted, and the human presence and absence environment is fitted. Finally, the model file is stored in the local folder.

[0088] Figure 9 It is a schematic diagram of the working process of the AI Inference Module. Using the 1X112-dimensional matrix IQ points of each packet as the model input value, and using the trained model to make inferences on a single packet. Finally, the mathematical operation - average is taken as the final prediction value.

[0089] The present invention also discloses an electronic device. Figure 10 It is a schematic diagram of the composition of the electronic device in an embodiment of the present invention; please refer to Figure 10 , at the hardware level, the electronic device includes a memory, a processor, and at least one network interface; the processor can be a microprocessor, and the memory can include an internal memory, such as a Random Access Memory (RAM), and can also include a non-volatile memory, etc. Of course, the electronic device can also be provided with other hardware according to needs.

[0090] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, etc.; the bus can include an address bus, a data bus, a control bus, etc. The memory is used to store programs (which can include an operating system program and application programs); the programs can include program codes, and the program codes can include computer operation instructions. The memory can include a memory and a non-volatile memory, and provide instructions and data to the processor.

[0091] In one embodiment, the processor can read the corresponding program from the non-volatile memory into the memory and then run it; the processor can execute the program stored in the memory and is specifically used to perform the following operations (as Figure 2 shown):

[0092]

Step S1

[0093]

Step S2

[0094]

Step S3

[0095]

Step S4

[0096] The present invention further discloses a storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the following steps of the method of the present invention are implemented (as Figure 2 shown):

[0097]

Step S1

[0098]

Step S2

[0099]

Step S3

[0100]

Step S4

[0101] In summary, for the human activity behavior recognition system, method, electronic device and storage medium proposed by the present invention, the present invention can utilize wireless signals to monitor human activities, ensuring the privacy of users; in addition, the entire system is simply built, and can be identified directly through a wifi machine and a receiving chip in the space, and overall the cost is extremely low.

[0102] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0103] The description and application of the present invention here are illustrative, and it is not intended to limit the scope of the present invention to the above embodiments. The effects or advantages involved in the embodiments may not be reflected in the embodiments due to various factors. The description of the effects or advantages is not used to limit the embodiments. The deformations and changes of the embodiments disclosed here are possible, and the substitutions and equivalent various components of the embodiments are well-known to those of ordinary skill in the art. Those skilled in the art should clearly understand that the present invention can be implemented in other forms, structures, arrangements, proportions, and with other components, materials and parts without departing from the spirit or essential characteristics of the present invention. Other deformations and changes can be made to the embodiments disclosed here without departing from the scope and spirit of the present invention.

Claims

1. A human activity behavior recognition system, characterized in that, The human activity behavior recognition system includes: A data collection module for collecting CSI signal data under a Wifi wireless network in both occupied and unoccupied environments to form a data set; A feature extraction module for extracting CSI signal features from the data collected by the data collection module. The CSI signal features are extracted from the points on the constellation diagram showing the relationship between the modulation signal distribution and the digital coding bits. The CSI signal features include IQ point data, and the IQ point data includes the real number I on the X-axis and the imaginary number Q on the Y-axis; An AI training module for annotating the data set collected by the data collection module in both occupied and unoccupied environments, and training the network using the characteristics of the sound classification deep network for model training through voice spectrograms; An AI inference module for making predictions using feature variables to form inference values, and using the inference values as the final prediction values to obtain human activity behaviors.

2. The human activity behavior recognition system according to claim 1, characterized in that: The feature extraction module includes a CSI signal feature extraction unit; the CSI signal feature extraction unit is used to extract CSI signal features from the data collected by the data collection module; The feature extraction module is also used to extract appropriate features in combination with the sound classification deep network under large-scale data collection; the sound classification network extracts features for the sound in the frequency domain segment, and CSI is the data in the frequency domain segment, and the sound classification network is used as the deep learning network architecture. The deep learning model uses the sound classification deep network because the input feature of the sound classification model itself is the spectrum, and the CSI signal is also presented as signals in different frequency segments, so its AI model architecture is applied to this system.

3. The human activity behavior recognition system according to claim 1, characterized in that: The AI training module uses the IQ point data of the m subcarriers of the CSI signal as the input of the model to train the model; the sum of the real number Q and the imaginary number I of the IQ point data is a total of 2m data. The AI training module is used to annotate the occupied environment and the unoccupied environment; for the data of different time periods when a person stands between the development board and the router, it is marked as occupied; if it is the data of the unoccupied time period, it is marked as unoccupied; through marking, the AI model can effectively distinguish the input environmental signal features into the occupied or unoccupied environmental states through real labels; The input data for training by the AI training module includes occupied data and unoccupied data, and the amounts of occupied data and unoccupied data are equivalent to balance the amount of category data, and appropriate proportions are extracted for different scenarios of the environment, so that the data will not be biased towards a specific scenario; During training, the data is put into the model in batches, and the scheme that minimizes the error is continuously found to update the weight values, and this is repeated several times. Finally, the test data of different scenarios are collected to confirm the effectiveness of the model accuracy.

4. The human activity behavior recognition system according to claim 1, characterized in that: The AI inference module uses the weights of each layer of the trained voice classification deep network to update the input data and predict the results. The details are as follows: The voice classification deep network extracts various convolution features from 2m IQ points of the input data, and finally extracts n features in total. Then, through the update of the feature weights of the fully connected layer and the activation function for non-linear classification, the inference result, whether there is a person or not, will be obtained finally.

5. A method for identifying human activity behaviors, characterized in that, The human activity behavior recognition method described above includes: A data collection step; collecting CSI signal data under the Wifi wireless network environment with and without people to form a data set. A feature extraction step; extracting CSI signal features from the data collected in the data collection step. The CSI signal features are extracted from the points on the constellation diagram showing the relationship between the modulation signal distribution and the digital coding bits. The CSI signal features include IQ point data, and the IQ point data includes the real number I on the X-axis and the imaginary number Q on the Y-axis. An AI training step; annotating the data set collected in the data collection step and training the network using the characteristics of the voice classification deep network to perform model training through the voice spectrogram. An AI inference step; using the feature variables to make predictions to form inference values, and using the inference values as the final prediction values to obtain the human activity behavior.

6. The human activity behavior recognition method according to claim 5, wherein: In the feature extraction step, appropriate features can be extracted by cooperating with the voice classification deep network under the condition of collecting data on a large scale. The voice classification network will extract features for the voice in the frequency domain segment. CSI is the data in the frequency domain segment, and the voice classification network is used as the deep learning network architecture. The deep learning model adopts the voice classification deep network because the input feature of the voice classification model itself is the spectrum, and the CSI signal is also presented as signals in different frequency segments. Therefore, its AI model architecture is applied to this system.

7. The human activity behavior recognition method according to claim 5, wherein: In the AI training step, the IQ point data of m subcarriers of the CSI signal is used as the input of the model to train the model; the sum of the real number Q and the imaginary number I of the IQ point data is a total of 2m data. In the AI training step, the environment with people and the environment without people are annotated. For the data of different time periods when a person stands between the development board and the router, it is marked as having a person; if it is the data of the time period without people, it is marked as without people. Through the marking, the AI model can effectively distinguish the input environmental signal features into the environmental states of having a person or not through the real labels. The input data for training includes data with people and data without people, and the number of data with people and data without people is equivalent to balance the number of category data, and appropriate proportions are extracted for different scenarios of the environment, so that the data will not be biased towards a specific scenario. During training, the data is put into the model in batches, and the scheme that minimizes the error is continuously found to update the weight values, and this is repeated several times. Finally, the test data of different scenarios are collected to confirm the effectiveness of the model accuracy.

8. The human activity behavior recognition method according to claim 5, wherein: In the AI inference step, the weights of each layer of the trained voice classification deep network are used to update the input data and predict the result; specifically, the voice classification deep network extracts various convolutional features from the m IQ points of the input data, and finally extracts a total of n features. Then, after the feature weights of the fully connected layer are updated and the activation function is used for non-linear classification, the inference result, whether there is a person or not, will be obtained.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 5 to 8.

10. A storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, it implements the steps of the method according to any one of claims 5 to 8.