Method and apparatus for counting number of people in area based on motion state, and medium
By extracting the two-dimensional effective peak features of radar echo data and training a random forest classifier, the problem of low accuracy caused by a single factor in radar people counting is solved, and high-accuracy people counting under different motion states is achieved.
Patent Information
- Application Number
- CN202111318801.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-09
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2041-11-09
AI Technical Summary
In existing technologies, radar-based regional population counting methods consider only one factor, resulting in low prediction accuracy and difficulty in adapting to the differences in echo signals and feature aliasing of targets under different motion states.
By extracting the two-dimensional effective peaks of the range-angle spectrum generated from radar echo data, and extracting the spacing features and energy block features, a random forest classifier is used to train a motion state and population classification model to determine the target motion state and the number of people in the area.
It improves the accuracy of regional population statistics, especially in environments with strong static clutter and multipath, it can accurately distinguish between 0 people, 1 person, and multiple people, which is better than traditional methods, with an average classification accuracy of 97.45%.
Smart Images

Figure CN114049604B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of area people counting, and particularly relate to an area people counting method and device based on motion state, equipment and medium. BACKGROUND
[0002] With the rapid development of the Internet of Things, applications based on area people information are becoming more and more widespread, and are becoming more and more important to daily life and work. Area people information can facilitate personnel diversion for relevant departments in intelligent transportation; in smart city management, it can help the government better grasp the dynamics of the crowd in the urban environment; in intelligent building heating, knowing the personnel distribution information in the building, the air conditioning system can intelligently adjust the appropriate temperature; in intelligent security, the security department can take different actions according to different area people. Area people information is essential. Area people counting has become a hot research topic.
[0003] In related technologies, the area people counting method based on radar considers relatively single and one-sided factors, and uses a relatively single prediction model, which may result in a relatively low accuracy of the result predicted by the model. It is very challenging and difficult to design a single optimal prediction model. SUMMARY
[0004] Embodiments of the present application provide an area people counting method and device based on motion state, equipment and medium, which can determine the motion state of the target in the area and the area people, and improve the accuracy of area people counting.
[0005] In a first aspect, embodiments of the present application provide an area people counting method based on motion state, which comprises: extracting a two-dimensional effective peak of a range-angle spectrum generated according to radar echo data in a preset area, and extracting a spacing feature and an energy block feature of the two-dimensional effective peak;
[0006] inputting the spacing feature and the energy block feature into a motion state classification model to determine a target motion state;
[0007] determining a target motion state people counting model according to the target motion state;
[0008] inputting the spacing feature and the energy block feature into the target motion state people counting model to obtain the number of people in the preset area.
[0009] In a second aspect, the embodiments of the present application also provide a device for counting the number of people in a region based on a motion state, which comprises: a feature extraction module, configured to extract a two-dimensional effective peak of a range-angle spectrum generated according to radar echo data in a preset region, and extract a spacing feature and an energy block feature of the two-dimensional effective peak;
[0010] a motion state determination module, configured to input the spacing feature and the energy block feature into a motion state classification model to determine a target motion state;
[0011] a motion state number classification model determination module, configured to determine a target motion state number classification model according to the target motion state;
[0012] a region number determination module, configured to input the spacing feature and the energy block feature into the target motion state number classification model to obtain the number of people in the preset region.
[0013] In a third aspect, the embodiments of the present application also provide an electronic device, which comprises:
[0014] one or more processors;
[0015] a storage device, configured to store one or more programs,
[0016] when the one or more programs are executed by the one or more processors, the one or more processors implement the method for counting the number of people in a region based on a motion state according to any one of the embodiments of the present application.
[0017] In a fourth aspect, the embodiments of the present application also provide a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement the method for counting the number of people in a region based on a motion state according to any one of the embodiments of the present application.
[0018] The technical scheme provided by the embodiments of the present application extracts a two-dimensional effective peak of a range-angle spectrum generated according to radar echo data in a preset region, and extracts a spacing feature and an energy block feature of the two-dimensional effective peak; inputs the spacing feature and the energy block feature into a motion state classification model to determine a target motion state; determines a target motion state number classification model according to the target motion state; and inputs the spacing feature and the energy block feature into the target motion state number classification model to obtain the number of people in the preset region. By executing the scheme, the motion state of a target in a region and the number of people in the region can be determined, and the accuracy of the counting of the number of people in the region can be improved. BRIEF DESCRIPTION OF DRAWINGS
[0019] Figure 1 is a flowchart of a method for counting the number of people in a region based on a motion state provided by the embodiments of the present application;
[0020] Figure 2 is a structure schematic diagram of a region people counting device based on a motion state provided by an embodiment of the present application;
[0021] Figure 3 is a structure schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0022] The present application will be further described below in conjunction with the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present application, but not to limit the present application. In addition, it should be noted that, for the convenience of description, only the parts related to the present application are shown in the drawings, but not all the structures.
[0023] Figure 1 is a flowchart of a region people counting method based on a motion state provided by an embodiment of the present application. The method can be executed by a region people counting device based on a motion state. The device can be realized by software and / or hardware. The device can be configured in an electronic device for region people counting. The method is applied to a scene of determining the number of people in a region. As shown in the flowchart, the technical scheme provided by the embodiment of the present application specifically includes the following steps. Figure 1
[0024] S110: extracting a two-dimensional effective peak of a range-angle spectrum generated according to radar echo data in a preset region, and extracting a spacing feature and an energy block feature of the two-dimensional effective peak.
[0025] Specifically, the radar can be a multiple-input multiple-output (MIMO) radar. The scheme can emit electromagnetic waves outward through a millimeter wave radio frequency module, and transmit the radar echo into an AD acquisition module after mixing to sample signals, so as to obtain Y k (n, m). Wherein, k represents a frame number dimension, is the kth frame echo data, m represents a slow time dimension, is the mth linear frequency modulation signal chirp, and n represents a fast time dimension, is the nth distance sampling unit.
[0026] The radar echo data is preprocessed, including generating a range-angle spectrum according to the radar echo data and clutter suppression. Specifically, the fast time dimension of each frame of collected radar echo is subjected to a fast Fourier transform (FFT) to generate a range image Y FFT , and then a distance unit of Y FFT is fixed, and a Capon spectrum is calculated. When the entire distance unit is traversed, a range-angle spectrum P k (l, θ). Wherein, θ is angle unit, θ ∈ [-60°, 60°], l is distance unit, and l ∈ [1, L], the value of L is equal to the number of set sampling points. Wherein, Capon spectrum is a kind of parameterized spatial spectrum, which can obtain three-dimensional spatial information of the target according to multiple digital target signals.
[0027] The clutter suppression is performed on the distance-angle spectrum P k (l, θ), and common clutter suppression methods include band-pass filtering, mean filtering, adaptive iterative filtering and the like. Here, adaptive iterative filtering is taken as an example for illustration, which can be represented by the following two steps:
[0028] D k (l, θ) = P k (l, θ) - C k (l, θ)
[0029] C k+1 (l, θ) = αC k (l, θ) + (1-α)P k (l, θ)
[0030] Wherein, D k (l, θ) represents filtered data after background subtraction, C k (l, θ) represents the clutter map of the current frame. 0 ≤ α ≤ 1 represents the update coefficient of the clutter map.
[0031] Two-dimensional effective peak extraction is performed on the clutter-suppressed D(n, θ), and the preset effective peak number N c and the size of the preset protection window l win and θ win are selected to be extracted. Specifically, N c = 12, l win = 2, and θ win = 4 can be determined according to experience. Then, the original filtered data can be recorded as D0(l, θ) to start the following steps: (1) the maximum value of D0(l, θ) is taken as the first maximum value, and the maximum value in the preset protection window centered on the first maximum value in D0(l, θ) is taken as the second maximum value, the first maximum value and the second maximum value are compared, if they are equal, the data point corresponding to the first maximum value is determined as the effective peak, and the amplitude and coordinates of the effective peak are saved, recorded as amplitude Z g , distance unit and angle unit (wherein, g ∈ [1, N cD (1, 0), and the target filtered data obtained by setting the data in the preset protection window centered on the first maximum value in D (1, 0) to 0 is recorded as D (1, 0). (2) The maximum value of D (1, 0) is recalculated as a third maximum value, and the maximum value in the preset protection window centered on the third maximum value is recalculated as a fourth maximum value, and then the third maximum value and the fourth maximum value are compared, if they are equal, the data point corresponding to the third maximum value can be determined as an effective peak, and can be recorded in the order of the effective peak as above, and then the data in the preset protection window centered on the third maximum value in D (1, 0) is set to 0 to update D (1, 0). (3) The step (2) is iterated until the preset number of effective peaks is extracted or the third maximum value is extracted again as 0, then the process of two-dimensional effective peak extraction is completed.
[0032] The distance feature and the energy block feature of the two-dimensional effective peak are extracted: the distance feature can reflect the relative position relationship between the person and the environment and the physical environment structure. The distance feature is specifically as follows:
[0033]
[0034] wherein, F1 represents the first distance feature, N c represents the preset number of effective peaks, represents the distance unit of the qth effective peak, represents the distance unit of the pth effective peak, represents the angle unit of the qth effective peak, represents the angle unit of the pth effective peak, and β represents a weight factor for balancing the distance and the angle;
[0035]
[0036] wherein, F2 represents the second distance feature, represents the distance unit of the maximum effective peak in the current frame, represents the angle unit of the maximum effective peak in the current frame;
[0037]
[0038] wherein, F3 represents the third distance feature, Z q represents the amplitude of the qth effective peak;
[0039]
[0040] wherein, F4 represents the fourth distance feature; and
[0041]
[0042] F5 represents the fifth distance feature. The distance feature can reflect the relative position relationship between the person and the environment and a certain physical environment structure.
[0043] The energy block feature of the two-dimensional effective peak is extracted, the preset area is divided into blocks according to the spatial size of the preset area and the maximum number of persons to be estimated, and the energy sum of the effective peak of each region after the division is counted respectively to determine the energy block feature. Specifically, due to the difference in the number of persons, the distribution of the effective peak energy is different, and according to the spatial size of the experimental environment (i.e. the preset area) and the maximum number of persons to be estimated, the experimental environment can be divided into N p blocks according to the distance and angle, and then the energy sum of the effective peak of N p blocks can be counted respectively. For example, N p may be 6, and the energy sum of the effective peak of 6 blocks can be obtained as the energy block feature of the two-dimensional effective peak, which can be recorded as [F6, F7, F8, F9, F 10 , F 11 ] T , further, the feature quantity used by the random forest classifier can be F = [F1, F2, …, F 11 ] T .
[0044] S120: input the distance feature and the energy block feature into a motion state classification model to determine a target motion state.
[0045] The motion state classification model can be trained and determined according to radar echo data of experimenters in an experimental area. For example, the experimental scene is a closed stainless steel metal gateway. The metal gateway is 2.4 meters long, 1.2 meters wide, and 2.4 meters high. The MIMO radar is installed at a height of 2.2 meters, and the angle between the main lobe direction and the horizon is designed to be about 50 degrees. The experiment is conducted by 3 male volunteers and 1 female volunteer. The data recording is recorded in batches according to different motion states of the experimenters. The radar echo data is divided into six types, namely, no one radar echo data (data-0p), one person static radar echo data (data-1p-s), one person motion radar echo data (data-1p-m), two people motion radar echo data (data-2p-mm), one person static one person motion radar echo data (data-2p-sm), and two people static radar echo data (data-2p-ss). The motion state can be set according to actual needs. For example, the motion state can be a background state, the motion state can be a static state, the motion state can be a confusion state, and the motion state can be a motion state. The correspondence between the radar echo data and the motion state is that the training data of the background state is composed of two-dimensional spatial attribute features calculated by data_0p, the training data of the static state is composed of two-dimensional spatial attribute features calculated by data_1p_s and data_2p_ss, the training data of the confusion state is composed of two-dimensional spatial attribute features calculated by data_1p_m and data_2p_sm, and the training data of the motion state is composed of two-dimensional spatial attribute features calculated by data_2p_mm. The two-dimensional spatial attribute features are distance features and energy block features. For each of the above six types of radar echo data, 20,000 frames of radar echo data can be recorded, of which 14,000 frames of radar echo data are used as training data, and the remaining 6,000 frames of radar echo data are used as test data. The distance features and the energy block features are extracted from the training data, and the obtained distance features and the energy block features are input into a random forest classifier for training. The motion state classification model is continuously optimized using test data, and the motion state classification model is finally determined.
[0046] In this embodiment, optionally, the motion state includes at least one of a background state, a static state, a confusion state, and a motion state; accordingly, the motion state person number classification model includes at least one of a background state person number classification model, a static state person number classification model, a confusion state person number classification model, and a motion state person number classification model.
[0047] The background state represents a case where no person is in the experimental area. The static state represents a case where one person is standing still or two persons are standing still simultaneously. The confusion state represents a case where one person is moving in the experimental area and two persons are in the experimental area, one of which is standing still and the other is moving. The moving state represents a case where two persons are moving simultaneously in the experimental area. The moving state number classification model is trained using a random forest classifier. The training data of the background state number classification model is composed of two-dimensional spatial attribute features calculated from data_0p and data_1p_s, and is used for classification of 0 persons and 1 person. The training data of the static state number classification model is composed of two-dimensional spatial attribute features calculated from data_1p_s and data_2p_ss, and is used for classification of 1 person and 2 persons. The training data of the confusion state number classification model is composed of two-dimensional spatial attribute features calculated from data_1p_m and data_2p_sm, and is used for classification of 1 person and 2 persons. The training data of the moving state number classification model is composed of two-dimensional spatial attribute features calculated from data_1p_m and data_2p_mm, and is used for classification of 1 person and 2 persons.
[0048] In this way, by classifying and determining the moving state and the moving state number classification model respectively, the number of persons in the preset area in the moving state can be determined, and the statistics of the number of persons in the area can be more accurate.
[0049] In a feasible implementation, optionally, the determination process of the moving state classification model includes: determining a first preset number of target radar echo data in a target area, and determining first feature data according to the target radar echo data; wherein the first feature data includes a distance feature and an energy block feature; the target radar echo data includes at least one of no-person radar echo data, one-person standing still radar echo data, one-person moving radar echo data, two-person moving radar echo data, one-person standing still and one-person moving radar echo data, and two-person standing still radar echo data; a first training data set and a first test data set are determined from the first feature data according to a first preset proportion; the first training data set and a moving state label associated with each feature data in the first training data set are input into a random forest classifier to obtain a to-be-trained moving state classification model and a first evaluation result; the accuracy of the to-be-trained moving state classification model is detected and optimized using the first test data set until the first evaluation result reaches a first preset threshold, and a trained moving state classification model is obtained.
[0050] The target region can be an experimental region for data statistics. The first preset number can be 2000, the first preset number can be 20000, and the first preset number can be set according to actual needs. The first feature data can be data obtained by extracting the interval feature and the energy block feature of the target radar echo data. The unmanned radar echo data represents radar echo data in an unmanned state in the experimental region. The one-person stationary radar echo data represents radar echo data in which one person is standing still in the experimental region. The one-person moving radar echo data represents radar echo data in which one person is moving (including small body shaking or walking, etc.) in the experimental region. The two-person moving radar echo data represents radar echo data in which two persons are moving (including small body shaking or walking, etc.) in the experimental region. The one-person stationary one-person moving radar echo data represents radar echo data in which one person is standing still and the other person is moving (including small body shaking or walking, etc.) in the experimental region. The two-person stationary radar echo data represents radar echo data in which two persons are standing still in the experimental region. The correspondence between the radar echo data and the motion state is as follows: the training data of the background state is composed of two-dimensional spatial attribute features calculated by data_0p, and the motion state label is 1. The training data of the stationary state is composed of two-dimensional spatial attribute features calculated by data_1p_s and data_2p_ss, and the motion state label is 2. The training data of the confusion state is composed of two-dimensional spatial attribute features calculated by data_1p_m and data_2p_sm, and the motion state label is 3. The training data of the motion state is composed of two-dimensional spatial attribute features calculated by data_2p_mm, and the motion state label is 4. The two-dimensional spatial attribute features are interval features and energy block features. The first preset ratio represents the proportion of the first training data set and the first test data set, for example, it can be 7:3, or it can also be 8:2. For example, for each of the above six kinds of radar echo data, 20000 frames of radar echo data can be recorded, and 14000 frames of radar echo data can be taken as the first training set, and the remaining 6000 frames of radar echo data can be taken as the first test set. The interval feature and the energy block feature are extracted from the first training set, and the obtained interval feature, energy block feature and motion state label matched with the interval feature and energy block feature are input into the random forest classifier for training to obtain a motion state classification model to be trained, and the accuracy of the motion state classification model to be trained is counted, that is, the first evaluation result. The first preset threshold can be 99%, the first preset threshold can also be 98%, and the first preset threshold can be set according to actual needs. The first test set is used to continuously optimize the motion state classification model until the first evaluation result of the motion state classification model reaches the first preset threshold, the training of the motion state classification model is stopped, and the motion state classification model is finally determined.
[0051] Thus, by training the random forest classifier using radar echo data in various motion states to determine a motion state classification model, the motion state of the target in the preset area can be determined, the problem that the area population statistical result in the prior art is biased towards a certain motion state can be solved, and thus the accuracy of area population statistics can be improved.
[0052] S130: Determine a target motion state population classification model according to the target motion state.
[0053] In this scheme, after determining the target motion state, the corresponding motion state population classification model is determined according to the target motion state. For example, if it is determined that the target motion state is the background state, the target motion state population classification model is determined to be the background state population classification model. If it is determined that the target motion state is the static state, the target motion state population classification model is determined to be the static state population classification model. If it is determined that the target motion state is the confusion state, the target motion state population classification model is determined to be the confusion state population classification model. If it is determined that the target motion state is the motion state, the target motion state population classification model is determined to be the motion state population classification model.
[0054] In another possible implementation, optionally, the determination process of the background state population classification model includes: determining a second preset amount of background state radar echo data, and determining second feature data according to the background state radar echo data; wherein the second feature data includes a distance feature and an energy block feature; the background state radar echo data includes no-person radar echo data and one-person static radar echo data; a second training data set and a second test data set are determined from the second feature data according to a second preset proportion; the second training data set and the area population associated with each feature data in the second training data set are input into the random forest classifier to obtain a background state population classification model to be trained and a second evaluation result; the accuracy of the background state population classification model to be trained is detected and optimized using the second test data set until the second evaluation result reaches a second preset threshold, and a trained background state population classification model is obtained.
[0055] The second preset quantity can be 2000, the second preset quantity can also be 20000, and the second preset quantity can be set according to actual needs. The second feature data can be data obtained by extracting the interval feature and the energy block feature of the background state radar echo data. The second preset ratio represents the proportion of the second training data set and the second test data set, and can be set according to actual needs. The second preset ratio can be 7:3, or can also be 8:2. For example, for the unmanned radar echo data and the one-person static radar echo data, 20000 frames of radar echo data can be recorded for each kind of data, and 14000 frames of radar echo data in each kind of data can be taken as the second training data, and the remaining 6000 frames of radar echo data can be taken as the second test data. The interval feature and the energy block feature are extracted from the second training set, and the interval feature, the energy block feature and the number of people in the region matched with the interval feature and the energy block feature are input into the random forest classifier for training to obtain the background state number of people classification model to be trained. For example, the number of people in the region matched with the feature data determined according to the unmanned radar echo data is 0, and the number of people in the region matched with the feature data determined according to the one-person static radar echo data is 1. The accuracy of the background state number of people classification model to be trained is counted, that is, the second evaluation result. The second preset threshold can be 99%, the second preset threshold can also be 98%, and the second preset threshold can be set according to actual needs. The background state number of people classification model is continuously optimized using the second test data until the second evaluation result of the background state number of people classification model reaches the second preset threshold, the training of the background state number of people classification model is stopped, and the background state number of people classification model is finally determined.
[0056] Therefore, by determining the background state number of people classification model, the number of people in the region in the background state can be determined, and the counting of the number of people in the region can be more accurate.
[0057] In another possible implementation, optionally, the determination process of the static state number of people classification model includes: determining a third preset quantity of static state radar echo data, and determining third feature data according to the static state radar echo data; wherein the third feature data includes an interval feature and an energy block feature; the static state radar echo data includes one-person static radar echo data and two-person static radar echo data; a third training data set and a third test data set are determined from the third feature data according to a third preset ratio; the third training data set and the number of people in the region associated with each feature data in the third training data set are input into a random forest classifier to obtain a static state number of people classification model to be trained and a third evaluation result; the accuracy of the static state number of people classification model to be trained is detected and optimized using the third test data set until the third evaluation result reaches a third preset threshold, and a trained static state number of people classification model is obtained.
[0058] The third preset quantity can be 2000, the third preset quantity can also be 20000, and the third preset quantity can be set according to actual needs. The third feature data can be data obtained by extracting the interval feature and the energy block feature from the static state radar echo data. The third preset ratio represents the proportion of the third training data set and the third test data set, and can be set according to actual needs. The third preset ratio can be 7:3, or can also be 8:2. For example, for one person static radar echo data and two person static radar echo data, 20000 frames of radar echo data can be recorded for each kind of data, 14000 frames of radar echo data are taken as the third training data, and the remaining 6000 frames of radar echo data are taken as the third test data. The interval feature and the energy block feature are extracted from the third training data, and the obtained interval feature, energy block feature and the number of people in the region matched with the interval feature and the energy block feature are input into the random forest classifier for training to obtain the static state number of people classification model to be trained. For example, the number of people in the region matched with the feature data determined according to the one person static radar echo data is 1, and the number of people in the region matched with the feature data determined according to the two person static radar echo data is 2. The accuracy of the static state number of people classification model to be trained is counted, that is, the third evaluation result. The third preset threshold can be 99%, the third preset threshold can also be 98%, and the third preset threshold can be set according to actual needs. The third test data is used to continuously optimize the static state number of people classification model until the third evaluation result of the static state number of people classification model reaches the third preset threshold, the training of the static state number of people classification model is stopped, and the static state number of people classification model is finally determined.
[0059] Therefore, by determining the static state number of people classification model, the number of people in the region in the static state can be determined, and the statistics of the number of people in the region can be more accurate.
[0060] In the embodiment, optionally, the determination process of the confusion state number of people classification model comprises: determining a fourth preset number of confusion state radar echo data, and determining fourth feature data according to the confusion state radar echo data; wherein the fourth feature data comprises a distance feature and an energy block feature; the confusion state radar echo data comprises one-person moving radar echo data and one-person static one-person moving radar echo data; a fourth training data set and a fourth test data set are determined from the fourth feature data according to a fourth preset proportion; the fourth training data set and the number of people in the region associated with each feature data in the fourth training data set are input into a random forest classifier to obtain a confusion state number of people classification model to be trained and a fourth evaluation result; the accuracy of the confusion state number of people classification model to be trained is continuously detected and optimized using the fourth test data set until the fourth evaluation result reaches a fourth preset threshold, and a trained confusion state number of people classification model is obtained.
[0061] The fourth preset number can be 2000, the fourth preset number can also be 20000, and the fourth preset number can be set according to actual needs. The fourth feature data can be data obtained by extracting the distance feature and the energy block feature from the confusion state radar echo data. The fourth preset proportion represents the proportion of the fourth training data set and the fourth test data set, and can be set according to actual needs. The fourth preset proportion can be 7:3, or can also be 8:2. For example, for one-person moving radar echo data and one-person static one-person moving radar echo data, 20000 frames of radar echo data can be recorded for each type of data, and 14000 frames of radar echo data are taken as the fourth training data, and the remaining 6000 frames of radar echo data are taken as the fourth test data. The distance feature and the energy block feature are extracted from the fourth training data, and the distance feature, the energy block feature and the number of people in the region matched with the distance feature and the energy block feature are input into a random forest classifier for training to obtain a confusion state number of people classification model to be trained. For example, the number of people in the region matched with the feature data determined according to the one-person moving radar echo data is 1, and the number of people in the region matched with the feature data determined according to the one-person static one-person moving radar echo data is 2. The accuracy of the confusion state number of people classification model to be trained, i.e. the fourth evaluation result, is calculated. The fourth preset threshold can be 99%, the fourth preset threshold can also be 98%, and the fourth preset threshold can be set according to actual needs. The confusion state number of people classification model is continuously optimized using the fourth test data until the fourth evaluation result of the confusion state number of people classification model reaches the fourth preset threshold, the training of the confusion state number of people classification model is stopped, and the confusion state number of people classification model is finally determined.
[0062] Therefore, by determining the confused state person number classification model, the number of people in the confused state can be determined, and the statistics of the number of people in the region can be more accurate.
[0063] In this embodiment, optionally, the determination process of the motion state person number classification model comprises: determining a fifth preset number of motion state radar echo data, and determining fifth feature data according to the motion state radar echo data; wherein the fifth feature data comprises a distance feature and an energy block feature; the motion state radar echo data comprises one-person motion radar echo data and two-person motion radar echo data; a fifth training data set and a fifth test data set are determined from the fifth feature data according to a fifth preset ratio; the fifth training data set and the number of people in the region associated with each feature data in the fifth training data set are input into the random forest classifier to obtain a motion state person number classification model to be trained and a fifth evaluation result; the accuracy of the motion state person number classification model to be trained is continuously detected and optimized using the fifth test data set until the fifth evaluation result reaches a fifth preset threshold, and a trained motion state person number classification model is obtained.
[0064] The fifth preset number can be 2000, the fifth preset number can also be 20000, and the fifth preset number can be set according to actual needs. The fifth feature data can be data obtained by extracting the distance feature and the energy block feature from the motion state radar echo data. The fifth preset ratio represents the proportion of the fifth training data set and the fifth test data set, which can be set according to actual needs. The fifth preset ratio can be 7:3, for example, or 8:2. For example, for one-person motion radar echo data and two-person motion radar echo data, 20000 frames of radar echo data can be recorded for each type of data, and 14000 frames of radar echo data can be used as the fifth training data, and the remaining 6000 frames of radar echo data can be used as the fifth test data. The distance feature and the energy block feature are extracted from the fifth training data, and the distance feature, the energy block feature, and the number of people in the region matched with the distance feature and the energy block feature are input into the random forest classifier for training to obtain a motion state person number classification model to be trained. For example, the number of people in the region matched with the feature data determined according to the one-person motion radar echo data is 1, and the number of people in the region matched with the feature data determined according to the two-person motion radar echo data is 2. The accuracy of the motion state person number classification model to be trained, i.e. the fifth evaluation result, is calculated. The fifth preset threshold can be 99%, the fifth preset threshold can also be 98%, and the fifth preset threshold can be set according to actual needs. The motion state person number classification model is continuously optimized using the fifth test data until the fifth evaluation result of the motion state person number classification model reaches the fifth preset threshold, the training of the motion state person number classification model is stopped, and the motion state person number classification model is finally determined.
[0065] Therefore, by determining the motion state person classification model, the number of people in the region in the motion state can be determined, and the statistics of the number of people in the region can be more accurate.
[0066] S140: inputting the interval feature and the energy block feature into the target motion state person classification model to obtain the number of people in the preset region.
[0067] Specifically, after the motion state classification is completed, the corresponding motion state person classification model is cascaded for person classification. The random forest classifier is selected for the model, and the feature quantity used is the interval feature and the energy block feature of the two-dimensional effective peak determined in S110, that is, F = [F1, F2, …, F 11 ] T The interval feature and the energy block feature are input into the determined target motion state person classification model to obtain the number of people in the preset region. For example, the number of people in the preset region is 0, 1 or more.
[0068] On the basis of the above technical solutions, the area person counting method provided is tested. Taking a 77GHz millimeter wave radar (with a bandwidth of 4GHz) as an example, the size of the stainless steel ramp is 2.4m x 1.2m x 2.4m, the installation height is 2.2m, and the method is tested for 16.67 minutes (20000 frames of data, frame rate is 20Hz) for 0 people, 1 person and multiple people. The results are shown in the following tables. As can be seen from Table 1, the average accuracy of motion state classification is 97.69%. As can be seen from Table 2, the person classification accuracy of the four motion state prediction models is more than 95%. As can be seen from Table 3, the average accuracy of person classification is 97.45%. In the strong static clutter and strong multipath metal gate, 0 people, 1 person and multiple people can be accurately distinguished, and the traditional access control system such as card swiping, fingerprint and face recognition can be combined to realize the anti-tailing function. The experimental results show that the average classification accuracy of 0 people, 1 person and multiple people is more than 97.45%, which is better than the existing area person counting method.
[0069] Table 1
[0070]
[0071] Table 2
[0072] Model Accuracy Background state model 98.46% Stationary state model 99.68% Confusion state model 95.39% Motion state model 97.45%
[0073] Table 3
[0074]
[0075] In the related art, the radar-based area people counting method rarely considers the motion state of the target and uses a single prediction model to classify the number of people. Because the echo signals of the target in different motion states are quite different, feature aliasing sometimes occurs between adjacent people. Therefore, it is difficult to find an optimal single prediction model to adapt to different motion states of the target.
[0076] The technical scheme provided by the embodiment of the present application extracts the two-dimensional effective peaks of the range-angle spectrum generated according to the radar echo data in the preset area, and extracts the interval feature and energy block feature of the two-dimensional effective peaks; inputs the interval feature and the energy block feature into a motion state classification model to determine the motion state of the target; determines a target motion state people counting model according to the motion state of the target; and inputs the interval feature and the energy block feature into the target motion state people counting model to obtain the number of people in the preset area. By executing the scheme, the motion state of the target in the area and the number of people in the area can be determined, and the accuracy of area people counting can be improved.
[0077] Figure 2 is a structure diagram of an area people counting device based on a motion state provided by the embodiment of the present application. The device can be configured in an electronic device for area people counting, as shown in Figure 2 The device includes:
[0078] The feature extraction module 210 is configured to extract the two-dimensional effective peaks of the range-angle spectrum generated according to the radar echo data in the preset area, and extract the interval feature and energy block feature of the two-dimensional effective peaks.
[0079] The motion state determination module 220 is configured to input the interval feature and the energy block feature into a motion state classification model to determine the motion state of the target.
[0080] The motion state people counting model determination module 230 is configured to determine a target motion state people counting model according to the motion state of the target.
[0081] The area people determination module 240 is configured to input the interval feature and the energy block feature into the target motion state people counting model to obtain the number of people in the preset area.
[0082] Optionally, the motion state includes at least one of a background state, a stationary state, a confusion state, and a motion state; and correspondingly, the motion state people counting model includes at least one of a background state people counting model, a stationary state people counting model, a confusion state people counting model, and a motion state people counting model.
[0083] Optionally, the determination process of the motion state classification model comprises: determining a first preset number of target radar echo data in the target area, and determining first feature data according to the target radar echo data; wherein the first feature data comprises a distance feature and an energy block feature; the target radar echo data comprises at least one of no-person radar echo data, one-person static radar echo data, one-person motion radar echo data, two-person motion radar echo data, one-person static one-person motion radar echo data, and two-person static radar echo data; a first training data set and a first test data set are determined from the first feature data according to a first preset proportion; the first training data set and a motion state label associated with each feature data in the first training data set are input into a random forest classifier to obtain a motion state classification model to be trained and a first evaluation result; the accuracy of the motion state classification model to be trained is detected and optimized using the first test data set until the first evaluation result reaches a first preset threshold, and a trained motion state classification model is obtained.
[0084] Optionally, the determination process of the background state person number classification model comprises: determining a second preset number of background state radar echo data, and determining second feature data according to the background state radar echo data; wherein the second feature data comprises a distance feature and an energy block feature; the background state radar echo data comprises no-person radar echo data and one-person static radar echo data; a second training data set and a second test data set are determined from the second feature data according to a second preset proportion; the second training data set and a region person number associated with each feature data in the second training data set are input into a random forest classifier to obtain a background state person number classification model to be trained and a second evaluation result; the accuracy of the background state person number classification model to be trained is detected and optimized using the second test data set until the second evaluation result reaches a second preset threshold, and a trained background state person number classification model is obtained.
[0085] Optionally, the determination process of the static state people number classification model comprises: determining a third preset quantity of static state radar echo data, and determining third feature data according to the static state radar echo data; wherein the third feature data comprises a distance feature and an energy block feature; the static state radar echo data comprises one-person static state radar echo data and two-person static state radar echo data; determining a third training data set and a third test data set from the third feature data according to a third preset proportion; inputting the third training data set and the area number of people associated with each feature data in the third training data set into a random forest classifier to obtain a static state people number classification model to be trained and a third evaluation result; and continuing to detect and optimize the accuracy of the static state people number classification model to be trained by using the third test data set until the third evaluation result reaches a third preset threshold, thereby obtaining a trained static state people number classification model.
[0086] Optionally, the determination process of the confusion state people number classification model comprises: determining a fourth preset quantity of confusion state radar echo data, and determining fourth feature data according to the confusion state radar echo data; wherein the fourth feature data comprises a distance feature and an energy block feature; the confusion state radar echo data comprises one-person motion radar echo data and one-person static one-person motion radar echo data; determining a fourth training data set and a fourth test data set from the fourth feature data according to a fourth preset proportion; inputting the fourth training data set and the area number of people associated with each feature data in the fourth training data set into a random forest classifier to obtain a confusion state people number classification model to be trained and a fourth evaluation result; and continuing to detect and optimize the accuracy of the confusion state people number classification model to be trained by using the fourth test data set until the fourth evaluation result reaches a fourth preset threshold, thereby obtaining a trained confusion state people number classification model.
[0087] Optionally, the determination process of the motion state people number classification model comprises: determining a fifth preset quantity of motion state radar echo data, and determining fifth feature data according to the motion state radar echo data; wherein the fifth feature data comprises a distance feature and an energy block feature; the motion state radar echo data comprises one-person motion radar echo data and two-person motion radar echo data; determining a fifth training data set and a fifth test data set from the fifth feature data according to a fifth preset proportion; inputting the fifth training data set and the area number of people associated with each feature data in the fifth training data set into a random forest classifier to obtain a motion state people number classification model to be trained and a fifth evaluation result; and continuing to detect and optimize the accuracy of the motion state people number classification model to be trained by using the fifth test data set until the fifth evaluation result reaches a fifth preset threshold, thereby obtaining a trained motion state people number classification model.
[0088] The device provided in the above embodiment can execute the motion state-based area people counting method provided in any embodiment of the application, and has the corresponding function modules and advantages of the execution method.
[0089] Figure 3 is a schematic diagram of an electronic device structure provided in an embodiment of the application, as shown in the figure, the device comprises: Figure 3
[0090] one or more processors 310, Figure 3 in which the processor 310 is taken as an example;
[0091] a memory 320;
[0092] The device can also include an input device 330 and an output device 340.
[0093] The processor 310, the memory 320, the input device 330 and the output device 340 in the device can be connected through a bus or other means, Figure 3 in which the connection through the bus is taken as an example.
[0094] The memory 320, as a kind of non-transient computer readable storage medium, can be used to store software programs, computer executable programs and modules, such as the program instructions / modules of a motion state-based area people counting method in an embodiment of the application. The processor 310 executes the software programs, instructions and modules stored in the memory 320, thereby executing various function applications and data processing of the computer device, i.e. implementing a motion state-based area people counting method in the above method embodiment, i.e.
[0095] extracting a two-dimensional effective peak of a range-angle spectrum generated according to radar echo data in a preset area, and extracting a spacing feature and an energy block feature of the two-dimensional effective peak;
[0096] inputting the spacing feature and the energy block feature into a motion state classification model to determine a target motion state;
[0097] determining a target motion state people counting model according to the target motion state;
[0098] inputting the spacing feature and the energy block feature into the target motion state people counting model to obtain the number of people in the preset area.
[0099] The memory 320 can include a program storage area and a data storage area, where the program storage area can store an operating system, at least one application required by a function, and the data storage area can store data created according to the use of the computer device, etc. In addition, the memory 320 can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory 320 can optionally include a memory disposed remotely relative to the processor 310, which can be connected to the terminal device through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0100] The input device 330 can be used to receive input digital or character information, and to generate key signal inputs related to the user settings and function control of the computer device. The output device 340 can include a display device such as a display screen.
[0101] The embodiment of the present application provides a computer readable storage medium, which stores a computer program, and the program is executed by a processor to realize a region people counting method based on a motion state, i.e.:
[0102] Extract a two-dimensional effective peak of a range-angle spectrum generated according to radar echo data in a preset region, and extract a spacing feature and an energy block feature of the two-dimensional effective peak;
[0103] Input the spacing feature and the energy block feature into a motion state classification model to determine a target motion state;
[0104] Determine a target motion state people counting model according to the target motion state;
[0105] Input the spacing feature and the energy block feature into the target motion state people counting model to obtain a number of people in the preset region.
[0106] Any combination of one or more computer readable medium can be utilized. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. A computer readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer readable storage medium can be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0107] A computer readable signal medium can include a propagated data signal with computer readable program code embodied therein, for example, in baseband or as part of a carrier wave. Such a propagated signal can take any of a variety of forms, including, but not limited to, electro-magnetic, optical, or any suitable combination thereof. A computer readable signal medium can be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0108] Program code embodied on a computer readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0109] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In an embodiment of the application, the remote computer can be a server or another desktop computer.
[0110] Note that the above merely describes preferred embodiments of the present application and the principles of the technology applied. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and that various obvious changes, modifications and substitutions can be made without departing from the scope of the present application. Therefore, although the present application has been described in detail through the above embodiments, the present application is not limited to the above embodiments, and can include more other equivalent embodiments without departing from the concept of the present application, and the scope of the present application is determined by the scope of the claims.
Claims
1. A method for regional population counting based on motion status, characterized in that, include: The effective peaks of the two-dimensional range-angle spectrum generated based on radar echo data within a preset area are extracted, and the spacing characteristics and energy block characteristics of the two-dimensional effective peaks are extracted. The spacing features and the energy block features are input into the motion state classification model to determine the target motion state; Determine the target motion state and the number of people in the target motion state classification model; The distance feature and the energy block feature are input into the target motion state number classification model to obtain the number of people in the preset area; The motion state includes at least one of the following: background state, static state, confused state, and motion state; Accordingly, the motion state people classification model includes at least one of the following: background state people classification model, stationary state people classification model, confused state people classification model, and motion state people classification model; The process of determining the confusion state population classification model includes: A fourth preset number of confused state radar echo data is determined, and a fourth feature data is determined based on the confused state radar echo data; wherein, the fourth feature data includes spacing features and energy block features; the confused state radar echo data includes radar echo data of one person moving and radar echo data of one person stationary and one person moving. The fourth training dataset and the fourth test dataset are determined from the fourth feature data according to the fourth preset ratio; The fourth training dataset and the number of people in the region associated with each feature data in the fourth training dataset are input into a random forest classifier to obtain a confused state population classification model to be trained and a fourth evaluation result. The accuracy of the confused state people classification model to be trained is further detected and optimized using the fourth test dataset until the fourth evaluation result reaches the fourth preset threshold, thus obtaining the trained confused state people classification model.
2. The method according to claim 1, characterized in that, The process of determining the motion state classification model includes: A first preset number of target radar echo data within the target area are determined, and first feature data is determined based on the target radar echo data; wherein, the first feature data includes spacing features and energy block features; the target radar echo data includes at least one of unmanned radar echo data, one-person stationary radar echo data, one-person moving radar echo data, two-person moving radar echo data, one-person stationary and one-person moving radar echo data, and two-person stationary radar echo data. A first training dataset and a first test dataset are determined from the first feature data according to a first preset ratio; The first training dataset and the motion state labels associated with each feature data in the first training dataset are input into a random forest classifier to obtain the motion state classification model to be trained and the first evaluation result. The accuracy of the motion state classification model to be trained is further detected and optimized using the first test dataset until the first evaluation result reaches the first preset threshold, thus obtaining the trained motion state classification model.
3. The method according to claim 2, characterized in that, The process of determining the background state population classification model includes: A second preset number of background state radar echo data is determined, and second feature data is determined based on the background state radar echo data; wherein, the second feature data includes spacing features and energy block features; the background state radar echo data includes unmanned radar echo data and single-person stationary radar echo data; The second training dataset and the second test dataset are determined from the second feature data according to the second preset ratio; The second training dataset and the number of people in the region associated with each feature data in the second training dataset are input into a random forest classifier to obtain a background state population classification model to be trained and a second evaluation result. The accuracy of the background state people classification model to be trained is further tested and optimized using the second test dataset until the second evaluation result reaches the second preset threshold, thus obtaining the trained background state people classification model.
4. The method according to claim 2, characterized in that, The process of determining the static population classification model includes: A third preset number of stationary radar echo data is determined, and a third characteristic data is determined based on the stationary radar echo data; wherein, the third characteristic data includes spacing characteristics and energy block characteristics; the stationary radar echo data includes stationary radar echo data of one person and stationary radar echo data of two people; The third training dataset and the third test dataset are determined from the third feature data according to the third preset ratio; The third training dataset and the number of people in the region associated with each feature data in the third training dataset are input into a random forest classifier to obtain a static population classification model to be trained and a third evaluation result. The accuracy of the stationary people classification model to be trained is further tested and optimized using the third test dataset until the third evaluation result reaches the third preset threshold, thus obtaining the trained stationary people classification model.
5. The method according to claim 2, characterized in that, The process of determining the classification model for moving populations includes: A fifth preset number of moving radar echo data is determined, and a fifth feature data is determined based on the moving radar echo data; wherein, the fifth feature data includes spacing features and energy block features; the moving radar echo data includes radar echo data of one person moving and radar echo data of two people moving; The fifth training dataset and the fifth test dataset are determined from the fifth feature data according to the fifth preset ratio; The fifth training dataset and the number of people in the region associated with each feature data in the fifth training dataset are input into a random forest classifier to obtain the moving-state number classification model to be trained and the fifth evaluation result; The accuracy of the motion-based people classification model to be trained is further tested and optimized using the fifth test dataset until the fifth evaluation result reaches the fifth preset threshold, thus obtaining the trained motion-based people classification model.
6. A device for counting the number of people in a region based on motion status, characterized in that, include: The feature extraction module is used to extract the two-dimensional effective peaks of the range-angle spectrum generated based on radar echo data within a preset area, and to extract the spacing features and energy block features of the two-dimensional effective peaks. A motion state determination module is used to input the spacing features and the energy block features into a motion state classification model to determine the target motion state. The motion state number of people classification model determination module is used to determine the target motion state number of people classification model based on the target motion state; The area population determination module is used to input the spacing features and the energy block features into the target motion state population classification model to obtain the population within the preset area; The motion state includes at least one of background state, stationary state, confused state, and motion state; correspondingly, the motion state people classification model includes at least one of background state people classification model, stationary state people classification model, confused state people classification model, and motion state people classification model. The process of determining the confused state people classification model includes: determining a fourth preset number of confused state radar echo data, and determining fourth feature data based on the confused state radar echo data; wherein, the fourth feature data includes spacing features and energy block features; the confused state radar echo data includes radar echo data of one person moving and radar echo data of one person stationary and one person moving; determining a fourth training dataset and a fourth test dataset from the fourth feature data according to a fourth preset ratio; inputting the fourth training dataset and the number of people in the region associated with each feature data in the fourth training dataset into a random forest classifier to obtain the confused state people classification model to be trained and the fourth evaluation result; using the fourth test dataset to continue to detect and optimize the accuracy of the confused state people classification model to be trained until the fourth evaluation result reaches the fourth preset threshold, thus obtaining the trained confused state people classification model.
7. An electronic device, characterized in that, include: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the motion-based regional population counting method as described in any one of claims 1-5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the regional population counting method based on motion status as described in any one of claims 1-5.
Citation Information
Patent Citations
Regional people counting method and device, computer equipment and storage medium
CN113311405A
Radar-based people counting method and device, equipment and storage medium
CN113313165A