A method for constructing optimal subcarrier combination based on LSTM model

Through the optimal subcarrier combination construction method based on the LSTM model, the problem of high key repetition ratio in the prior art is solved, the effect of reducing subcarrier correlation and improving the key update rate is achieved, and communication security is ensured.

CN119182528BActive Publication Date: 2025-05-16JIANGSU JINWANG TESTING CERTIFICATION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411641814.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-18
Publication Date
2025-05-16
Estimated Expiration
2044-11-18

AI Technical Summary

Technical Problem

In the prior art, the CSI-based key generation method has a high key duplication ratio due to the correlation between subcarriers, which is easily cracked by eavesdroppers, resulting in extremely poor key availability.

Method used

The optimal subcarrier combination construction method based on the LSTM model is adopted, and the data set is formed by capturing the original CSI data, pre-processing it, and the optimal cluster number is determined using the improved elbow rule and the K-means algorithm, and the data label is generated, and the optimal subcarrier combination is output through the LSTM subcarrier classification model.

Benefits of technology

Effectively reduce the correlation between subcarriers, significantly reduce the key duplication ratio, improve the key update rate, and ensure communication security, especially when communication interactions are frequent.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119182528B_ABST
    Figure CN119182528B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for constructing an optimal subcarrier combination based on an LSTM model, including capturing CSI raw data and preprocessing to form an original data set containing a plurality of data packets; processing the original data set using an improved elbow rule based on a cost function, determining the value range of the clustering number M, and determining the optimal clustering number M' from this range. The present invention can effectively reduce the correlation between subcarriers, so that the key repetition ratio is significantly reduced. In theory, there will be no subcarrier sampling sequence with high similarity, which can effectively prevent a large number of repeated bits from appearing in the key sequence, improve the key update rate, and effectively ensure communication security, especially in the case of frequent communication interactions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of wireless channel keys, and in particular to a method for constructing an optimal subcarrier combination based on an LSTM model. Background Art

[0002] As a supplement to existing encryption technologies, physical layer key generation technology has received extensive attention. Physical layer key generation uses the inherent randomness of the radio channel to provide dynamic shared keys. Based on the physical observation properties of the time-varying multiple-input multiple-output (MIMO) channel, time- and space-related channel samples are used to establish keys. When legitimate communication nodes conduct wireless communication, the channel characteristics of the communication parties have channel reciprocity, uniqueness in time and space, rapid time variation, and randomness. Among the key generation technologies that analyze channel characteristics, key generation technology based on channel state information (CSI) has broad application prospects and research value. CSI provides multi-dimensional and refined information about the transmission characteristics of wireless channels, including parameters such as signal amplitude, phase, and frequency. Slight changes in these parameters can be used to generate keys, making it difficult for attackers to crack the keys through simple measurements or predictions. Compared with simple signal strength indicators such as RSSI, CSI describes the channel more comprehensively and accurately, contains rich channel state information, and can remain relatively stable even in the presence of environmental changes or interference.

[0003] At present, common communication systems all use OFDM modulation. OFDM uses subcarriers of different frequencies to send information, and the CFR of different subcarriers can be obtained. CSI is actually the discrete sampling of CFR on different subcarriers, that is, CFR is a form of CSI. The key generation method based on traditional CSI is to use the CFR of multiple subcarriers measured simultaneously in a data packet as a shared random source, and the reciprocity is good, which means that the CFR of a single subcarrier corresponding to the communicating party has a similar value. Based on the CSI-based different-numbered subcarrier quotient model, the quotient of the sampling sequence of different subcarriers is used as a random source. Compared with the traditional key generation method that uses CSI as a random source for keys in the frequency domain dimension, it achieves the effect of improving CSI information utilization and key generation rate.

[0004] Although the subcarrier signal frequencies in the OFDM channel are orthogonal, adjacent subcarriers have very close frequencies, resulting in similar channel responses in the frequency domain, and the CSI measured from them may have a strong correlation. Therefore, in the different-numbered subcarrier quotient model, the continuous groups of keys generated by quantizing the quotient sequence of the CSI measured by the same subcarrier and several adjacent subcarriers will produce a large ratio of repeated segments, which can be easily cracked by eavesdroppers, resulting in extremely poor key availability. Summary of the invention

[0005] The purpose of the present invention is to provide a method for constructing an optimal subcarrier combination based on an LSTM model to solve the technical problems raised in the background technology.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] S1. Capture CSI raw data and preprocess it to form a raw data set containing several data packets. ,in, represents the data packet of the i-th original data, , the data packet contains several valid subcarrier CSI sampling information;

[0008] S2, using the improved elbow rule based on the cost function to process the original data set, determine the value range of the cluster number M, and determine the optimal cluster number M' from this range;

[0009] S3. When the number of classifications is the optimal number of clusters M', the subcarrier CSI sampling information is clustered and data labels are generated. The subcarrier CSI sampling information in each data packet contained in the original data set is labeled using the data labels to form a labeled data set. ,in, represents the i-th data packet with data label, ;

[0010] S4. Labeled dataset The model is divided into a training set and a test set. After the training set is input into the LSTM subcarrier classification model for training, the trained LSTM subcarrier classification model is tested with the test set. Finally, in each category output by the LSTM subcarrier classification model, any subcarrier corresponding to the category is selected to form the optimal subcarrier combination. The structure of the LSTM subcarrier classification model is as follows from input to output: LSTM layer, fully connected layer and output layer.

[0011] Furthermore, the specific steps of step S2 are:

[0012] S21, select several CSI raw data packets from the raw data set, and use the elbow rule to determine the value range U of the clustering number M;

[0013] S22. Construct a cost function, calculate the function values ​​corresponding to all cluster numbers in U through the cost function, and take the cluster number corresponding to the minimum function value as the optimal cluster number M'.

[0014] Furthermore, the specific steps of determining the value range U of the number of clusters by the elbow rule in step S21 are:

[0015] S211, using the number T of subcarrier CSI sampling information contained in the selected several CSI original data packets as the value range of the classification number K of the K-means algorithm, which is expressed as K∈[1,T], and using the K-means algorithm to perform multiple subcarrier clustering within this range to generate multiple clustering results, wherein the clustering results are within the value range of the classification number K;

[0016] S212, calculate the intra-cluster sum of squares based on each clustering result , we get the value range U of the number of clusters M, where U∈K, then The calculation formula is:

[0017] ,

[0018] In the above formula, represents the cluster center of the i-th class, P represents the data points belonging to the i-th class, K represents the number of classifications, represents the i-th cluster.

[0019] Furthermore, the cost function in step S22 is expressed by the following formula:

[0020] ,

[0021] In the above formula, F represents the cost function value, D represents the intra-class distance, and L represents the inter-class distance. represents the mean value of the subcarrier CSI sampling information in the kth class, represents the number of subcarrier CSI sampling information in the kth class, represents the mean of all sample data, represents the subcarrier CSI sampling information of the kth class, represents the CSI sampling information of the jth subcarrier in the i-th class, , represents the kth cluster, and M represents the number of clusters.

[0022] Furthermore, the output layer of the LSTM subcarrier classification model adopts The activation function is The activation function outputs the predicted probability of each category , predicted probability The calculation formula is:

[0023] ,

[0024] In the above formula, is the predicted probability of the subcarrier class, and b represent the weight matrix and bias respectively, and Y represents the output data after the fully connected layer maps the output of the LSTM layer to the classification space.

[0025] Furthermore, the data packet with data tags includes several valid data tags and their corresponding subcarrier CSI sampling information.

[0026] Furthermore, the training set accounts for 80% of the labeled data set, and the test set accounts for 20% of the labeled data set.

[0027] Beneficial effects:

[0028] 1. The present invention processes the original data set based on the improved elbow rule of the cost function and the K-means algorithm, determines the value range of the cluster number, and determines the optimal cluster number from this range. When the classification number is the optimal cluster number, the subcarrier CSI sampling information is subcarrier clustered and data labels are generated by the K-means algorithm. The LSTM subcarrier classification model is used to output the optimal subcarrier combination, which can effectively reduce the correlation between subcarriers, significantly reduce the key repetition ratio, and theoretically there will be no subcarrier sampling sequence with high similarity, which can effectively prevent a large number of repeated bits from appearing in the key sequence, and improve the key update rate, especially in the case of frequent communication interactions, which can effectively ensure communication security.

[0029] 2. The present invention determines the value range of the number of clusters through the elbow rule, and then determines the optimal number of clusters through the cost function. While maintaining the effectiveness of clustering, the selected data can be optimally classified, which is conducive to the subsequent LSTM subcarrier classification model outputting the optimal subcarrier combination. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions implemented in the present invention, the accompanying drawings required for describing the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0031] Figure 1 A flowchart of a method for constructing an optimal subcarrier combination according to the present invention;

[0032] Figure 2 The calculation result of the elbow rule of the present invention;

[0033] Figure 3 The calculation result of the cost function of the present invention;

[0034] Figure 4 Data tags corresponding to different subcarriers of the present invention;

[0035] Figure 5 : The subcarrier CSI amplitude waveform under different data labels of the present invention;

[0036] Figure 6 This is a structural diagram of the LSTM subcarrier classification model based on the present invention;

[0037] Figure 7 It is the training result of the LSTM subcarrier classification model based on the present invention;

[0038] Figure 8 The test results of the LSTM subcarrier classification model test set of the present invention are as follows;

[0039] Fig. 9 The optimal subcarrier combination data table of the present invention;

[0040] Fig.10 The figure is a comparison of the CSI quotient waveforms of the optimal subcarrier of the present invention. DETAILED DESCRIPTION

[0041] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0042] like Figure 1-Figure 10 As shown, the present invention provides an optimal subcarrier combination construction method based on the LSTM model, and the specific steps are as follows:

[0043] S1. Capture CSI raw data and preprocess it to form a raw data set containing several data packets. ,in, represents the i-th original data packet, , the data packet contains several valid subcarrier CSI sampling information. In this embodiment, two ESP32-S3 development boards are used as the transmitter and the receiver respectively to transmit and receive wireless signals, and the original CSI data is captured in an open indoor environment. The preprocessing adopted in this step is to remove outliers from the captured original data and apply smoothing filtering and other preprocessing steps, which are conventional data preprocessing schemes and will not be elaborated here. In the data collection process, this embodiment captured a total of 700 CSI original data packets, each of which contained 52 valid subcarrier CSI sampling information.

[0044] S2. The original data set is processed using the improved elbow rule based on the cost function to determine the value range of the cluster number M, and the optimal cluster number M' is determined from this range.

[0045] In this embodiment, in order to illustrate the implementation method in step S2, the specific steps are as follows:

[0046] S21. Select several CSI raw data packets from the raw data set, and use the elbow rule to determine the value range U of the clustering number M.

[0047] In this embodiment, in order to illustrate the implementation method of determining the value range U of the number of clusters by the elbow rule in step S21, the specific steps are as follows:

[0048] S211, using the number T of subcarrier CSI sampling information contained in the selected several CSI original data packets as the value range of the classification number K of the K-means algorithm, which is expressed as K∈[1,T], and using the K-means algorithm to perform multiple subcarrier clustering within this range to generate multiple clustering results, wherein the clustering results are within the value range of the classification number K;

[0049] S212, calculate the intra-cluster sum of squares based on each clustering result , we get the value range U of the number of clusters M, where U∈K, then The calculation formula is:

[0050] ;

[0051] In the above formula, represents the cluster center of the i-th class, P represents the data points belonging to the i-th class, K represents the number of classifications, represents the i-th cluster;

[0052] When applying a clustering algorithm to group the selected data, the Sum of Squared Errors within Cluster (SSE) will gradually decrease as the number of categories K increases. However, when the number of categories K increases to a certain value, the rate of decrease of SSE will slow down significantly, forming an obvious "elbow" feature on the graph. This feature point is usually considered to represent the optimal number of clusters M'.

[0053] The core value of the elbow rule is to find a balance between inaccurate classification caused by insufficient number of categories K and overfitting caused by too many categories K. In clustering problems, choosing the number of categories K corresponding to the "elbow" can balance the number of clusters and the homogeneity of data within the clusters. SSE will gradually decrease with the increase of the number of categories K, but will experience a rapid decline during the reduction process. As the number of categories K increases further, the rate of decline of SSE will become gentle. In this case, the number of categories K corresponding to the end point of the area where the SSE drops the steepest is usually selected as the optimal number of clusters M', which can best classify the selected data while maintaining the effectiveness of clustering.

[0054] like Figure 2 As shown in the figure, taking a group of (100 CSI raw data packets) CSI amplitude data as an example, each packet contains 52 subcarrier sampling information. First, when the classification number K takes the value [1,52], the K-means algorithm is used to perform multiple clustering within this range. According to the SSE calculated based on each clustering result, it can be seen that an "elbow" is formed between K=41 and K=45, and [41,45] is the value range U of the cluster number M. Therefore, it can be determined that when the optimal cluster number M' value is between 41-45, the clustering effect is better.

[0055] S22. Construct a cost function, calculate the function values ​​corresponding to all cluster numbers in U through the cost function, and take the cluster number corresponding to the minimum function value as the optimal cluster number M'.

[0056] In this embodiment, the cost function in step S22 is expressed by the following formula:

[0057] ,

[0058] In the above formula, F represents the cost function value, D represents the intra-class distance, and L represents the inter-class distance. represents the mean value of the subcarrier CSI sampling information in the kth class, represents the number of subcarrier CSI sampling information in the kth class, represents the mean of all sample data, represents the subcarrier CSI sampling information of the kth class, represents the CSI sampling information of the jth subcarrier in the kth class, , Represents the kth cluster, and M represents the number of clusters. The elbow rule only measures the effectiveness of K-means clustering through the intra-class distance. The sum of the squares of the distances from the samples within the class to the center of their class is recorded as the degree of distortion of the class, and the sum of the distortion degrees of K classes is recorded as the intra-class distance. The smaller the value, the closer the distance between the data within the class and the better the clustering effect. Conversely, the clustering effect is poor. At this time, if the optimal number of clusters k is determined only by the intra-class distance, it may be quite different from the actual number of clusters, resulting in unstable clustering division, poor clustering effect and other problems. Since the ultimate goal of the K-means clustering algorithm is to make the intra-class distance D as small as possible and the inter-class distance L as large as possible, the cost function method comprehensively considers the relationship between the two, constructs the optimal cost function, and calculates the F value under different cluster numbers M. The number of clusters M corresponding to the minimum F value is the optimal number of clusters M', such as Figure 3 As shown in the figure, the horizontal axis is the number of clusters, and the vertical axis is the calculation result of the cost function. When M=42, the value of the cost function is 0.09, and the other cost function values ​​are all above 0.10. Therefore, the improved elbow rule based on the cost function finally determines that M=42 is the optimal number of clusters.

[0059] S3. When the number of classifications is the optimal number of clusters M', the subcarrier CSI sampling information is clustered and data labels are generated. The subcarrier CSI sampling information in each data packet contained in the original data set is labeled using the data labels to form a labeled data set. ,in, represents the i-th data packet with data label, The data packet with data label contains several valid data labels and their corresponding subcarrier CSI sampling information; the clustering result of K-means algorithm when M=42 is 52 subcarriers labeled with data labels, and the labeling result is as follows Figure 4 shown.

[0060] like Figure 5 As shown in the figure, the CSI amplitudes of all subcarriers under data label 1 and data label 4 are plotted and displayed to verify the effectiveness of K-means clustering. Figure 5It can be seen that the CSI amplitude waveforms with data labels 1 and 4 are quite different, and the CSI amplitudes of subcarriers with the same label have similar change trends. Therefore, the subcarriers can be classified into different categories using the K-means clustering algorithm. At the same time, the data labels assigned to the subcarriers are effective. In this embodiment, in order to classify the subcarriers, 42 data labels are used to mark each subcarrier in the original data set. Each subcarrier has 200 sample data. The entire original data set contains 10,400 samples to form the final labeled data set.

[0061] CSI describes the attenuation factor of the wireless signal on each transmission path, that is, the value of each element in the channel gain matrix. It provides detailed information about how the signal propagates in a specific environment and what obstacles or interference may affect it. However, this information is not directly converted into specific labels or classifications, so labeling the data labels is a necessary condition for subsequent LSTM subcarrier classification model training.

[0062] S4. Labeled dataset The training set is divided into a training set and a test set. After the training set is input into the LSTM subcarrier classification model for training, the trained LSTM subcarrier classification model is tested with the test set. In each category output by the LSTM subcarrier classification model, a subcarrier corresponding to the category is randomly selected to form the optimal subcarrier combination. The training set accounts for 80% of the labeled data set, and the test set accounts for 20% of the labeled data set. Figure 6 As shown in Figure 2, the structure of the LSTM subcarrier classification model is as follows from input to output: LSTM layer, fully connected layer (FC) and output layer. The model processes the input labeled data through the LSTM layer, maps the output of the LSTM layer to the classification space through the fully connected layer and inputs it to the output layer. The output layer of the model uses The activation function is The activation function outputs the predicted probability of each category , predicted probability The calculation formula is:

[0063] ,

[0064] In the above formula, is the predicted probability of the subcarrier class, and b represent the weight matrix and bias respectively, and Y represents the output data after the fully connected layer maps the output of the LSTM layer to the classification space.

[0065] LSTM neural network introduces the concept of cell gate and has the same chain structure. It can pass context information through the gate structure, thus realizing the transmission of information over long distances. LSTM has two state transmission mechanisms: one is the cell state, denoted as ; The other is the hidden state, denoted as . Using the input of the current time step and the hidden state passed down from the previous time step LSTM can generate four key states through training, namely the forget gate, input gate, candidate memory, and updated hidden state. At each moment, the LSTM unit receives input and the hidden state at the previous moment , and calculate the hidden state at the current moment according to the state of the gate control unit and output In this example, the labeled dataset The input is fed into LSTM, and finally a high-dimensional feature vector is output, which contains the long-term dependencies in the subcarrier CSI sampling information.

[0066] Use MATLAB to define a cell array XTrain containing 7280 42-dimensional time series, corresponding to a cell array YTrain containing subcarrier data labels from 1 to 42, representing 42 categories respectively. The learning rate is 0.00125, the number of iterations is 600, and the output layer uses the softmax activation function. The training process is as follows Figure 7 As shown, it can be seen that after 600 iterations, the accuracy of the training model can reach up to about 86.1%.

[0067] In this example, the labeled dataset 20% of the data is divided into a test set, a total of 1560 samples, of which each subcarrier has 30 samples. The test results are as follows: Figure 8 As shown in the figure, according to the test results, the average test accuracy of the LSTM subcarrier classification model in 42 categories is 85.6%, among which the accuracy of categories "8", "19", "20", "38", "39" and "41" is close to 90%. Fig. 9 As shown, each data label represents a category, and a subcarrier corresponding to the category is randomly selected from each category to form an optimal subcarrier combination, a total of 42 subcarriers, forming the optimal subcarrier combination.

[0068] LSTM, as a variant of recurrent neural network, has the ability to process sequence data and performs well in dealing with long-term dependencies in time series, while subcarrier classification is a task of signal classification in the frequency domain. Data labels are assigned, and the long-term dependencies in the CSI labeled data set are effectively captured through the LSTM neural network, so the classification of CSI similar sequences can be completed more efficiently and accurately.

[0069] In order to verify the availability of the optimal subcarrier combination in the present invention, the CSI quotient waveforms of Alice and Bob are taken as an example to observe the CSI quotient waveforms of the two communicating parties. If the waveforms are completely different, the keys generated after data quantization will not be too similar, as described in detail as follows:

[0070] The different-numbered subcarrier quotient model uses the quotient between subcarriers with different numbers to obtain reciprocal data. However, the continuous groups of keys generated by quantizing the quotient sequence of CSI measured on the same subcarrier and several adjacent subcarriers will produce a large proportion of repeated segments, which can be easily cracked by eavesdroppers, resulting in extremely poor key availability.

[0071] The optimal subcarrier combination is obtained by the present invention and used as a random source of the different-numbered subcarrier quotient model, that is, the optimal subcarrier combination is used as the data source, and the adjacent data labels are 2, 3, 4 and their corresponding subcarriers numbered 11, 26 and -4 are taken as examples, and the CSI of subcarrier No. 11 is taken as the quotient with that of subcarrier No. 26 and -4 respectively, and the CSI quotient waveforms of the two communicating parties are observed. The results are as follows: Fig.10 As shown in the figure, the waveforms obtained by taking the amplitude of subcarrier No. 11 and subcarrier No. 26 and No. -4 as quotients are completely different. Therefore, the keys generated by the quantization of the two sets of data will not be too similar. It has been verified that the number of inconsistent bits of the CSI amplitude quotients of Alice and Bob after quantization is within the correctable range, that is, less than 25% specified in the BCH algorithm. Therefore, the optimal subcarrier combination algorithm can effectively avoid the occurrence of similar CSI quotients.

[0072] The present invention is applicable to a CSI-based different-numbered subcarrier quotient model, solves the problem of too many key repetition segments due to the correlation of adjacent subcarriers, and can effectively improve the key update rate. The present invention discloses an optimal subcarrier combination construction method based on an LSTM model, uses an improved elbow rule based on a cost function and a K-means algorithm to process an original data set, determines a value range of a clustering number, and determines an optimal clustering number from the range, performs subcarrier clustering on subcarrier CSI sampling information by using the K-means algorithm when the classification number is the optimal clustering number and generates data labels, and uses an LSTM subcarrier classification model to output an optimal subcarrier combination. The optimal subcarrier combination in the present invention can effectively reduce the correlation between subcarriers, so that the key repetition ratio is significantly reduced. In theory, there will be no subcarrier sampling sequence with a high similarity, which can effectively prevent a large number of repeated bits from appearing in the key sequence, improve the key update rate, and effectively ensure communication security in the case of frequent communication interactions, especially in the case of frequent communication interactions. The present invention is also applicable to scenes with frequent communication interactions such as sufficient data volume and the need to quickly generate and distribute a large number of keys.

[0073] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.

[0074] In addition, it should be understood that although the present specification is described according to implementation modes, not every implementation mode contains only one independent technical solution. This description of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation modes that can be understood by those skilled in the art.

Claims

1. A method for constructing an optimal subcarrier combination based on an LSTM model, characterized in that: The following steps are involved: S1. Capture CSI raw data and preprocess to form a raw data set containing n data packets, each of which contains several valid subcarrier CSI sampling information; S2, select several data packets from the original data set, and use the elbow rule to determine the value range U of the cluster number M; S3. Construct a cost function, calculate the function values ​​corresponding to all cluster numbers in U through the cost function, and take the cluster number corresponding to the minimum function value as the optimal cluster number M'. The cost function is expressed as: In the above formula, F represents the cost function value, D represents the intra-class distance, and L represents the inter-class distance. represents the mean value of the CSI sampling information of the subcarrier in the kth class, n k represents the number of subcarrier CSI sampling information in the kth class, represents the mean of all sample data, p kj Indicates the j-th subcarrier CSI sampling information of the k-th class, j = 1, 2, ... n k , C k represents the kth cluster, and M represents the number of clusters; S4. When the number of clusters is the optimal number of clusters M', the subcarrier CSI sampling information is clustered and a data label is generated. The subcarrier CSI sampling information in each data packet in the original data set is labeled using the data label to form a labeled data set X' = {X1', X2', X3'..., X i ′}, where X i ′ represents the i-th data packet with data label, i = 1, 2, 3...n; S5. Divide the labeled data set into a training set and a test set. After the training set is input into the LSTM subcarrier classification model for training, the trained LSTM subcarrier classification model is tested using the test set. Finally, in each category output by the LSTM subcarrier classification model, any subcarrier corresponding to the category is selected to form an optimal subcarrier combination. The structure of the LSTM subcarrier classification model is, in the order from input to output, an LSTM layer, a fully connected layer, and an output layer.

2. The optimal subcarrier combination construction method according to claim 1, characterized in that: In step S2, the elbow rule is used to determine the value range U of the number of clusters, and the specific steps are: S21, taking the number of subcarrier CSI sampling information T contained in the selected data packets as the value range of the classification number K of the K-means algorithm, expressing it as K∈[1,T], and performing multiple subcarrier clustering within this range using the K-means algorithm to generate multiple clustering results, wherein the clustering results are within the value range of the classification number K; S22. Calculate the intra-cluster sum of squares SSE according to each clustering result to obtain the value range U of the number of clusters M, where U∈K.

3. The method for constructing an optimal subcarrier combination according to claim 1, characterized in that: The output layer of the LSTM subcarrier classification model adopts a softmax activation function to output the predicted probability of each category. Prediction probability The calculation formula is: In the above formula, is the predicted probability of the subcarrier category, W T and b represent the weight matrix and bias respectively, and Y represents the output data after the fully connected layer maps the output of the LSTM layer to the classification space.

4. The method for constructing an optimal subcarrier combination according to claim 1, characterized in that: The data packet with data tags contains several valid data tags and their corresponding subcarrier CSI sampling information.

5. The method for constructing an optimal subcarrier combination according to claim 1, characterized in that: The training set accounts for 80% of the labeled data set, and the test set accounts for 20% of the labeled data set.

Citation Information

Patent Citations

  • WiFi identity recognition method fused with deep learning model

    CN110288018A

  • Channel state information indoor positioning method based on subcarrier selection

    CN115278518A