Semi-supervised representation contrastive learning method for massive MIMO positioning
Through the semi-supervised representation contrast learning method, large-scale MIMO positioning is performed using uplink received signals, which solves the problems of high labeling costs and frequent database updates, and achieves efficient position estimation and improved positioning accuracy.
Patent Information
- Application Number
- CN202310376215.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-11
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2043-04-11
AI Technical Summary
Existing massive MIMO positioning methods require a large number of CSI samples with ground truth locations, which has high labeling costs and requires constant updating of the database, affecting performance.
A semi-supervised representation contrastive learning method is adopted to use the uplink received signal for position estimation. The encoder and contrastive loss function are pre-trained to avoid precise channel estimation, and a self-supervised model is used for downstream positioning tasks.
It achieves efficient position estimation, reduces annotation costs, and improves positioning accuracy, which is superior to traditional methods.
Smart Images

Figure CN116383656B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of wireless communication, and relates to a wireless positioning method of a large-scale MIMO system. BACKGROUND
[0002] In recent decades, accurate positioning has been paid more and more attention as an important driving factor for many location-based services such as navigation, intelligent robots and Internet of Things. The massive multiple-input multiple-output (MIMO) technology is widely used in 5G and later wireless networks, which provides stronger sensing and positioning capabilities for the system. Most of the existing large-scale MIMO positioning methods take channel state information (CSI) as the starting point. In recent years, the mapping from CSI to user terminal position (UT) has been completed with the help of deep neural networks (DNN). These methods need to supervise the training of DNN in offline mode so as to predict the related position coordinates of UT in online mode.
[0003] The existing large-scale MIMO fingerprint positioning method needs a special training dataset containing a large number of CSI samples labeled with ground truth positions. However, firstly, accurate channel estimation results are required for CSI samples, and secondly, although the cost of labeling is linearly related to the size of the dataset (a constant time is required to label each example), the performance of the model is only sub-linearly related to it. This means that labeling more samples will become less cost-effective. Thirdly, due to the continuous change of CSI, the database needs to be updated constantly, which further increases the cost of labeling one by one. The above reasons may be the performance limiting factors of the existing method. SUMMARY
[0004] Technical problem: In order to solve the above problems, the object of the application is to give a semi-supervised representation contrast learning method for large-scale MIMO positioning, which only uses the position estimation method of uplink received signals, such signals are readily available at the base station (BS) without the need for accurate channel estimation and a large number of ground truth position labels, which can make the database easier and faster to update. This method can achieve excellent performance, avoid accurate channel estimation, realize labeling efficiency, and is worth popularization and application.
[0005] Technical solution: The semi-supervised representation contrast learning method for large-scale MIMO positioning of the application comprises the following steps:
[0006] Step 1: According to the configuration of the large-scale MIMO system, the beam domain channel representation is given, and the received signal representation form is obtained.
[0007] Step 2, in order to pre-train the encoder in the pre-training stage, first according to the available received signal of the reference point RP(Reference Point) in different positions Create positive and negative samples;
[0008] Step 3, use the encoder F(·) to convert the positive and negative samples to the feature representation space;
[0009] Step 4, update the encoder weights using the contrastive loss function and the optimizer;
[0010] Step 5, append a randomly initialized fully connected regression layer f(·) on top of the encoder to complete the downstream positioning task.
[0011] Wherein,
[0012] Step 1, the large-scale MIMO system configuration includes a base station, K users; the base station is configured with a large-scale uniform linear array antenna, and the antenna spacing is half a wavelength; the user is configured with a single antenna; the number of antennas on the base station side is N r ; Orthogonal Frequency Division Multiplexing OFDM modulation is used to convert the frequency selective fading channel into multiple parallel channels; the number of subcarriers in the large-scale MIMO-OFDM system is N c , N p pilot subcarriers are used for uplink pilot signal transmission; the length of the cyclic prefix is represented as N g , and the sampling interval is represented as T s ; the subcarrier spacing is Let Θ i , τ j be the direction cosine and delay of the sampling, N a and N d are called the number of samples in the spatial and frequency domains, a(Θ i ), b(τ j ) are called the sampling steering vectors in the spatial and frequency domains, respectively; in order to ensure the accuracy of quantization, N a ≥N r , N d ≥N g , is uniformly distributed between (-1, 1], is uniformly distributed between (0, N g T s ]; N r is the number of antennas, and in addition, the matrices A and B are defined as
[0013]
[0014]
[0015] By using the refined-based double-beam channel model, in the t-th OFDM symbol, the spatial-frequency domain channel matrix between the k-th user and the base station can be modeled as the channel model:
[0016] H k,t = A(Ξ k ⊙V k,t )B T (1)
[0017] where is a complex Gaussian random matrix, each element is independent and identically distributed (i.i.d.) with zero mean and unit variance, and the non-negative matrix remains unchanged in different OFDM symbols; define G k,t = Ξ k ⊙V k,t is called the refined-based double-beam channel matrix, and the channel power matrix of the k-th user is defined as Ω k = Ξ k ⊙Ξ k , which is a sparse matrix because most of the channel power is distributed in a limited number of distinguishable spatial directions and time delays; the superscript T is the transpose;
[0018] The signal of the t-th OFDM symbol at the base station is given by the following received signal model
[0019] Y t = H k,t X k + Z t
[0020] where Z t is a complex Gaussian noise matrix composed of i.i.d. elements with zero mean and variance ; X k is the user uplink pilot signal, and by substituting the channel model (1) into the above received signal model, it can be rewritten as
[0021] Y t = AG k,t B T X k + Z t = AG k,t P+ Z t
[0022] By left-multiplying Y t by the sampling matrix A H , A H is the sampling matrix composed of steering vectors on formula (1), and the superscript H represents the conjugate transpose. Right-multiply by the sampling matrix P H, P = B r X k , resulting in the received pilot signal in the refined beam domain as
[0023] P H , P = B T X k
[0024] A H Y t P H = A H AG k,t PP H + A H Z t P H
[0025] Let Φ denote the expected value of the received power matrix in the refined beam domain as follows, where E{.} denotes the expectation operation, the superscript * represents the conjugate operation, and denotes the Hadamard product of matrices
[0026] Φ = E{(A H Y t P H )⊙(A H Y t P H ) *}
[0027] resulting in the received model for the channel power matrix Ω k
[0028] Φ k = T a Ω k T d + N
[0029] where T a , T d , and N are deterministic matrices defined as
[0030] T a = (A H A)⊙(A H A) *
[0031] T d = (P H P)⊙(P H P) *
[0032]
[0033] To estimate the position coordinate vector k of the k-th user in the two-dimensional plane from the received signal Φ where Coordinates representing x-axis and y-axis respectively; a self-supervised model is used to obtain an accurate solution.
[0034] The creation of positive and negative samples is described in step 2. In the pre-training phase, the base station obtains the received signals of all reference points The subscript represents the serial number of the reference point. For a small batch of reference point received signals Assume that the received signal of the i-th reference point is It is considered as an "anchor point", and its positive sample is denoted as After different data augmentation, a small batch of The received signals of other reference points in the batch are negative samples of the i-th reference point and form a set, denoted as
[0035] Step 3 describes the use of an encoder F(·) to convert positive and negative samples to a feature representation space. The encoder block F(·) : is composed of four two-dimensional convolutional layers and one fully connected feature output layer, each followed by an activation layer, where d is the output dimension, R is the real number space, and the role of the encoder is to convert positive and negative samples from an N a ×N d dimensional real number space to a d-dimensional real number space; ReLU functions are used for all activation layers, and batch normalization (BN) layers are added in the middle to minimize overfitting and gradient vanishing or explosion. In the pre-training phase, a nonlinear projection head g(·) is connected to the top of the encoder to improve the representation quality of the encoder. In the downstream task, the nonlinear projection head g(·) is abandoned, and only the trained encoder is used.
[0036] Step 4 describes the use of a contrastive loss function and an optimizer to update the encoder weights. The encoder is pre-trained using a contrastive loss function with unlabeled received signal data from different reference points Consider an encoded anchor q = F(Φ i ) ∈ R d×1 is a real number vector with dimension d, and a batch of encoded negative samples {k0 = F(Φ0), k1 = F(Φ1), k2 = F(Φ2),...} comes from the set Let there be an encoded positive sample k + = F(A(Φ i )) that matches q, where q is a real number vector with dimension d, which is the output of the anchor point through the encoder. The contrastive loss is a function that is minimized when q and k +Similarity with all other {k0, k1, k2,...} is low, and the value is low; the similarity is measured by dot product, and a form of contrastive loss function is considered, called information noise contrastive loss:
[0037]
[0038] Where tau is a temperature hyperparameter, the result is calculated on a positive sample and K negative samples, and the loss is based on the log loss of a (K+1)-classification softmax classifier that tries to classify q as k + .
[0039] Step 5 described above adds a randomly initialized fully connected regression layer f(·) on top of the encoder to complete the downstream positioning task, and a randomly initialized fully connected regression layer f(·) : R d →R 2 Is connected to the top of the encoder to complete the downstream positioning task; for 5% of all reference points, the received signal of this part of the reference point Is marked with the ground truth position coordinate vector , Using this part of the labeled data set, the trained encoder F(·) and the regression module f(·) are fine-tuned, and the loss function is as follows:
[0040] The predicted Is calculated using the mean square error MSE loss function, and the output value of the regression network is a 2-dimensional real vector, representing the distance between the network's position prediction and the actual position coordinate vector p i , and the loss function with L2 regularization is as follows
[0041]
[0042] Where N train Is the number of training data, w is the vector of all trainable parameters of the DNN, and gamma is a hyperparameter.
[0043] Beneficial effects: In the present application, a semi-supervised positioning method based on contrastive learning is studied for large-scale MIMO systems. A large number of unlabeled received signals easily obtained by base stations are used to pretrain the encoder. Through the contrastive loss function, the encoder can distinguish between positive and negative samples in the representation space. Simulation results show that compared with the baseline method of supervised training, the entire network can well complete the downstream positioning task after fine-tuning. Compared with other existing methods, this method can obtain excellent performance, avoid accurate channel estimation, realize labeling efficiency, and is worth popularization and application. BRIEF DESCRIPTION OF DRAWINGS
[0044] Figure 1 A planar illustration of a positioning scenario for a massive MIMO system in an embodiment of the application.
[0045] Figure 2 A plot comparing the positioning estimation performance of the application with other algorithms in an embodiment of the application. DETAILED DESCRIPTION
[0046] The technical solutions provided by the application will be described in detail below in conjunction with specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the application and not to limit the scope of the application.
[0047] As shown in Figure 1 , the semi-supervised representation contrast learning method for massive MIMO positioning disclosed in an embodiment of the application uses a two-dimensional geometric-based propagation model to simulate the wireless transmission environment. Figure 1 A two-dimensional planar layout is included to explain the simulation setup; the coordinates (X, Y) of the plane correspond to the X and Y axes. It is assumed that the base station is located at the (0, 0) m coordinate origin, and its equipped uniform linear array is parallel to the Y axis, with 128 antennas and 256 beams. The area considered is a square with a center at (500, 0) m and a side length of 50 m. There are 50 scatterers per square kilometer. A path is any unobstructed transmission between the user and the base station that is not blocked by other scatterers. A single bounce (GBSB) propagation based on geometry is considered, which is used to simplify the model without losing generality. The bandwidth of the uplink OFDM channel is 20 MHz, with 1024 subcarriers.
[0048] The area to be positioned is evenly divided into multiple reference points. The base station collects 10,000 samples of received pilot signals as a training dataset, and uses 5% of them with real positions as a labeled subset for fine-tuning. 500 user terminals are randomly distributed within the positioning area, and the base station collects as a validation dataset for the fine-tuning phase. 500 randomly distributed user terminals are regenerated, and their are collected for position prediction in online mode. The encoder consists of four identical CNN layers and a projection layer, with the CNN consisting of 16 3x3 kernels. The feature representation size of the encoder is d = 1024, and the projection head g(·) is 128. For the given method, MATLAB 2020a is used to calculate the received signal and coordinates. The network is trained and tested using TensorFlow 2.6. The simulation is performed on a computer equipped with an Intel Core i7-8700k CPU and a Geforce GTX 3080 10GB GPU.
[0049] The following is a description of the most important hyperparameters: batch size 32: Since the objective can be interpreted as a classification (broadly speaking) of a batch of Φ i , the size of the batch is actually a more important hyperparameter than usual. The higher the better. Temperature 0.1: The temperature defines the "softness" of the softmax distribution used for the cross-entropy loss and is an important hyperparameter. Lower values usually lead to higher contrast accuracy. Optimizer: Adam was used because it provided good performance at a learning rate of 0.0005 and other default parameters.
[0050] Pre-training procedure:
[0051] In addition to the InfoNCE loss function described above, the following metrics were used to monitor the performance of the pre-training: Contrast accuracy (c_acc): A self-supervised metric, i.e. the ratio of cases in which the encoded representation of a reference point data is more similar to its different augmented versions than to the representations of any other reference point in the current batch. Even without labeled samples, the contrast accuracy can be used for hyperparameter tuning. Linear probe accuracy (p_acc): Linear probe is a popular metric for evaluating self-supervised models. It is computed as the accuracy of a logistic classifier trained on top of the encoder representations. In the case, this was done by training a single fully connected on top of the frozen encoder. The 5% labeled reference points were divided into 25 classes, which were trained during pre-training. This way, its value can be monitored during training, which helps with experimentation and debugging.
[0052] The semi-supervised representation contrast learning method for massive MIMO positioning of the application comprises the following steps:
[0053] Step 1, according to the configuration of the massive MIMO system, the beam domain channel representation is given, and the received signal representation form is obtained;
[0054] Step 2, in order to pre-train the encoder in the pre-training stage, first, according to the available received signals of the reference points RP (Reference Point) at different positions Create positive and negative samples;
[0055] Step 3, use the encoder F(·) to convert the positive and negative samples to the feature representation space;
[0056] Step 4, update the encoder weights using the contrast loss function and the optimizer;
[0057] Step 5, append a randomly initialized fully connected regression layer f(·) on top of the encoder to complete the downstream positioning task.
[0058] wherein,
[0059] The large-scale MIMO system configuration described in step 1 includes one base station, K users; the base station is configured with a large-scale uniform linear array antenna, and the antenna spacing is half a wavelength; the user is configured with a single antenna; the number of antennas on the base station side is N r ; orthogonal frequency division multiplexing (OFDM) modulation is used to convert a frequency-selective fading channel into multiple parallel channels; the number of subcarriers in the large-scale MIMO-OFDM system is N c , N p pilot subcarriers are used for uplink pilot signal transmission; the length of the cyclic prefix is denoted as N g , and the sampling interval is denoted as T s ; the subcarrier spacing is Let be the sampled direction cosine and delay, N a and N d are referred to as the number of samples in the spatial and frequency domains, respectively, and a(Θ i ), b(τ j ) are referred to as the sampled steering vectors in the spatial and frequency domains, respectively; in order to ensure the accuracy of quantization, N a ≥ N r , N d ≥ N g , is uniformly distributed between (-1, 1], is uniformly distributed between (0, N g T s ]; N r is the number of antennas, and in addition, the matrices A and B are defined as
[0060]
[0061]
[0062] By using the refined-based double-beam channel model, the spatial-frequency domain channel matrix between the kth user and the base station in the tth OFDM symbol can be modeled as the channel model:
[0063] H k,t = A(Ξ k ⊙V k,t )B T (1)
[0064] where is a complex Gaussian random matrix, each element is independently and identically distributed (i.i.d.) with zero mean and unit variance, and the non-negative matrix remains unchanged in different OFDM symbols; define as the refined-based double-beam channel matrix, and the channel power matrix of the kth user is defined as Ωk k k This is a sparse matrix because most of the channel power is distributed in a limited number of resolvable spatial directions and time delays; the superscript T is the transpose;
[0065] The signal at the tth OFDM symbol at the base station is given by the received signal model
[0066] Y t = H k,t X k + Z t
[0067] where Z t is a complex Gaussian noise matrix with i.i.d. elements of zero mean and variance ; X k is the user uplink pilot signal. Substituting the channel model (1) into the received signal model above, we have
[0068] Y t = AG k,t B T X k + Z t = AG k,t P + Z t
[0069] By left-multiplying Y t by the sampling matrix A H , A H is the sampling matrix composed of steering vectors on formula (1), the superscript H represents the conjugate transpose. Right-multiplying by the sampling matrix P H , P = B T X k , the received pilot signal in the refined beam domain is
[0070] P H , P = B T X k
[0071] A H Y t P H = A H AG k,t PP H + A H Z t P H
[0072] Let Φ denote the expected value of the received power matrix in the refined beam domain as follows, where E{.} denotes the expectation operation, the superscript * represents the conjugate operation and the symbol represents the Hadamard product of matrices
[0073] Φ=E{(A H Y t P H )⊙(A H Y t P H ) *}
[0074] Get the channel power matrix Ω k The receiving model is
[0075] Φ k =T a Ω k T d +N
[0076] Where T a 、T d and N is a deterministic matrix, defined as
[0077] T a =(A H A)⊙(A H A) *
[0078] T d =(P H P)⊙(P H P) *
[0079]
[0080] In order to receive the signal Φ k Estimate the position coordinate vector of the kth user on the two-dimensional plane in Represent the coordinates of the x-axis and y-axis respectively; a self-supervised model is used to obtain an accurate solution.
[0081] Step 2 creates positive and negative samples. During the pre-training phase, the base station obtains the received signals of all reference points. The subscript represents the serial number of the reference point, for a small batch of reference points receiving signals Assume that the received signal at the i-th reference point is It is regarded as an "anchor point", and its positive samples are recorded as data augmentation. After different data augmentations, small batches The received signals of other reference points in are all negative samples of the i-th reference point, and form a set, recorded as
[0082] The principle is explained below: Contrastive learning is considered to be a dictionary-style query problem; in contrastive learning, each sample input into the neural network can be regarded as a query, and other samples in the dataset can be regarded as entries in the dictionary. Usually the "query" point (query), also called the "anchor" point (anchor), is compared with other samples. The goal of contrastive learning is to project the query sample into the feature space and compare it with the entries in the dictionary to find the dictionary entry that is most similar to the query sample. In this way, contrastive learning can learn effective feature representations and achieve good performance in many machine learning tasks.
[0083] The "anchor" can be regarded as an object of focus, which is used to divide other samples in the dictionary into two categories: positive samples that are similar to the anchor point and negative samples that are not similar to the anchor point. Usually, the anchor point and the positive samples form a group of sample pairs, and the anchor point and the negative samples form another group of sample pairs. Then, the model is trained by comparing the similarity of the two groups of sample pairs through the contrast loss function. In the present invention, if the signal received by the reference point 1 is is the "anchor", and the "positive" sample is the "anchor" Data augmentation, using means "anchor" and “positive” samples are positive pairs of each other; “negative” samples are mini-batches randomly selected from the received signals of other reference points Self-supervised learning can train an encoder to perform a proxy task of dictionary lookup: the "anchor" encoded by the neural network encoder should be similar to the encoded output of its matching "positive" sample and dissimilar to other samples; the learning process is formulated as minimizing the contrastive loss function; the main purpose of self-supervised learning is to pre-train the encoder to output feature representations and then transfer this encoder to downstream tasks through fine-tuning. Contrastive learning obtains positive and negative samples through data augmentation. The two most important data augmentation methods A(·) are as follows:
[0084] Cropping: Randomly crop the same reference point different parts of the model, forcing the model to the same reference point Encode different parts of
[0085] Jitter: A principled approach is to jitter the reference point Perform affine transformation;
[0086] In addition, random horizontal flipping is added in data augmentation. The above three operations together constitute our data augmentation method, using hyperparameters to control the strength of data augmentation; strong data augmentation is suitable for contrastive learning, and weak data augmentation is suitable for supervised regression to avoid overfitting on a small number of labeled examples;
[0087] The following summary is made for positive and negative samples: assume that the received signal of the 1st reference point is "anchor", its positive sample is The received signals of other reference points after different data augmentations are negative samples of the first reference point, and form a set, denoted as set
[0088] Step 3 uses the encoder F(·) to convert positive and negative samples into a feature representation space, and the encoder block is composed of four two-dimensional convolutional layers and a fully connected feature output layer, each layer is followed by an activation layer, where d is the dimension of the output, R is the real space, and the role of the encoder is to convert positive and negative samples from an N a ×N d dimensional real space to a d-dimensional real space; ReLU function is used for all activation layers, and batch normalization BN layer is added in the middle to minimize overfitting and gradient vanishing or explosion, in the pre-training stage, a nonlinear projection head g(·) is connected to the top of the encoder to improve the representation quality of the encoder, and the nonlinear projection head g(·) is abandoned in the downstream task, and only the trained encoder is used.
[0089] Step 4 uses a contrastive loss function and an optimizer to update the encoder weights, pre-trains the encoder using a contrastive loss function, and uses unlabeled received signal data from different reference points Consider an encoded anchor q=F(Φ i )∈R d×1 is a real vector with dimension d, and a batch of encoded negative samples {k0=F(Φ0), k1=F(Φ1), k2=F(Φ2),...} comes from set There is an encoded positive sample k + =F(A(Φ i )) that matches q, q is a real vector with dimension d, which is the output of the anchor point after the encoder, and the contrastive loss is a function, when q is similar to k + and dissimilar to all other {k0, k1, k2,...}, its value is low; the similarity is measured by dot product, and a form of contrastive loss function is considered, called information noise contrastive loss:
[0090]
[0091] where τ is a temperature hyperparameter, the result is computed on a positive sample and K negative samples, this loss is the log loss of a (K+1)-classification softmax-based classifier that tries to classify q as k + .
[0092] Step 5 described in the top of the encoder attached a random initialization of a fully connected regression layer f(·) to complete the downstream positioning task, a random initialization of a fully connected regression layer f(·): R d →R 2 Connected to the top of the encoder to complete the downstream positioning task; for 5% of all reference points, the received signal of this part of the reference point With the ground true position coordinate vector Marked, Using this part of the labeled data set, the encoder F(·) and the regression module f(·) that have been trained are fine-tuned, and the loss function is as follows:
[0093] The predicted is the output value of the regression network, which is a 2-dimensional real vector, representing the network's prediction of the user's position and the actual position coordinate vector p i The distance between them, the loss function with L2 regularization is as follows
[0094]
[0095] Where N train is the number of training data, w is the vector of all trainable parameters of the DNN, and γ is a hyperparameter.
[0096] Implementation effect
[0097] In order to make the personnel in the technical field better understand the scheme of the present application, the performance results of a semi-supervised representation comparison learning method for large-scale MIMO positioning in the embodiment under the specific system configuration and the existing positioning method are compared.
[0098] Using 5% of the labeled data set The encoder F(·) and the regression module f(·) that have been trained are fine-tuned. A baseline supervised model uses random initialization, uses the same encoder architecture and the same labeled data set For training. In Figure 2 The cumulative distribution function (CDF) of the online position prediction error is plotted. From Figure 2To see, the position regression performance of the pre-trained encoder + regression network is compared with the baselines and existing fingerprint-based methods. Simulation results show that, using only a small amount of labeled data, the encoder can achieve an RMSE of 1.3373 in the downstream localization task, outperforming the baseline methods with RMSEs of 1.7066 and 1.6524.
[0099] In the embodiments provided by the present application, it should be understood that the disclosed method can be implemented in other ways than those described in the embodiments without departing from the spirit and scope of the application. The present embodiments are only exemplary and should not be used to limit the purpose of the application. For example, some features can be ignored or not implemented.
[0100] The technical means disclosed in the present application scheme is not limited to the technical means disclosed in the above-mentioned embodiments, but also includes the technical scheme composed by any combination of the above technical features. It should be pointed out that for ordinary skilled in the art, without departing from the principle of the present application, a number of improvements and refinements can also be made, which are considered as the protection scope of the present application.
Claims
1. A semi-supervised representation contrastive learning method for massive MIMO positioning, characterized by: The method comprises the following steps: Step 1: Based on the massive MIMO system configuration, a beam-domain channel representation is given to obtain the received signal representation. Step 2: In order to pre-train the encoder in the pre-training phase, firstly, the available received signals at different reference points RP are Create positive and negative samples; Step 3: Use encoder F(·) to convert positive and negative samples into feature representation space; Step 4: Update the encoder weights using the contrastive loss function and optimizer. Step 5: Add a randomly initialized fully connected regression layer f(·) on top of the encoder to complete the downstream positioning task; add a randomly initialized fully connected regression layer f(·): R d →R 2 Connected to the top of the encoder to complete the downstream positioning task; for 5% of all reference points, the received signal of this part of the reference points Use the ground truth position coordinate vector Marking, Using this partially labeled dataset, we fine-tune the trained encoder F(·) and regression module f(·), and the loss function is as follows: The predicted value is calculated using the mean squared error (MSE) loss function. It is the output value of the regression network, which is a 2-dimensional real number vector representing the network's prediction of the user's location and the actual location coordinate vector The distance between, the loss function with L2 regularization is described as follows where N train is the number of training data, w is the vector of all trainable parameters of the DNN, and γ is a hyperparameter.
2. The semi-supervised representation contrastive learning method for massive MIMO positioning according to claim 1, characterized in that: The massive MIMO system configuration described in step 1 includes one base station and K users; the base station is equipped with a massive uniform linear array antenna with an antenna spacing of half a wavelength; the user is equipped with a single antenna; the number of antennas on the base station side is N r Orthogonal frequency division multiplexing (OFDM) modulation is used to convert the frequency selective fading channel into multiple parallel channels. The number of subcarriers in a massive MIMO-OFDM system is N. c , N p pilot subcarriers are used for uplink pilot signal transmission; the length of the cyclic prefix is expressed as N g , the sampling interval is expressed as T s ; The subcarrier spacing is Let Θ i , τ j is the direction cosine and delay of the samples, N a and N d It is called the number of samples in the spatial domain and frequency domain, a(Θ i ), b(τ j ) are respectively called the sampling rudder vectors in the spatial domain and the frequency domain; in order to ensure the accuracy of quantization, N a ≥N r , N d ≥N g , is a uniform distribution between (-1,1], is (0, N g T s ] uniform distribution between; N r is the number of antennas, and matrices A and B are defined as By using the dual-beam channel model based on refinement, in the t-th OFDM symbol, the space-frequency channel matrix H between the k-th user and the base station is k,t It can be modeled as a channel model: H k,t =A(Ξ k ⊙V k,t )B T (1) in is a complex Gaussian random matrix, each element is independent and identically distributed (iid) with zero mean and unit variance, non-negative matrix Remains unchanged in different OFDM symbols; define G k,t =Ξ k ⊙V k,t It is called the refinement-based dual-beam domain channel matrix, and the channel power matrix of the kth user is defined as Ω k =Ξ k ⊙Ξ k , which is a sparse matrix because most of the channel power is distributed in a limited number of resolvable spatial directions and time delays; where the superscript T is the transpose and ⊙ represents the Hadamard product of the matrix; The received signal of the tth uplink OFDM symbol at the base station is The received signal model is given by Y t =H k,t X k +Z t where Z t is composed of a mean of zero and a variance of The complex Gaussian noise matrix composed of iid elements; X k is the user uplink pilot signal. Substituting the channel model (1) into the above received signal model, it can be rewritten as Y t =AG k,t B T X k +Z t =AG k,t P+Z t By adding Y t Left multiply by the sampling matrix A H , where A H is the sampling matrix formed by the rudder vector in formula (1), and the superscript H represents the conjugate transpose; Right multiply by the sampling matrix p H , P=B T X k , the received pilot signal in the refined beam domain is obtained as p H ,P=B T X k A H Yes t P H =A H AG k,t PP H +A H Z t P H Let Φ k The expected value of the received power matrix in the refined beam domain is as follows, where E{·} represents the expectation operation, the superscript * represents the conjugate operation, and ⊙ represents the Hadamard product of the matrix. Φ k =E{(A H AND t P H )⊙(A H AND t P H ) * } Get the channel power matrix Ω k The receiving model is Φ k =T a Ω k T d +N Where T a 、T d and N is a deterministic matrix, defined as T a =(A H A)⊙(A H A) * T d =(P H P)⊙(P H P) * To receive the signal Φ from the kth user k Estimate the position coordinate vector of the kth user on the two-dimensional plane in Represent the coordinates of the x-axis and y-axis respectively; a self-supervised model is used to obtain an accurate solution.
3. The semi-supervised representation contrastive learning method for massive MIMO positioning according to claim 1, characterized in that: Step 2 creates positive and negative samples. During the pre-training phase, the base station obtains the received signals of all reference points. The subscript represents the serial number of the reference point. For a small batch of reference points receiving signals The batch size is N RP , assuming that the received signal at the i-th reference point is It is regarded as an "anchor point" and its positive samples are recorded as data augmented After different data augmentations, small batches The received signals of other reference points in are all negative samples of the i-th reference point, and form a set, recorded as Among them A i (·) Generally refers to a data augmentation method.
4. The semi-supervised representation contrastive learning method for massive MIMO positioning according to claim 2, characterized in that: The encoder F(·) described in step 3 converts positive and negative samples into feature representation space, and the encoder block It consists of four two-dimensional convolutional layers and a fully connected feature output layer. Each layer is followed by an activation layer, where d is the output dimension and R is the real number space. The role of the encoder is to transform positive and negative samples from N a ×N d The real space of dimension d is transformed into the real space of dimension d; the ReLU function is used for all activation layers, and a batch normalization BN layer is added in the middle to minimize overfitting and gradient disappearance or explosion. In the pre-training stage, a nonlinear projection head g(·) is connected to the top of the encoder to improve the representation quality of the encoder. The nonlinear projection head g(·) is abandoned in downstream tasks, and only the trained encoder is used.
5. The semi-supervised representation contrastive learning method for massive MIMO positioning according to claim 3, characterized in that: Update the encoder weights using the contrastive loss function and optimizer as described in step 4, pre-train the encoder using the contrastive loss function, and use unlabeled received signal data from different reference points Consider an encoded anchor is a real vector of dimension d, and a batch of encoded negative samples From the collection Suppose there is a positive sample k + =F(A i (Φ i )) matches q, which is a real vector of dimension d. It is the output of the anchor point after passing through the encoder. The contrast loss is a function. When q is equal to k + When it is similar to {k1, k2, k3, ...} but dissimilar to all other {k1, k2, k3, ...}, its value is low; using the dot product to measure the similarity, a form of contrast loss function is considered, called information noise contrast loss: where τ is a temperature hyperparameter calculated over one positive example and K negative examples. The loss is the logarithmic loss of a (K+1)-classification softmax classifier that tries to classify q as k. + .
Citation Information
Patent Citations
Text recognition system training method in self-supervised contrast learning natural scene
CN114973226A
Image classification method and system based on label propagation contrast semi-supervised learning
CN115410026A