Indoor positioning methods, devices and equipment
By using linear projection of channel state information and context feature extraction, combined with preset embedding mapping to a low-dimensional space, the problem of high data acquisition cost in large-scale indoor positioning scenarios is solved, and a high-precision and scalable indoor positioning system is realized.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- UNIV OF SCI & TECH OF CHINA
- Filing Date
- 2026-02-28
- Publication Date
- 2026-04-21
AI Technical Summary
Existing deep learning models have high data acquisition and annotation costs in large-scale indoor positioning scenarios, resulting in low positioning accuracy. Furthermore, traditional indoor positioning systems experience a decrease in positioning accuracy in non-line-of-sight situations, and their versatility and scalability are limited.
By acquiring channel state information, linear projection and contextual feature extraction are performed using a localization model. Combined with pre-defined classification embedding and location embedding, local embedding fusion is performed, mapping to a low-dimensional embedding space to determine location information, thereby reducing the overhead of the localization system and improving accuracy.
It achieves high-precision positioning in large-scale indoor positioning scenarios, has good scalability and adaptability, avoids dependence on the number of fixed receivers, and improves the recognition accuracy of the positioning system.
Smart Images

Figure CN121751329B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of signal processing technology, and more specifically to an indoor positioning method, apparatus, and device. Background Technology
[0002] Indoor Positioning Systems (IPS) aim to provide high-precision positioning services in indoor environments where GPS and other satellite positioning technologies have low or no positioning accuracy. With the development of computer technology, deep learning-based methods have been widely applied in indoor positioning, demonstrating higher positioning accuracy compared to traditional methods. However, in large-scale positioning scenarios, such as entire buildings or campuses, existing deep learning models require significant time and manpower for data collection and annotation, thus limiting the model's feature extraction capabilities and resulting in lower positioning accuracy. Summary of the Invention
[0003] In view of the above problems, the present invention provides an indoor positioning method, apparatus and device.
[0004] According to a first aspect of the present invention, an indoor positioning method is provided, comprising: acquiring multiple channel state information and a mapping relationship between location and channel state, wherein the channel state information represents the state information of the channel when a transmitter transmits a wireless signal to a receiver via a channel, and the mapping relationship between location and channel state includes multiple preset location embeddings and preset channel state embeddings corresponding to the preset location embeddings; for any channel state information, linearly projecting the channel state information using a positioning model to obtain an input sequence; extracting context features from the superimposed vector obtained by superimposing the input sequence, preset classification embeddings, and preset location embeddings to obtain a local embedding of any channel state information, wherein the preset classification embedding is used to extract preset location embeddings and the input sequence associated with other channel state information, and the other channel state information is used to indicate channel state information other than any channel state information among the multiple channel state information; determining a target channel state embedding from the multiple preset channel state embeddings based on the global embedding obtained by fusing the local embeddings of the multiple channel state information; and mapping the preset location embedding corresponding to the target channel state embedding to obtain location information for the transmitter.
[0005] A second aspect of the present invention provides an indoor positioning device, comprising: an acquisition module, configured to acquire multiple channel state information and a mapping relationship between location and channel state, wherein the channel state information represents the channel state information when a transmitter transmits a wireless signal to a receiver via a channel, and the mapping relationship between location and channel state includes multiple preset location embeddings and preset channel state embeddings corresponding to the preset location embeddings; a first obtaining module, configured to linearly project any channel state information using a positioning model to obtain an input sequence; a second obtaining module, configured to extract context features from a superimposed vector obtained by superimposing the input sequence, preset classification embeddings, and preset location embeddings to obtain a local embedding of any channel state information, wherein the preset classification embedding is used to extract preset location embeddings and the input sequence associated with other channel state information, and the other channel state information is used to indicate channel state information other than any channel state information among the multiple channel state information; a determination module, configured to determine a target channel state embedding from multiple preset channel state embeddings based on a global embedding obtained by fusing the local embeddings of the multiple channel state information; and a third obtaining module, configured to map the preset location embedding corresponding to the target channel state embedding to obtain location information for the transmitter.
[0006] A third aspect of the present invention provides an electronic device comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.
[0007] A fourth aspect of the present invention also provides a computer-readable storage medium having a computer program or instructions stored thereon, wherein the computer program or instructions, when executed by a processor, implement the steps of the above-described method.
[0008] A fifth aspect of the present invention also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method.
[0009] According to the indoor positioning method provided by the present invention, the preset channel state embedding is obtained by mapping preset channel state information to the target dimension. The global embedding is obtained by projecting the channel state information to the target dimension using a positioning model. The input sequence obtained by linear projection on the target dimension is superimposed with the preset classification embedding and the preset position embedding to obtain the superimposed vector. Then, the context of the superimposed vector is extracted to obtain the local embedding. Then, multiple local embeddings are fused together. Since the global embedding and the preset channel state embedding have the same dimension, that is, both are the target dimension, the channel state information of the real physical space dimension is mapped to the low-dimensional embedding space. Then, the mapping relationship between the position and the channel state in the low-dimensional embedding space is used to determine the position information of the transmitting end. The channel state information is mapped to the low-dimensional embedding space, thereby reducing the overhead of the positioning system. At the same time, it makes the channel state information of adjacent positions more distinguishable, avoids the dependence on the fixed number of receivers or the display coordinate information, has good scalability and adaptability for large-scale deployment, and improves the accuracy of indoor positioning. Attached Figure Description
[0010] The above-described features, other objects, and advantages of the present invention will become clearer from the following description of embodiments of the invention with reference to the accompanying drawings, in which:
[0011] Figure 1 An application scenario diagram of the indoor positioning method according to an embodiment of the present invention is shown.
[0012] Figure 2 A flowchart of an indoor positioning method according to an embodiment of the present invention is shown.
[0013] Figure 3 An architecture diagram of a positioning model according to an embodiment of the present invention is shown.
[0014] Figure 4 A flowchart illustrating a training method for a localization model according to an embodiment of the present invention is shown.
[0015] Figure 5 A flowchart of a training method for a localization model according to another embodiment of the present invention is shown.
[0016] Figure 6 A structural block diagram of an indoor positioning device according to an embodiment of the present invention is shown.
[0017] Figure 7 A block diagram of an electronic device suitable for implementing an indoor positioning method according to an embodiment of the present invention is shown. Detailed Implementation
[0018] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the invention. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the invention for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concept of the invention.
[0019] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0020] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0021] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0022] In the technical solution of this invention, the user information (including but not limited to user personal information, user image information, user device information, such as location information) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of related data all comply with relevant laws, regulations, and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entry points for users to choose to authorize or refuse.
[0023] In scenarios involving automated decision-making using personal information, the methods, devices, and systems provided in this invention offer users corresponding entry points for choosing to agree to or reject the automated decision-making results. If the user chooses to reject, the process proceeds to the expert decision-making stage. Here, "automated decision-making" refers to the activity of automatically analyzing and evaluating an individual's behavioral habits, interests, or economic, health, and credit status through computer programs, and then making a decision. Here, "expert decision-making" refers to the activity of making decisions by personnel who specialize in a particular field, possess specialized experience, knowledge, and skills, and have reached a certain level of professional expertise.
[0024] Indoor positioning systems (IPS) aim to provide high-precision positioning services in indoor environments where GPS and other satellite positioning technologies lack accuracy or fail. IPS is a crucial foundational task with significant value in commerce, retail, and inventory tracking. However, vision-based indoor positioning is susceptible to lighting conditions and presents privacy concerns; radar-based indoor positioning is costly to deploy. In contrast, widely deployed Wireless Fidelity (WiFi) devices are more cost-effective. The Channel State Information (CSI) in WiFi devices describes how the signal propagates from the transmitter to the receiver and represents the combined effects of scattering, fading, and power attenuation with distance. Related technologies utilize information such as Received Signal Strength (RSS) and Direction of Arrival (AoA) from the CSI to implement indoor positioning systems. These systems use signal processing methods to model and estimate the geometric parameters of the direct path, followed by triangulation or trilateration for position estimation. However, in non-line-of-sight (LOS) conditions, the positioning accuracy of these indoor positioning systems significantly decreases. As deep learning technology matures and achieves breakthroughs in image recognition and natural language processing with its superior performance, deep learning-based methods are also being applied to indoor positioning. Deep learning-based positioning methods can automatically extract features directly from raw signals, achieving higher positioning accuracy compared to traditional methods. However, deep learning-based positioning methods typically rely on large-scale labeled data to drive model training and learn effective feature representations. In a positioning scenario where most receivers can receive the wireless signal transmitted by the transmitter, this can be called a small-scale positioning scenario (typically with 3-4 receivers). In this scenario, each snapshot can be represented as a fixed-size, dense tensor. In a positioning scenario where only a small number of receivers near the transmitter can receive the wireless signal, this can be called a large-scale positioning scenario (with hundreds of receivers, for example, over 400 receivers). In this scenario, snapshots become a variable-size, sparse set, and the set of available receivers belongs to this set. Furthermore, the set of available receivers changes dynamically over time due to factors such as user movement, power control, and network scheduling. In large-scale positioning scenarios in practical applications, such as throughout an entire building or park, data collection and labeling require a significant investment of manpower and time, resulting in high costs and thus limiting the model's feature extraction capabilities.To alleviate the problem of high annotation costs, related technologies attempt to use auxiliary data, such as automatic collection by mobile robots, or user feedback information and transfer learning techniques to improve localization performance. However, these often require additional hardware support or rely on specific scenario assumptions, thus limiting their versatility and scalability in practical applications.
[0025] In view of this, embodiments of the present invention provide an indoor positioning method, comprising: acquiring multiple channel state information and a mapping relationship between location and channel state, wherein the channel state information represents the state information of the channel when the transmitter sends a wireless signal to the receiver via the channel, and the mapping relationship between location and channel state includes multiple preset location embeddings and preset channel state embeddings corresponding to the preset location embeddings; for any channel state information, linearly projecting the channel state information using a positioning model to obtain an input sequence; extracting context features from the superimposed vector obtained by superimposing the input sequence, preset classification embeddings, and preset location embeddings to obtain a local embedding of any channel state information, wherein the preset classification embedding is used to extract the preset location embeddings and the input sequence associated with other channel state information, and the other channel state information is used to indicate the channel state information other than any channel state information among the multiple channel state information; determining a target channel state embedding from the multiple preset channel state embeddings based on the global embedding obtained by fusing the local embeddings of the multiple channel state information; and mapping the preset location embedding corresponding to the target channel state embedding to obtain location information for the transmitter.
[0026] Figure 1 An application scenario diagram of the indoor positioning method according to an embodiment of the present invention is shown.
[0027] like Figure 1 As shown, application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0028] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).
[0029] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0030] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0031] It should be noted that the indoor positioning method provided in the embodiments of the present invention can generally be executed by server 105. Correspondingly, the indoor positioning device provided in the embodiments of the present invention can generally be installed in server 105. The indoor positioning method provided in the embodiments of the present invention can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the indoor positioning device provided in the embodiments of the present invention can also be installed in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.
[0032] It should be understood that Figure 1 The number of first terminal devices, second terminal devices, third terminal devices, networks, and servers shown in the diagram is merely illustrative. Depending on implementation needs, any number of first terminal devices, second terminal devices, third terminal devices, networks, and servers can be included.
[0033] Figure 2 A flowchart of an indoor positioning method according to an embodiment of the present invention is shown.
[0034] like Figure 2 As shown, the indoor positioning method 200 of this embodiment includes operations S210 to S250.
[0035] In operation S210, multiple channel state information and the mapping relationship between location and channel state are obtained.
[0036] In operation S220, for any channel state information, the localization model is used to linearly project any channel state information to obtain the input sequence.
[0037] In operation S230, context features are extracted from the superimposed vector obtained by superimposing the input sequence, preset classification embedding, and preset position embedding to obtain the local embedding of any channel state information.
[0038] In operation S240, the target channel state embedding is determined from multiple preset channel state embeddings based on the global embedding obtained by fusing the local embeddings of multiple channel state information.
[0039] In operation S250, the preset position embedding corresponding to the target channel state embedding is mapped to obtain the position information for the transmitting end.
[0040] Channel state information characterizes the state of the channel when a transmitter sends a wireless signal to a receiver via the channel. The transmitter may include one or more devices, and the receiver may include multiple devices. In one implementation, the transmitter can be a user device, such as a smartphone, laptop, or other location-enabled electronic device. In this case, A set of senders can be formed. ,in, Represents the set of senders The first sender in the process, Represents the set of senders The second sender in the process, Represents the set of senders The first in A transmitter. In one implementation, the receiver can be a Wireless Fidelity Access Point (AP), which includes... In this case, A set of receivers can be formed. ,in, Represents the set of receivers The first receiver in the process, Represents the set of receivers The second receiver in the process, Represents the set of receivers The first in One receiver.
[0041] In large-scale positioning scenarios, such as an entire building or campus, as the user moves within the positioning area carrying the transmitter, the receiving end communicating with the transmitter will also change. The wireless signal transmitted by the transmitter may only be collected by the receiving end. Some receivers in the signal receive signals; these receivers can be called usable receivers. Multiple usable receivers can form a set of usable receivers. Generally, the set of usable receivers is much smaller than the set of receivers. For example, in the receiver set... In this scenario, if the user's transmitter is located on the first floor of a building, the set of available receivers may be: If the user is carrying a transmitter located on the fifth floor of a building, the set of available receivers may be: .
[0042] The mapping relationship between location and channel state can include multiple preset location embeddings and their associated preset channel state embeddings. Preset location embeddings can be obtained by mapping preset location information to an embedding space of the target dimension. Preset channel state embeddings can be obtained by mapping preset channel state information to an embedding space of the target dimension.
[0043] By discretizing time, multiple discrete moments can be obtained. At moment t, the transmitter sends a wireless signal. Multiple receivers in the receiver set that can receive the wireless signal transmitted by the transmitter are considered available receivers. Each channel between the transmitter and each available receiver has its own channel state information; therefore, multiple channel state information values can be included. After sampling and detecting the channel using hardware devices, the corresponding channel state information measurement value can be obtained. The channel state information measurement value can include the channel state information measurement values of multiple available receivers communicating with a single transmitter. At moment t, the multiple channel state information measurement values associated with a single transmitter can constitute a snapshot. As shown in formula (1), each channel state information measurement value at time t can include any available receiver at time t and its corresponding communication mode. The communication mode can characterize various configuration factors of the positioning system, including the positioning model, during operation, such as antenna parameters, subcarrier allocation method, or transmission parameters. In actual positioning scenarios, snapshots have the characteristics of sparsity and variable size. Affected by user mobility, power control, and scheduling mechanisms, the communication mode will also change over time, resulting in the number and structure of channel state information contained in each snapshot being variable.
[0044] (1);
[0045] in, This represents the channel state information measurement value of any available receiver in the set of available receivers at time t. This represents the set of available receivers at time t. Let represent any available receiver in the set of available receivers at time t. Represents a set of preset communication modes. It represents any preset communication mode in the preset communication mode set.
[0046] For any channel state information among multiple channel state information, a positioning model can be used to perform linear projection on any channel state information to obtain an input sequence, including: for any channel state information, using a preset projection matrix in the positioning model to perform linear projection on the channel state information to obtain an input sequence.
[0047] A pre-defined projection matrix maps channel state information to an embedding space of a target dimension, where the target dimension is smaller than the dimension of the channel state information in the real physical space. This involves taking the channel state information measurement value for the r-th available receiver from the channel state information measurements. Construct the input matrix The input matrix includes input features corresponding to multiple subcarriers. Then, based on the preset projection matrix, the input features corresponding to each subcarrier in the input matrix are linearly projected to obtain the input sequence, as shown in formula (2).
[0048] (2);
[0049] in, Represents the input sequence. Let R represent the preset projection matrix, d represent the target dimension of the embedding space, and F represent the dimension of the features extracted from each subcarrier. This represents the input features corresponding to each subcarrier.
[0050] The predefined classification embedding can be used to extract a predefined position embedding and input sequence associated with other channel state information. The other channel state information can be used to indicate channel state information other than any single channel state information among multiple channel state information sets. In one implementation, the predefined classification embedding can be used to extract a predefined position embedding and input sequence associated with input features corresponding to other subcarriers. The input features corresponding to other subcarriers can be used to indicate input features corresponding to subcarriers in the input matrix corresponding to multiple subcarriers, excluding the input features corresponding to any single subcarrier.
[0051] The input sequence, the preset classification embedding, and the preset position embedding are superimposed to obtain the superimposed vector, as shown in formula (3).
[0052] (3);
[0053] in, Represents a superimposed vector. This indicates the embedding of a preset category. T represents the number of subcarriers. This represents the input feature corresponding to the first subcarrier in the input sequence. This represents the input feature corresponding to the T-th subcarrier in the input sequence. , This indicates that multiple preset positions are embedded. , This indicates the first preset position for embedding. This indicates that the second preset position is embedded. This indicates that the (T+1)th preset position is embedded.
[0054] In one implementation, context features are extracted from the superimposed vector to obtain the context features as shown in formula (4). Based on the effective mask, the output of the first input feature corresponding to the subcarrier can be used as a local embedding of any channel state information. As shown in formula (5).
[0055] (4);
[0056] (5);
[0057] in, Indicates encoder, Indicates a valid mask. The output represents the first input feature corresponding to the subcarrier.
[0058] The process of obtaining the local embeddings of other channel state information is similar to the process of obtaining the local embeddings of any channel state information described above, and will not be repeated here. Therefore, the set of local embeddings corresponding to all channel information in a snapshot is... As shown in formula (6).
[0059] (6);
[0060] in, This represents the channel state information measurement value for the r-th available receiver in the snapshot.
[0061] By fusing the local embeddings of multiple channel state information, a global embedding can be obtained. The dimension of the preset channel state embedding is the same as that of the global embedding, which is the target dimension. Based on the global embedding, the target state embedding is determined from multiple preset channel state embeddings, and then the target state embedding is projected. This projection process projects the target state embedding from the target dimension to the dimension of the real physical space, thereby obtaining the location information for the transmitting end.
[0062] The process described above maps channel state information to the embedding space of the target dimension using a localization model to obtain an input sequence. Contextual features are extracted from the superimposed vector obtained by superimposing the input sequence, preset classification embeddings, and preset location embeddings to obtain a local embedding for any channel state information. The process of fusing the local embeddings of multiple channel state information to obtain a global embedding can be called the unsupervised representation learning process for channel state information. The process of determining the target channel state embedding from multiple preset channel state embeddings based on the global embedding is called the fingerprint matching process. Through the unsupervised representation learning process and the fingerprint matching process, the target channel state embedding is determined. Mapping the target channel state embedding then yields the location information for the transmitting end.
[0063] Preset channel state embedding is obtained by mapping preset channel state information to the target dimension. Global embedding is obtained by projecting channel state information to the target dimension using a positioning model. Preset classification embedding and preset position embedding are superimposed on the input sequence obtained by linear projection on the target dimension to obtain a superimposed vector. Then, context extraction is performed on the superimposed vector to obtain local embeddings. Finally, multiple local embeddings are fused. Since global embedding and preset channel state embedding have the same dimension, that is, the target dimension, the channel state information of the real physical space dimension is mapped to the low-dimensional embedding space. Then, the mapping relationship between the position and the channel state in the low-dimensional embedding space is used to determine the position information of the transmitter. The channel state information is mapped to the low-dimensional embedding space, thereby reducing the overhead of the positioning system. At the same time, the channel state information of adjacent positions is more distinguishable, avoiding dependence on a fixed number of receivers or display coordinate information. It has good scalability and adaptability for large-scale deployment and improves the accuracy of indoor positioning.
[0064] Based on the global embedding obtained by fusing the local embeddings of multiple channel state information, the target channel state embedding is determined from multiple preset channel state embeddings, including: superimposing the preset embedding table with multiple local embeddings to obtain multiple local enhanced representations; superimposing the multiple local enhanced representations to obtain a context matrix; decoding the vector obtained after cross-attention processing of the query vector and the context matrix to obtain the global embedding of multiple channel state information; and determining the target channel state embedding from multiple preset channel state embeddings based on the similarity between the global embedding and the multiple preset channel state embeddings.
[0065] The preset embedding table can include a receiver identifier, which can be used to determine the receiver identifier for receiving wireless signals associated with channel state information. The receiver is used to receive wireless signals transmitted by the transmitter. Given that available receivers are determined, the channel state information measurements associated with them are also determined. Therefore, by assigning a receiver identifier to each receiver, the receiver identifier for receiving wireless signals associated with that channel state information can be determined based on the preset embedding table. It is represented by the following formula (7).
[0066] (7);
[0067] The preset embedding table is superimposed on multiple local embeddings to obtain multiple local augmented representations. The preset embedding table is superimposed on each local embedding to obtain the local augmented representation, as shown in formula (8).
[0068] (8);
[0069] in, This indicates local enhancement. This represents the number of the r-th receiver.
[0070] By superimposing multiple local enhancements, a context matrix can be obtained, as shown in formula (9).
[0071] (9);
[0072] in, Represents the context matrix. This represents the maximum number of receivers participating in the positioning calculation in a snapshot of the indoor positioning model.
[0073] The query vector represents the context features corresponding to each subcarrier in the context matrix. The context matrix can include context features corresponding to multiple subcarriers. Decoding the vector obtained after cross-attention processing using the query vector and the context matrix yields the global representation. The following formula (10) represents the global representation. The output of the context features corresponding to the first subcarrier is used as a global embedding of multiple channel state information. As shown in formula (11).
[0074] (10);
[0075] (11);
[0076] in, The decoder is represented by Q, which represents a preset feature vector based on the number of input features. Indicates the target mask. Represents global representation The output of the context features corresponding to the first subcarrier.
[0077] By superimposing a preset embedding table with multiple local embeddings, multiple local enhanced representations can be obtained. Then, by superimposing multiple local enhanced representations, a context matrix can be obtained. Finally, by decoding the vector obtained after cross-attention processing of the query vector and the context matrix, a global embedding of multiple channel state information can be obtained. This enables the channel state information to be mapped to a low-dimensional embedding space, thereby reducing the overhead of the positioning system and making the channel state information of adjacent locations more distinguishable, thus improving positioning accuracy.
[0078] Determining a target channel state embedding from multiple preset channel state embeddings based on the similarity between the global embedding and multiple preset channel state embeddings includes: sorting the similarity between the global embedding and multiple preset channel state embeddings; determining at least one target similarity based on a preset similarity threshold; and determining the target channel state embedding from multiple preset channel state embeddings based on at least one target similarity.
[0079] The mapping relationship between location and channel state can include multiple preset location embeddings and corresponding preset channel state embeddings. The preset location embeddings, preset channel state embeddings, and global embeddings are all mapped to a low-dimensional embedding space of the target dimension. Mapping the preset location embeddings and preset channel state embeddings to this low-dimensional embedding space yields the mapping relationship between location and channel state. The following formula (12) can be used to abstract the process of mapping the snapshot to a low-dimensional embedding space using the positioning model to obtain the global embedding as follows (13). By comparing the similarity between the global embedding and multiple preset channel state embeddings and sorting the similarities, a similarity ranking can be obtained. Based on a preset similarity threshold, at least one target similarity can be determined from multiple similarities.
[0080] (12);
[0081] (13);
[0082] in, This represents a preset snapshot corresponding to multiple preset channel state information. Indicates that a preset snapshot will be used. or snapshot The representation function mapped to the low-dimensional embedding space. This indicates the embedding of the preset channel state. Indicates embedding at a preset position. This indicates the mapping relationship between location and channel state.
[0083] In addition, the nearest neighbor method can be used to determine k target preset channel state embeddings from multiple preset channel state embeddings based on similarity. As shown in formula (14). In one implementation, the similarity can be cosine similarity. In another implementation, k can be 3.
[0084] (14);
[0085] in, A similarity measure representing a low-dimensional embedding space. This indicates the embedding of the i-th preset channel state. This indicates the embedding at the i-th preset position.
[0086] By comparing the similarity between the global embedding and multiple preset channel state embeddings, and determining the target preset channel state embedding based on a preset similarity threshold, the location information for the transmitter is obtained based on the target preset channel state embedding. This avoids dependence on a fixed number of receivers or display coordinate information, and has good scalability and adaptability for large-scale deployment.
[0087] Based on multiple channel state information, the snapshot is mapped to a low-dimensional embedding space to obtain a global embedding, which satisfies the following formula (15). The global embedding is compared with the multiple preset channel state embeddings in the mapping relationship between location and channel state to obtain the target channel state embedding. Based on the target channel state embedding, the location information for the transmitting end is obtained, realizing the mapping from the snapshot to the real physical space, which satisfies the following formula (16).
[0088] (15);
[0089] (16);
[0090] in, Indicates snapshot Global embedding in a low-dimensional embedding space This represents all available snapshots for multiple channel state information. This represents the mapping function that maps snapshots to the real physical space. The dimension representing the real physical space, such as two-dimensional or three-dimensional. This indicates the location information of the sending end in the real physical space.
[0091] Figure 3An architecture diagram of a positioning model according to an embodiment of the present invention is shown.
[0092] like Figure 3 As shown, the localization model includes an observation encoder and a localization decoder, which calculates the channel state information for the r-th available receiver by taking the channel state information measurement values from the channel state information measurements. The input observation encoder is used to linearly project the channel state information measurements using its linear projection layer, mapping the high-dimensional channel state information measurements to a target-dimensional input sequence. The input sequence, a preset classification embedding (CLS), and a preset position embedding are then superimposed to obtain a superimposed vector. This vector is then used for self-attention computation via a multi-head attention mechanism in the observation encoder's attention layer to learn the correlation features between multiple subcarriers. Finally, the connection and normalization layers of the observation encoder are utilized. Residual connections and layer normalization are performed to avoid gradient vanishing. Then, a nonlinear transformation is applied using the feedforward network layer of the observation encoder to obtain local embeddings. Multiple connection and normalization layers can be included to enhance feature extraction capabilities. These multiple local embeddings output by the observation encoder are input to the localization decoder. A pre-defined embedding table is superimposed on each of the multiple local embeddings to obtain multiple enhanced local representations. Then, the attention layer of the localization decoder performs self-attention and cross-attention processing on the query vector and the context matrix composed of these enhanced local representations. The vector obtained after self-attention and cross-attention processing is then decoded to obtain the global representation. The connection and normalization layers of the localization decoder... Similar to encoders, it will not be described in detail here.
[0093] Figure 4 A flowchart illustrating a training method for a localization model according to an embodiment of the present invention is shown.
[0094] like Figure 4 As shown, the training method 400 for the localization model in this embodiment includes operations S410 to S470.
[0095] During operation of S410, multiple sample channel status information are acquired.
[0096] In operation S420, based on the target masking strategy, multiple target sample channel state information is determined from multiple sample channel state information.
[0097] In operation S430, channel state information of multiple target samples is input into the initial localization model to obtain the first sample global embedding for the channel state information of multiple target samples.
[0098] During operation of S440, the model parameters of the initial positioning model are adjusted using momentum update parameters to obtain the adjusted initial positioning model.
[0099] In operation S450, channel state information of multiple target samples is input into the adjusted initial positioning model to obtain the second sample global embedding for the channel state information of multiple target samples.
[0100] In operation S460, the first sample global embedding and the second sample global embedding are used as positive sample pairs, and the positive sample pairs and the preset negative sample global embedding sequence are processed using the contrastive loss function to obtain the loss value.
[0101] In operation S470, the network parameters of the initial localization model are iteratively adjusted using the loss value to obtain the localization model.
[0102] Each transmitter can communicate with multiple receivers, thus corresponding to multiple sample channel state information. Therefore, the sample channel state information can include multiple values. Based on a masking strategy, the multiple sample channel state information are masked, and the masked sample channel state information is used as the target sample channel state information. Therefore, multiple target sample channel state information can be obtained.
[0103] By inputting the target sample channel state information into the initial localization model, a first sample global embedding for the target sample channel state information can be obtained. As shown in formula (17). The model parameters of the initial positioning model are adjusted using the momentum update parameters to obtain the adjusted initial positioning model. The following formula (18) is satisfied. The target sample channel state information is input into the adjusted initial positioning model to obtain the second sample global embedding for the target sample channel state information. The following formula (19) is used. The global embeddings of the first and second samples are taken as positive sample pairs, and the contrastive loss function is used as shown in the following formula (20). The loss value is obtained by processing positive sample pairs and the preset negative sample global embedding sequence.
[0104] (17);
[0105] (18)
[0106] (19);
[0107] (20);
[0108] in, This indicates the channel state information of the first target sample. This indicates the channel state information of the second target sample. and These are channel state information of two target samples at the same location at different times. The basic representation of the global embedding of the first sample. The basic representation representing the global embedding of the second sample. This represents the momentum update factor. express Normalization This represents the model parameters of the initial localization model. This represents the model parameters of the adjusted initial positioning model. This refers to the projection head that is bound to the initial positioning model. This indicates the projection head that is bound to the adjusted initial positioning model. Indicates the temperature coefficient. Let K represent the global embedding of the i-th negative sample, and K represent the number of global embeddings of negative samples. This represents the transpose of the global embedding of the first sample.
[0109] In some implementations, the negative sample global embedding sequence can be preset, or it can be added to the queue in each iteration, while the earliest first or second sample global embedding is removed and used as the negative sample global embedding in the negative sample global embedding queue. The negative sample global embedding queue is a first-in-first-out queue.
[0110] Multiple target sample channel state information can be obtained by masking the acquired sample channel state information. Then, the channel state information of multiple target samples is input into the initial positioning model to obtain the first sample global embedding for the channel state information of multiple target samples. The model parameters of the initial positioning model are then adjusted using momentum update parameters to obtain the adjusted initial positioning model. The channel state information of multiple target samples is input into the adjusted initial positioning model to obtain the second sample global embedding for the channel state information of multiple target samples. The first sample global embedding and the second sample global embedding are used as positive sample pairs, and the positive sample pairs and the preset negative sample global embedding sequence are processed using a contrastive loss function to obtain the loss value. The network parameters of the initial positioning model are iteratively adjusted using the loss value to obtain the positioning model. Since it does not rely on the location information of the receiver, building floor plan or other topology information, the additional acquisition cost is reduced.
[0111] Inputting channel state information of multiple target samples into an initial localization model to obtain a first global embedding of the channel state information of multiple target samples can include: using the learnable projection matrix of the initial localization model to perform linear projection on the channel state information of any target sample among the multiple target sample channel state information to obtain a sample input sequence; extracting context features from the sample superposition vector obtained by superimposing the sample input sequence, the learnable classification embedding, and the learnable position embedding to obtain a sample local embedding of the channel state information of any target sample; and fusing the sample local embeddings of the channel state information of multiple target samples to obtain the first global embedding of the sample.
[0112] The localization model may include a preset projection matrix, preset location embeddings, preset classification vectors, and a preset embedding table. The initial localization model may include a learnable projection matrix, learnable location embeddings, learnable classification vectors, and a learnable embedding table. The learnable projection matrix, learnable location embeddings, learnable classification vectors, and learnable embedding table are gradually adjusted during the localization model's training process. Once the localization model is trained, the preset projection matrix, preset location embeddings, preset classification vectors, and preset embedding table are obtained.
[0113] The training process of the indoor positioning model will be described in detail below.
[0114] A learnable projection matrix can map the channel state information of a target sample to a low-dimensional embedding space. By using the learnable projection matrix of the initial localization model, a linear projection can be performed on the channel state information of any target sample from multiple target sample channel state information to obtain the sample input sequence.
[0115] A sample input matrix can be constructed from the channel state information measurements of the target sample channel state information for each available receiver. This sample input matrix includes sample input features corresponding to multiple subcarriers. Then, a linear projection is performed on the sample input features corresponding to each subcarrier in the sample input matrix based on a learnable projection matrix to obtain a sample input sequence. A learnable classification embedding is used to extract a preset position embedding and sample input sequence associated with other target sample channel state information. The other target sample channel state information is used to indicate the target sample channel state information other than any single target sample channel state information among the multiple target sample channel state information. In one implementation, the learnable classification embedding can be used to extract a preset position embedding and sample input sequence associated with sample input features corresponding to other subcarriers. The sample input features corresponding to other subcarriers can be used to indicate the sample input features corresponding to subcarriers in the input matrix corresponding to multiple subcarriers, excluding the input features corresponding to any single subcarrier.
[0116] By extracting context features from the superimposed sample vector obtained by superimposing the sample input sequence, learnable classification embedding, and learnable position embedding, a sample local embedding of the channel state information of any target sample can be obtained. The first global embedding of the target sample channel state information is obtained by fusing the sample local embeddings of multiple target samples. This process can include: superimposing the learnable embedding table with multiple sample local embeddings to obtain multiple sample local enhanced representations; superimposing the multiple sample local enhanced representations to obtain a sample context matrix; and decoding the vector obtained after cross-attention processing of the sample query vector and the sample context matrix to obtain the first global embedding of the channel state information of multiple target samples.
[0117] Learnable embedding tables can be used to determine the transmitter number for transmitting target sample channel state information. By superimposing the learnable embedding table with multiple sample local embeddings, multiple sample local augmented representations can be obtained. By superimposing multiple sample local augmentations, the sample context matrix can be obtained.
[0118] The sample query vector represents the sample context features corresponding to each subcarrier in the sample context matrix. The sample context matrix can include sample context features corresponding to multiple subcarriers. Decoding the vector obtained after cross-attention processing of the sample query vector and the sample context matrix yields the global sample representation. The output of the sample context features corresponding to the first subcarrier in the global sample representation is used as the global embedding of channel state information for multiple samples.
[0119] The above-described process of processing the channel state information of multiple target samples using the learnable projection matrix of the initial localization model yields a first global embedding for the channel state information of multiple target samples. The process of fusing these first global embeddings to obtain the first global embedding can be termed the unsupervised representation learning process for sample channel state information. Since unsupervised representation learning can extract spatially proximate feature representations without requiring extensive annotation, it avoids dependence on a fixed number of receivers or explicit coordinate information, and exhibits good scalability and adaptability for large-scale deployments.
[0120] Based on a target masking strategy, multiple target sample channel state information is determined from multiple sample channel state information, including: a first target masking strategy characterized by selecting sample channel state information associated with non-overlapping receivers for adjacent time periods to form a first target sample channel state information set; a second target masking strategy characterized by selecting sample channel state information associated with randomly selected receivers from the multiple sample channel state information to form a second target sample channel state information set; and determining multiple target sample channel state information from at least one of the first target masking strategy and the second target masking strategy and a preset probability.
[0121] The target masking strategy may include a first target masking strategy and a second target masking strategy. The first target masking strategy can be characterized by selecting sample channel state information associated with non-overlapping receivers for adjacent time periods to form a first target sample channel state information set; the second target masking strategy can be characterized by selecting sample channel state information associated with randomly selected receivers from multiple sample channel state information sets to form a second target sample channel state information set.
[0122] Based on at least one of the first target masking strategy and the second target masking strategy and a preset probability, multiple target sample channel state information are determined from at least one of the first target sample channel state information set or the second target sample channel state information set. Sample snapshots can be constructed based on the target sample channel state information. The sample snapshot based on the first target masking strategy satisfies the following formula (21), and the sample snapshot based on the second target masking strategy satisfies the following formula (22).
[0123] (twenty one);
[0124] (twenty two);
[0125] Where t represents time t, Represents the adjacent times of time t. This represents a snapshot of the samples at time t. Only the set of available receivers is retained. A sample snapshot composed of relevant sample channel state information. Indicates will Snapshot of samples at time Only the set of available receivers is retained. A snapshot composed of relevant sample channel state information; Represents a snapshot of the sample from time t. Randomly select the set of available receivers A sample snapshot composed of relevant sample channel state information. Indicates will Snapshot of samples at time Randomly select the set of available receivers A sample snapshot composed of relevant sample channel state information.
[0126] Based on at least one of a first target masking strategy and a second target masking strategy, and a preset probability, multiple target sample channel state information are determined from at least one of a first target sample channel state information set or a second target sample channel state information set. In one implementation, a masking strategy combining the first and second target masking strategies can be used. The first target masking strategy is selected with a preset probability α, and the second target masking strategy is selected with 1-α to mask the sample channel state information, thereby obtaining the target sample channel state information.
[0127] By masking the sample channel state information using a target masking strategy, the problem of shortcut learning can be effectively alleviated while ensuring that the localization model learns a robust spatial representation, thereby significantly improving the localization performance of the localization model.
[0128] Figure 5 A flowchart of a training method for a localization model according to another embodiment of the present invention is shown.
[0129] like Figure 5 As shown, this is a snapshot of samples taken from adjacent time points for the same user. and sample snapshot By using either the first target masking strategy or the second target masking strategy for masking, a sample snapshot can be determined. , , or Sample snapshot or It can be used This indicates a sample snapshot. or It can be used This indicates that the initial localization model is used to process sample snapshots. Generate the first sample global embedding q, and process the sample snapshot using the adjusted initial localization model. Generate the first sample global embedding The default global embedding sequence for negative samples is... The preset negative sample global embedding sequence can be based on the first sample global embedding. What is obtained is a first-in-first-out queue, where the global embedding q of the first sample is compared with the global embedding q of the first sample based on the contrastive loss function value. The similarity between the model and the preset negative sample global embedding sequence is maximized, and the similarity between the model and the preset negative sample global embedding sequence is minimized. The network parameters of the initial localization model are iteratively adjusted to obtain the localization model.
[0130] To verify the effectiveness of the indoor positioning method proposed in this application, the following detailed explanation is based on three datasets and seven indoor positioning scenarios as shown in Tables 1 and 2. These seven indoor positioning scenarios cover small-scale scenarios with 3-4 receivers and large-scale scenarios with hundreds of receivers. The small-scale positioning scenario experiments included six regional-level small scenarios, such as laboratories, conference rooms, and offices, with a relatively small number of receivers. The large-scale positioning scenario experiments included a large building with a total area of approximately 5100 square meters spanning four floors. Data collection was conducted using robots and handheld devices, employing 10 commercial smartphones and a leave-one-device cross-validation protocol to simulate the situation in real-world deployments where the positioning model needs to generalize on new, unseen devices.
[0131] Table 1 shows that the unlabeled data for scenarios S1 and S2 were collected using a robot, while the unlabeled data for scenarios S1-S7 were collected using a handheld device. The unlabeled data for scenarios S1, S2, and S7 is split based on the collection date, while the unlabeled data for scenarios S3-S6 is split based on user ID. The training / test set split is performed as follows: for example, July 28th for scenario S1 means using data from July 28th as training data and the remaining data as test data; similarly, for scenario S3, user IDs 4-6 / 7 means using data from user IDs 4-6 as training data and data from user ID 7 as test data.
[0132] Scenarios S1 to S7 in Table 2 correspond to S1 to S7 in Table 1. In Table 2, the scenario scale S represents a small-scale positioning scenario, M represents a medium-scale positioning scenario, L represents a large-scale positioning scenario, Sifi represents the positioning error of the supervised positioning method based on Sifi in related technologies, LocGPT represents the positioning error of the supervised positioning method based on LocGPT in related technologies, RobLoc represents the positioning error of the supervised positioning method proposed in this application, and the supervised error reduction ratio represents the relative reduction ratio of the positioning error based on the supervised positioning method proposed in this application relative to the minimum error in that scenario. CSS represents the positioning error of the CSS-based unsupervised positioning method in related technologies; AFN represents the positioning error of the AFN-based unsupervised positioning method in related technologies; Glow represents the positioning error of the Glow-based unsupervised positioning method in related technologies; ULoc represents the positioning error of the unsupervised positioning method proposed in this application; relative error difference represents the relative difference of the ULoc-based unsupervised positioning method relative to the minimum error in this scenario; unsupervised error reduction ratio represents the relative reduction ratio of the positioning error of the proposed unsupervised positioning method relative to the minimum error in this scenario; underlined data represents the minimum positioning error in a scenario of the same scale; and bold data represents the positioning error of the best positioning method in this scenario.
[0133] As shown in Tables 1 and 2, the unsupervised localization method ULoc proposed in this application reduces the average localization error by 10.4% compared to the best unsupervised localization method in related technologies. Furthermore, the ULoc architecture reduces the average localization error by 9.8% compared to the best unsupervised localization method in related technologies. More importantly, even with scarce labeled samples, the proposed unsupervised localization method ULoc maintains high localization performance, with its average localization error reduced by 10.5% compared to the best supervised localization method in related technologies. This indicates that the proposed unsupervised localization method ULoc has good applicability in both small-scale and large-scale localization scenarios and can achieve robust generalization to unseen devices in real-world environments.
[0134] Table 1 Data Collection Table under Different Positioning Scenarios
[0135]
[0136] Table 2 Comparison of positioning errors in different positioning scenarios
[0137]
[0138] Based on the above-described indoor positioning method, this invention also provides an indoor positioning device. The following will be combined with... Figure 6 The device is described in detail.
[0139] Figure 6 A structural block diagram of an indoor positioning device according to an embodiment of the present invention is shown.
[0140] like Figure 6 As shown, the indoor positioning device 600 of this embodiment includes an acquisition module 610, a first acquisition module 620, a second acquisition module 630, a determination module 640, and a third acquisition module 650.
[0141] The acquisition module 610 is used to acquire multiple channel state information and the mapping relationship between location and channel state. The channel state information represents the state of the channel when the transmitting end transmits a wireless signal to the receiving end via the channel. The mapping relationship between location and channel state includes multiple preset location embeddings and preset channel state embeddings corresponding to the preset location embeddings. In one embodiment, the acquisition module 610 can be used to perform the operation S210 described above, which will not be repeated here.
[0142] The first obtaining module 620 is used to linearly project any channel state information using a positioning model to obtain an input sequence. In one embodiment, the first obtaining module 620 can be used to perform the operation S220 described above, which will not be repeated here.
[0143] The second obtaining module 630 is used to extract context features from the superimposed vector obtained by superimposing the input sequence, the preset classification embedding, and the preset position embedding, to obtain a local embedding of any channel state information. The preset classification embedding is used to extract the preset position embedding and the input sequence associated with other channel state information, and the other channel state information is used to indicate the channel state information other than any single channel state information among multiple channel state information. In one embodiment, the second obtaining module 630 can be used to perform the operation S230 described above, which will not be repeated here.
[0144] The determining module 640 is used to determine the target channel state embedding from multiple preset channel state embeddings based on the global embedding obtained by fusing the local embeddings of multiple channel state information. In one embodiment, the determining module 640 can be used to perform the operation S240 described above, which will not be repeated here.
[0145] The third obtaining module 650 is used to map the preset position embedding corresponding to the target channel state embedding to obtain the position information for the transmitting end. In one embodiment, the third obtaining module 650 can be used to perform the operation S250 described above, which will not be repeated here.
[0146] According to an embodiment of the present invention, the determining module 640 includes: a first determining submodule, configured to superimpose a preset embedding table with multiple local embeddings to obtain multiple local enhanced representations, wherein the preset embedding table includes a receiver identifier; a second determining submodule, configured to superimpose the multiple local enhanced representations to obtain a context matrix; a third determining submodule, configured to decode the vector obtained after cross-attention processing using a query vector and a context matrix to obtain a global embedding of multiple channel state information; and a fourth determining submodule, configured to determine a target channel state embedding from the multiple preset channel state embeddings based on the similarity between the global embedding and the multiple preset channel state embeddings.
[0147] According to an embodiment of the present invention, the fourth determining submodule includes: a first determining unit, configured to sort the similarity between the global embedding and multiple preset channel state embeddings, and determine at least one target similarity based on a preset similarity threshold; and a second determining unit, configured to determine a target channel state embedding from multiple preset channel state embeddings based on at least one target similarity.
[0148] According to an embodiment of the present invention, the first obtaining module 620 includes: a first obtaining submodule, used to linearly project the channel state information using a preset projection matrix in the positioning model for any channel state information to obtain an input sequence, wherein the preset projection matrix is used to map the channel state information to a target dimension.
[0149] According to embodiments of the present invention, any plurality of modules among the acquisition module 610, the first obtaining module 620, the second obtaining module 630, the determining module 640, and the third obtaining module 650 may be combined into one module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in one module. According to embodiments of the present invention, at least one of the acquisition module 610, the first obtaining module 620, the second obtaining module 630, the determining module 640, and the third obtaining module 650 may be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any one of the three implementation methods or a suitable combination of any of them. Alternatively, at least one of the acquisition module 610, the first acquisition module 620, the second acquisition module 630, the determination module 640, and the third acquisition module 650 may be implemented at least partially as a computer program module, which can perform corresponding functions when the computer program module is run.
[0150] Figure 7 A block diagram of an electronic device suitable for implementing an indoor positioning method according to an embodiment of the present invention is shown.
[0151] like Figure 7 As shown, an electronic device 700 according to an embodiment of the present invention includes a processor 701, which can perform various appropriate actions and processes according to a program stored in ROM 702 or a program loaded from storage portion 708 into RAM 703. The processor 701 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 701 may also include onboard memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0152] RAM 703 stores various programs and data required for the operation of electronic device 700. Processor 701, ROM 702, and RAM 703 are interconnected via bus 704. Processor 701 executes various operations of the method flow according to embodiments of the present invention by executing programs in ROM 702 and / or RAM 703. It should be noted that the programs may also be stored in one or more memories other than ROM 702 and RAM 703. Processor 701 may also execute various operations of the method flow according to embodiments of the present invention by executing programs stored in said one or more memories.
[0153] According to an embodiment of the present invention, the electronic device 700 may further include an input / output (I / O) interface 705, which is also connected to a bus 704. The electronic device 700 may also include one or more of the following components connected to the input / output (I / O) interface 705: an input section 706 including a keyboard, mouse, etc.; an output section 707 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the input / output (I / O) interface 705 as needed. A removable medium 711, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 710 as needed so that computer programs read from it can be installed into the storage section 708 as needed.
[0154] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of the present invention.
[0155] According to embodiments of the present invention, a computer-readable storage medium may be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, RAM (Random Access Memory), ROM (Read Only Memory), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of the present invention, a computer-readable storage medium may include ROM 702 and / or RAM 703 and / or one or more memories other than ROM 702 and RAM 703 described above.
[0156] Embodiments of the present invention also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to enable the computer system to implement the indoor positioning method provided in the embodiments of the present invention.
[0157] When the computer program is executed by the processor 701, it performs the functions defined in the system / apparatus of this invention. According to embodiments of the invention, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0158] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 709, and / or installed from a removable medium 711. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.
[0159] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 709, and / or installed from the removable medium 711. When the computer program is executed by the processor 701, it performs the functions defined in the system of this embodiment of the invention. According to embodiments of the invention, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0160] According to embodiments of the present invention, program code for executing the computer programs provided in the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0161] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0162] Those skilled in the art will understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, the features described in the various embodiments of the present invention can be combined and / or combined in various ways without departing from the spirit and teachings of the present invention. All such combinations and / or combinations fall within the scope of the present invention.
[0163] The embodiments of the present invention have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of the invention. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of the invention, and all such substitutions and modifications should fall within the scope of the invention.
Claims
1. An indoor positioning method, characterized in that, The method includes: Multiple channel state information and the mapping relationship between location and channel state are obtained. The channel state information represents the channel state information when the transmitting end sends a wireless signal to the receiving end through the channel. The mapping relationship between location and channel state includes multiple preset location embeddings and preset channel state embeddings corresponding to the preset location embeddings. For any channel state information, the localization model is used to perform a linear projection on the channel state information to obtain the input sequence. Context features are extracted from the superimposed vector obtained by superimposing the input sequence, the preset classification embedding, and the preset position embedding to obtain the local embedding of any channel state information. The preset classification embedding is used to extract the preset position embedding and input sequence associated with other channel state information. The other channel state information is used to indicate the channel state information other than any channel state information among the multiple channel state information. Based on the global embedding obtained by fusing the local embeddings of multiple channel state information, the target channel state embedding is determined from multiple preset channel state embeddings. The preset location embedding corresponding to the target channel state embedding is mapped to obtain the location information for the transmitting end.
2. The method according to claim 1, characterized in that, The step of determining the target channel state embedding from multiple preset channel state embeddings based on the global embedding obtained by fusing the local embeddings of multiple channel state information includes: The preset embedding table is superimposed on multiple local embeddings to obtain multiple local enhanced representations, wherein the preset embedding table includes a receiver identifier; The multiple local enhancement representations are superimposed to obtain the context matrix; The vector obtained after cross-attention processing of the query vector and the context matrix is decoded to obtain the global embedding of multiple channel state information. The target channel state embedding is determined from the multiple preset channel state embeddings based on the similarity between the global embedding and the multiple preset channel state embeddings.
3. The method according to claim 2, characterized in that, The step of determining the target channel state embedding from the multiple preset channel state embeddings based on the similarity between the global embedding and the multiple preset channel state embeddings includes: The similarity between the global embedding and the multiple preset channel state embeddings is sorted, and at least one target similarity is determined based on a preset similarity threshold; The target channel state embedding is determined from multiple preset channel state embeddings based on the at least one target similarity.
4. The method according to claim 3, characterized in that, For any given channel state information, the localization model is used to perform a linear projection on the given channel state information to obtain an input sequence, including: For any channel state information, the channel state information is linearly projected using a preset projection matrix in the positioning model to obtain the input sequence, wherein the preset projection matrix is used to map the channel state information to the target dimension.
5. The method according to claim 1, characterized in that, The localization model was trained using the following operations: Acquire channel status information for multiple samples; Based on the target masking strategy, multiple target sample channel state information is determined from the multiple sample channel state information; The channel state information of the multiple target samples is input into the initial localization model to obtain the first sample global embedding for the channel state information of the multiple target samples. The model parameters of the initial positioning model are adjusted using momentum update parameters to obtain the adjusted initial positioning model; The channel state information of the multiple target samples is input into the adjusted initial positioning model to obtain the second sample global embedding for the channel state information of the multiple target samples. The first sample global embedding and the second sample global embedding are used as positive sample pairs, and the positive sample pairs and the preset negative sample global embedding sequence are processed using the contrastive loss function to obtain the loss value; The network parameters of the initial localization model are iteratively adjusted using the loss value to obtain the localization model.
6. The method according to claim 5, characterized in that, The localization model includes a preset projection matrix, a preset location embedding, a preset classification vector, and a preset embedding table; The initial localization model includes a learnable projection matrix, a learnable location embedding, a learnable classification vector, and a learnable embedding table; Specifically, the channel state information of the multiple target samples is input into the initial localization model to obtain a first sample global embedding for the channel state information of the multiple target samples, including: Using the learnable projection matrix of the initial positioning model, a linear projection is performed on the channel state information of any one of the multiple target sample channel state information to obtain the sample input sequence; Context feature extraction is performed on the sample superimposed vector obtained by superimposing the sample input sequence, learnable classification embedding, and learnable position embedding to obtain the sample local embedding of the channel state information of any target sample. The learnable classification embedding is used to extract the preset position embedding and sample input sequence associated with the channel state information of other target samples. The channel state information of other target samples is used to indicate the channel state information of target samples other than the channel state information of any target sample among the multiple target sample channel state information. The local embeddings of the channel state information of multiple target samples are fused to obtain the global embedding of the first sample.
7. The method according to claim 6, characterized in that, The step of fusing the local embeddings of the channel state information of multiple target samples to obtain the first global embedding includes: The learnable embedding table is superimposed with multiple sample local embeddings to obtain multiple sample local enhanced representations, wherein the learnable embedding table is used to determine the transmitter number for transmitting the target sample channel state information; The local augmented representations of the multiple samples are superimposed to obtain the sample context matrix; The vector obtained by performing cross-attention processing on the query vector and the sample context matrix is decoded to obtain the first sample global embedding of channel state information of multiple target samples.
8. The method according to claim 6, characterized in that, The target masking strategy includes a first target masking strategy and a second target masking strategy, wherein determining multiple target sample channel state information from the multiple sample channel state information based on the target masking strategy includes: The first target masking strategy represents the selection of sample channel state information associated with non-overlapping receivers for adjacent time periods to form a first target sample channel state information set. The second target masking strategy characterizes the sample channel state information associated with the randomly selected receiver among the multiple sample channel state information to form a second target sample channel state information set; Based on at least one of the first target masking strategy and the second target masking strategy and a preset probability, a plurality of target sample channel state information is determined from at least one of the first target sample channel state information set or the second target sample channel state information set.
9. An indoor positioning device, characterized in that, The device includes: The acquisition module is used to acquire multiple channel state information and the mapping relationship between location and channel state. The channel state information represents the channel state information when the transmitting end transmits a wireless signal to the receiving end through the channel. The mapping relationship between location and channel state includes multiple preset location embeddings and preset channel state embeddings corresponding to the preset location embeddings. The first obtaining module is used to linearly project any channel state information using a positioning model to obtain an input sequence. The second obtaining module is used to extract context features from the superimposed vector obtained by superimposing the input sequence, the preset classification embedding, and the preset position embedding to obtain the local embedding of any channel state information. The preset classification embedding is used to extract the preset position embedding and the input sequence associated with other channel state information. The other channel state information is used to indicate the channel state information other than any channel state information among the multiple channel state information. The determination module is used to determine the target channel state embedding from multiple preset channel state embeddings based on the global embedding obtained by fusing the local embeddings of multiple channel state information. The third module is used to map the preset position embedding corresponding to the target channel state embedding to obtain the position information for the transmitting end.
10. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Positioning method based on mixed fingerprints and line-of-sight / non-line-of-sight recognition
CN117278934A
Indoor positioning method and device and readable storage medium
CN117354917A