Perception positioning method and electronic equipment

By using beamforming information instead of channel state information, combined with lightweight models and meta-learning architecture, the problems of large amount of deep model parameters and poor environmental adaptability are solved, and efficient and accurate close-range positioning is achieved on resource-constrained devices.

CN120302235BActive Publication Date: 2025-08-26BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510773391.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-08-26
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

The existing machine learning-based perceptual positioning method has a large amount of deep model parameters, which makes it difficult to promote on ordinary devices with resource limitations, and lacks the ability to quickly adapt to environmental changes, resulting in a decrease in positioning accuracy.

Method used

Beamforming information is used instead of channel state information, and trained and deployed through a lightweight positioning model, combining the domain enhancement meta-learning architecture and online fine-tuning mechanism to achieve real-time processing of beamforming information and ensure positioning accuracy.

Benefits of technology

It significantly reduces the burden of data processing and storage, improves the scalability and positioning accuracy of close-aware positioning, and is suitable for devices with resource-constrained, and has the ability to quickly adapt to environmental changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302235B_ABST
    Figure CN120302235B_ABST
Patent Text Reader

Abstract

The present application provides a perception positioning method and electronic device, which relate to the field of positioning technology. The perception positioning method is applied to a wireless access terminal with a beamforming function turned on in a target scene, and the target scene contains a target object to be located. The beamforming information fed back to the access terminal by a fixed terminal connected to the wireless access terminal in the target scene is collected. The beamforming information is input into a trained positioning model, and the positioning position of the target object to be located is output through the positioning model. The trained positioning model can accurately predict the positioning position of the target object. The beamforming information contains a low amount of data parameters, but the information granularity is fine, which can take into account the real-time nature of position prediction and the perception positioning accuracy. The data storage space and transmission bandwidth are significantly compressed, while the burden of data processing and positioning position prediction is reduced, and the scalability of short-range perception positioning deployment is improved, which is particularly suitable for deployment on resource-constrained devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of positioning technology, and in particular to a perception positioning method and electronic equipment. Background Art

[0002] Wireless signal sensing and positioning can be applied in a variety of fields, including energy-efficient communications and health monitoring. Positioning efficiency and accuracy directly impact the experience of integrating the physical and digital worlds. Satellite positioning technology provides good accuracy outdoors, but accuracy is difficult to guarantee indoors. With the development of intelligent society, the demand for generalized, low-cost, and accurate indoor positioning technology is becoming increasingly prominent.

[0003] However, the existing machine learning-based perception methods have a large number of deep model parameters, and the inference computing power and storage requirements are relatively high, which is not conducive to promotion on ordinary equipment with limited resources, resulting in limited development of close-range perception and positioning technology. Summary of the Invention

[0004] In view of this, the purpose of this application is to propose a perception positioning method and electronic device to solve the problem of limited development of close-range perception positioning technology due to the large number of depth model parameters.

[0005] Based on the above objectives, a first aspect of the present application provides a perception positioning method, which is applied to a wireless access terminal with a beamforming function enabled in a target scene, where the target scene includes a target object to be positioned. The method includes:

[0006] Collect beamforming information fed back by fixed terminals connected to the wireless access point in the target scenario;

[0007] Inputting the beamforming information into a trained positioning model, and outputting the positioning position of the target object to be positioned through the positioning model;

[0008] Among them, the positioning model is obtained by training the beamforming information collected in different scenarios.

[0009] Optionally, beamforming information fed back by a fixed terminal connected to the wireless access point in the target scenario is collected, including:

[0010] A downlink data detection packet is sent to the fixed terminal, so that the fixed terminal performs channel estimation to obtain channel state information after receiving the data detection packet, decomposes and compresses the channel state information to obtain beamforming information, and feeds the beamforming information back to the wireless access terminal.

[0011] Optionally, inputting the beamforming information into a trained positioning model, and outputting the positioning position of the target object to be positioned through the positioning model, includes:

[0012] The beamforming information is input into the trained positioning model, and the beamforming information is fused through the feature fusion network in the positioning model to extract the low-dimensional position feature vector.

[0013] The low-dimensional position feature vector is input into the position probability regression network in the positioning model, and the position probability distribution of the target object in the predefined grid cells in the target scene is output through the position probability regression network;

[0014] The positioning position of the target object is determined according to the position probability distribution, and the positioning position is output through the positioning model.

[0015] Optional training methods for positioning models include:

[0016] Initialize the lightweight network to obtain the initial model;

[0017] For each of the multiple scenarios, collecting beamforming information fed back by the fixed terminal as a first data set;

[0018] constructing a plurality of combined datasets based on the first dataset of each scenario;

[0019] The initial model is trained using multiple combined data sets, and a positioning model is obtained after the training is completed.

[0020] Optionally, for each of the multiple scenarios, beamforming information fed back by the fixed terminal is collected as a first data set, including:

[0021] For each of the multiple scenarios, the area corresponding to each scene is divided into multiple grid units, a reference object is placed in each grid unit respectively, and the beamforming information fed back by the fixed terminal is collected. The beamforming information collected in each scenario and the grid unit position of the corresponding reference object are used as the first data set.

[0022] Optionally, constructing multiple combined data sets based on the first data set of each scene includes:

[0023] Different first data sets are split and combined to obtain multiple combined data sets.

[0024] Optionally, the initial model is trained using multiple combined datasets. After training, a positioning model is obtained, including:

[0025] Divide each combined dataset into a first training set and a first test set;

[0026] For each first training set, a replica model is constructed based on the initial model, and the replica model has the same model structure as the initial model;

[0027] The initial model is trained for multiple rounds of iterations. The training process for each round of iterations is as follows:

[0028] For each replica model, train the replica model using the corresponding first training set and first test set, and calculate the corresponding loss value;

[0029] Update the model parameters of the initial model according to the loss values ​​of all replica models;

[0030] In response to determining that the initial model does not meet a preset convergence criterion on at least one first test set, entering a next round of iterative training;

[0031] In response to determining that the initial model meets a preset convergence criterion on each first test set, the multiple rounds of iterative training are exited.

[0032] Optionally, for each replica model, the replica model is trained using the corresponding first training set and first test set to calculate a corresponding loss value, including:

[0033] For each replica model, the replica model is trained using the corresponding first training set to obtain the error loss, and the model parameters of the replica model are updated using the back propagation algorithm based on the error loss;

[0034] The updated replica model is tested using the corresponding first test set, and the corresponding loss value is calculated.

[0035] Optionally, before collecting beamforming information fed back by a fixed terminal connected to the wireless access terminal in the target scenario, the following steps are included:

[0036] The target scene area is divided into multiple grid cells, a reference object is placed in each grid cell, and beamforming information fed back by the fixed terminal is collected;

[0037] The collected beamforming information and the corresponding grid unit locations of the reference objects are used as sample data sets;

[0038] The localization model is fine-tuned using a sample dataset.

[0039] Based on the same inventive concept, the second aspect of the present application also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable by the processor, wherein the processor implements the method described above when executing the computer program.

[0040] As can be seen from the above, the present application provides a perception positioning method and electronic device, wherein the method is applied to a wireless access terminal with beamforming enabled in a target scene, wherein the target scene includes a target object to be located. Beamforming information fed back by a fixed terminal connected to the wireless access terminal in the target scene is collected. The beamforming information includes waveform information caused by the disturbance of the target object. The beamforming information is input into a trained positioning model, and the positioning model outputs the positioning position of the target object to be located. The positioning model is trained using beamforming information collected between the wireless access terminal and the fixed terminal in different scenarios. The trained positioning model can accurately predict the positioning position of the target object. Beamforming information is lightweight data with a low number of data parameters but fine information granularity, which can balance the real-time position prediction and perception positioning accuracy. Compared with complete channel state information, it can significantly compress data storage space and transmission bandwidth, while reducing the burden of data processing and positioning position prediction, improving the scalability of short-range perception positioning deployment, and is particularly suitable for deployment on resource-constrained devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in this application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are merely embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0042] Figure 1 A flowchart of a sensing positioning method according to an embodiment of the present application is shown;

[0043] Figure 2 A schematic diagram of the initial model training process of an embodiment of the present application;

[0044] Figure 3 This is a schematic structural diagram of a sensing and positioning device according to an embodiment of the present application;

[0045] Figure 4 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0046] In order to make the objectives, technical solutions and advantages of this application more clear, this application is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.

[0047] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should have the usual meanings understood by people with ordinary skills in the field to which this application belongs. The "first", "second" and similar words used in the embodiments of the present application do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0048] Classic long-range detection and positioning technology, which uses antenna arrays to detect the direction of arrival of signals and combines it with rigorous mathematical algorithms for signal estimation, has gradually matured. However, in close-range scenarios with complex environments and multipath reflections, conventional channel models struggle to accurately represent the diversity of wireless transmissions, and rigorous mathematical analysis struggles to achieve good performance. In recent years, with the widespread adoption of hotspot coverage communications, leveraging regular WiFi communication signals for close-range perception and positioning has become a viable solution. This technology leverages the Channel State Information (CSI) (CSI) of WiFi device communications to locate intruding objects (such as humans and drones). The key approach is to first collect CSI time series from target terminals in real time through WiFi access points (APs) in a specific environment. This data is then used to train and fit a deep neural network, ultimately achieving real-time positioning.

[0049] Wi-Fi technology, with its ubiquity, convenience, and low privacy concerns, combined with the perception of electromagnetic channel information generated during communication and the rapid response capabilities of deep learning, holds promise as a promising solution for close-range sensing and positioning. It will play a significant role in human perception and low-altitude object detection. However, existing machine learning-based perception methods require large deep model parameters, high inference computing power, and high storage capacity, hindering their application to resource-constrained, low-volume devices. Furthermore, these technologies typically collect specific CSI data in a laboratory or simulation scenario and train and evaluate model performance, failing to fully consider the diversity of environmental factors such as close-range environment layout, object materials, and object distribution. When the surrounding environment changes, wireless channel characteristics drift significantly, leading to increased perception and positioning errors in the model output. Consequently, there is a lack of effective solutions for adapting to new scenarios. Adapting to new scenarios often requires time-consuming retraining using large datasets, posing significant challenges to both real-time perception and inference and resilience to changing environments. Therefore, maintaining high accuracy while compressing the model and rapidly adapting to changing scenarios are key challenges facing close-range positioning technology.

[0050] There are two major problems with existing indoor and other short-range positioning technologies based on deep learning: (1) the computational and storage hardware overhead of predictive positioning is large, making it difficult to deploy in real time on ordinary devices with limited resources; (2) the data analysis method lacks the ability to quickly adapt to environmental changes, and the accuracy often drops significantly in new scenarios. In view of this, this application proposes a perceptual positioning method that replaces CSI with beamforming information (BFI) during short-range positioning. BFI is mainly used to adjust the direction of communication signal beam energy in WiFi systems. Applying it to perceptual positioning can reduce the number of feature parameters and provide finer information granularity, which can balance the real-time nature of inference and the accuracy of perceptual positioning. For example, under 4×4 MIMO (multiple-in multipleout) conditions, the number of parameters of a single BFI data is 256×12, and the number of parameters of CSI data is 256×16×2. The former is only 37.5% of the latter. In deep learning tasks, when the number of input data parameters is reduced, the corresponding model inference amount can also be reduced. Replacing CSI with BFI can significantly reduce the data processing and inference burden. Through the domain-enhanced meta-learning architecture and online fine-tuning mechanism, it can correct channel drift in real time using a small number of labeled samples, ensuring positioning accuracy and improving the scalability of perception deployment.

[0051] The embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0052] This application proposes a perception positioning method, referring to Figure 1, applied to a wireless access terminal with beamforming function enabled in a target scene, the target scene including a target object to be located, wherein the method comprises the following steps:

[0053] Step 102: Collect beamforming information fed back by a fixed terminal connected to the wireless access point in the target scene.

[0054] Specifically, the target scenario is a scenario with wireless network coverage. Examples include offices, rooftops, basements, and low-altitude locations. The wireless access point is a wireless router (WiFi device) with beamforming enabled, and the fixed terminal is connected to this wireless router. The fixed terminal can be one or more fixed wireless routers or a fixed mobile terminal (such as a mobile phone or computer) with signal transmission and reception capabilities. Based on the WiFi protocol process, the wireless access point can extract beamforming information.

[0055] Step 104: Input the beamforming information into a trained positioning model, and output the position of the target object to be positioned through the positioning model. The positioning model is trained by beamforming information collected in different scenarios.

[0056] Specifically, the positioning model is a trained machine learning model, specifically a deep neural network model. Exemplarily, the positioning model is a meta-learning agent. The deep neural network model is trained using beamforming information collected in different scenarios, and a positioning model is obtained after training. The positioning model can accurately predict the position of a target object in the target scene. The target object can be a human body or another object. When the target object is at different locations in the target scene, the corresponding collected beamforming information is different, and the target object perturbs the beamforming information. Therefore, real-time inference and prediction based on the beamforming information can determine the target object's location.

[0057] In addition, because the human body and objects have different effects on CSI, if the positioning position of the human body in the target scene is predicted, it is necessary to collect training data with the help of the human body's position transformation when training the positioning model, and the reference object used in the training or fine-tuning process is the human body. If the positioning position of the object in the target scene is predicted, it is necessary to collect training data with the help of the object's position transformation when training the positioning model, and the reference object used in the training or fine-tuning process is the object. It should be noted that due to the different materials of objects, there are also different effects on CSI. Therefore, the positioning model trained with training data collected from objects of the same type of material can only accurately locate objects of the same type of material.

[0058] By replacing CSI with BFI, positioning only requires beamforming information for prediction, which reduces the number of parameters and avoids the need for high-precision measurement and storage of raw amplitude and phase data. This significantly reduces the number of data parameters and storage overhead. Under the same model parameters, the inference workload is significantly reduced, alleviating the real-time inference computing overhead of edge devices.

[0059] Based on steps 102 to 104 above, the perception positioning method provided in this embodiment is applied to a wireless access point with beamforming enabled in a target scene, where the target scene contains a target object to be located. Beamforming information is collected from a fixed terminal connected to the wireless access point in the target scene. The beamforming information contains waveform information caused by the disturbance of the target object. This beamforming information is input into a trained positioning model, which then outputs the location of the target object to be located. The positioning model is trained by collecting beamforming information in different scenarios. The trained positioning model can accurately predict the location of the target object. Beamforming information is lightweight data with a low number of data parameters but fine information granularity, balancing real-time location prediction with perception positioning accuracy. This significantly reduces data storage space and transmission bandwidth, while reducing the burden of data processing and positioning prediction. This improves the scalability of short-range perception positioning deployment, making it particularly suitable for deployment on resource-constrained devices.

[0060] In some embodiments, collecting beamforming information of a fixed terminal connected to a wireless access terminal in a target scene includes:

[0061] A downlink data detection packet is sent to the fixed terminal, so that the fixed terminal performs channel estimation to obtain channel state information after receiving the data detection packet, decomposes and compresses the channel state information to obtain beamforming information, and feeds the beamforming information back to the wireless access terminal.

[0062] Specifically, after the wireless access point has enabled the beamforming function, it can periodically perform downlink detection on the fixed terminal connected to it. In terms of placement, the wireless access point and the fixed terminal are deployed separately, making the regional positioning between the two devices more accurate. Specifically, a data detection packet is sent to the fixed terminal. For example, the data detection packet is an NDP detection packet. The NDP detection packet refers to an information packet used in the Neighbor Discovery Protocol (NDP) to detect and discover other devices in the network. The wireless access device sends a pilot signal known to both the transmitter and receiver by sending an NDP detection packet. After receiving the NDP detection packet, the fixed terminal compares the received pilot signal with the original information and calculates the channel state information of the channel. Next, according to the IEEE Std 802.11 Part 11: Wireless LAN Medium Access Control (MAC) and Physical Layer (PHY) Specifications, Section 19.3.12.3.6, Compressed beamforming feedback matrix, Pages 2398-2400, the channel state information is subjected to singular value decomposition (SVA). The first spatial stream columns of the Zk matrix are extracted as the precoding matrix Vk. This is then compressed according to the protocol process to produce BFI data, which has relatively few parameters and is represented as an angle scalar. After obtaining the BFI data, the fixed terminal packages the beamforming information and feeds it back to the wireless access point. The wireless access point reads the corresponding fields in the data packet to obtain the compressed BFI data. Because BFI data is smaller than SCI data, subsequent machine learning using BFI data is faster. The downlink detection and BFI feedback process runs periodically in the wireless router with beamforming enabled, obtaining a BFI data source that can be used for model training and subsequent model inference.

[0063] In some embodiments, inputting the beamforming information into a trained positioning model, and outputting the positioning position of the target object to be positioned through the positioning model, includes:

[0064] The beamforming information is input into the trained positioning model, and the beamforming information is fused through the feature fusion network in the positioning model to extract the low-dimensional position feature vector.

[0065] The low-dimensional position feature vector is input into the position probability regression network in the positioning model, and the position probability distribution of the target object in the predefined grid cells in the target scene is output through the position probability regression network;

[0066] The positioning position of the target object is determined according to the position probability distribution, and the positioning position is output through the positioning model.

[0067] Specifically, the localization model includes a feature fusion network and a position probability regression network. The feature fusion network is used to extract low-dimensional position feature vectors from the beamforming information. The position probability regression network can calculate the position probability of the target object in each predefined grid cell. After receiving the beamforming information, the localization model inputs the beamforming information in parallel into the feature fusion network. The feature fusion network fuses the multi-channel time-space-frequency information and extracts a low-dimensional position feature vector. The low-dimensional position feature vector is then fed into a subsequent position probability regression network, which outputs a position probability distribution of the target object. The position probability distribution contains multiple position probabilities, each corresponding to a grid cell. The position probability represents the probability of the target object appearing in that grid cell. The target scene is pre-divided into multiple grid cells. Finally, based on the position probability distribution, strategies such as maximum probability or weighted average are used to determine the location of the target object.

[0068] The maximum probability strategy determines the grid cell with the highest probability in the location probability distribution as the location. The weighted average strategy takes a weighted average of the location probabilities to obtain a predicted expected value, which is used as the location. After determining the location, the positioning model outputs the location. The wireless access point can also send the location to upper-layer applications or user terminals through an interface, enabling visualization of the target object's location and subsequent service invocation.

[0069] In some embodiments, a method for training a positioning model includes:

[0070] Initialize the lightweight network to obtain the initial model;

[0071] For each of the multiple scenarios, collecting beamforming information fed back by the fixed terminal as a first data set;

[0072] constructing a plurality of combined datasets based on the first dataset of each scenario;

[0073] The initial model is trained using multiple combined data sets, and a positioning model is obtained after the training is completed.

[0074] Exemplary lightweight networks include convolutional networks and linear layer networks. During initialization, the weights and biases of each layer of the lightweight network are randomly initialized to construct an initial model. Multiple scenarios can include office scenes, outdoor rooftop scenes, or low-altitude scenes.

[0075] For each scenario, the beamforming information fed back by the fixed terminal is collected as the first data set, which specifically includes:

[0076] For each of the multiple scenarios, the area corresponding to each scene is divided into multiple grid units, a reference object is placed in each grid unit respectively, and the beamforming information fed back by the fixed terminal is collected. The beamforming information collected in each scenario and the grid unit position of the corresponding reference object are used as the first data set.

[0077] Exemplarily, the area where each scene is located is divided into m discrete location points, and each discrete location point corresponds to a grid unit. When collecting beamforming information, the reference object collects beamforming information once in each grid unit, and the amount of beamforming information finally collected is equal to the number of grid units. The coordinates of the discrete location points in the grid unit are the location tags corresponding to the beamforming information. In each scene, after the beamforming information corresponding to all grid units is collected, all the beamforming information and the corresponding location tags are used as the first data set. The reference object can be a human body or an object.

[0078] Constructing multiple combined data sets based on the first data set of each scene specifically includes: splitting and combining different first data sets to obtain multiple combined data sets.

[0079] For each scene, the first data set is divided into several subsets based on dimensions such as spatial position, time or object position. The subsets under different scenes are then combined across scenes to obtain multiple combined data sets. For example, the two scenes are scene one and scene two, the first data set of scene one is divided into three subsets {A, B, C}, and the first data set of scene two is divided into three subsets {a, b, c}. In order to obtain the common generalization features of different scenes, the subsets of the two scenes are combined to obtain several combined data sets such as {A, a}, {B, b}, {C, c}, thereby realizing domain enhanced data construction. Each combined data set corresponds to a combined scene. By training the initial model with the combined data set, the robustness of the model to changes in channel distribution can be enhanced. In the process of learning based on the combined data set, the model learns the generalized common features of different scenes, thereby enhancing the applicability of the model.

[0080] Furthermore, the initial model is trained using multiple combined datasets. After training, a positioning model is obtained, including:

[0081] Divide each combined dataset into a first training set and a first test set;

[0082] For each first training set, a replica model is constructed based on the initial model, and the replica model has the same model structure as the initial model;

[0083] The initial model is trained for multiple rounds of iterations. The training process for each round of iterations is as follows:

[0084] For each replica model, train the replica model using the corresponding first training set and first test set, and calculate the corresponding loss value;

[0085] Update the model parameters of the initial model according to the loss values ​​of all replica models;

[0086] In response to determining that the initial model does not meet a preset convergence criterion on at least one first test set, entering a next round of iterative training;

[0087] In response to determining that the initial model meets a preset convergence criterion on each first test set, the multiple rounds of iterative training are exited.

[0088] Specifically, each combined dataset is divided into a first training set and a first test set. For example, if the combined dataset contains 100 data items, 70 of them can be used as the first training set, and the remaining 30 as the first test set. Each data item in the combined dataset includes a location tag for a point and beamforming information collected by the wireless access point when a reference object is at that point, i.e., a location tag-beamforming information pair.

[0089] For each first training set, that is, for each combined scenario, a replica model is constructed, and the replica model has the same model structure as the initial model. The initial model is iteratively trained for multiple rounds, including: for each replica model, the replica model is trained using the first test set and the first training set to obtain a loss value. The model parameters of the initial model are updated based on the loss values ​​of all replica models. If the loss value obtained when the updated initial model is tested on at least one first test set is significantly lower than that of the previous test, it is determined that the preset convergence criteria are not met, and the initial model needs to be iteratively trained for the next round. If the loss value obtained when the updated initial model is tested on all first test sets no longer decreases or the rate of decrease is extremely low, it is determined that the preset convergence criteria are met. At this point, the initial model completes the training process, and after fixing the model parameters, a positioning model is obtained. By combining data, the differences between scene domains can be amplified, and a positioning model with efficient migration capabilities for unknown environments can be obtained.

[0090] Figure 2A schematic diagram of the initial model training process is shown. Multiple combined datasets include combined dataset 1, combined dataset 2, ..., combined dataset n. Each combined dataset is input into a corresponding replica model. The feature distributions of different combined datasets may differ. The initial model can serve as a meta-model or a meta-model agent. Each replica model outputs a corresponding loss value. Specifically, replica model 1 outputs a loss value of 1, replica model 2 outputs a loss value of 2, ..., and replica model n outputs a loss value of n. A comprehensive deviation loss is determined based on all loss values. For example, the comprehensive deviation loss can be an average loss value. The comprehensive deviation loss is input into the initial model, which is then updated and optimized through gradient learning. After the initial model parameters are updated, the initial model parameters are copied to each replica model to ensure consistency between the replica model parameters and the initial model parameters. The next round of iterative learning is then performed until the preset convergence conditions are met. The initial model at this point is also the positioning model. Parallel learning of multiple combined datasets in different replica models improves learning efficiency.

[0091] Furthermore, for each replica model, the replica model is trained using the corresponding first training set and the first test set, and the corresponding loss value is calculated, including:

[0092] For each replica model, the replica model is trained using the corresponding first training set to obtain an error loss, and the model parameters of the replica model are updated using a back propagation algorithm based on the error loss;

[0093] The updated replica model is tested using the corresponding first test set, and the corresponding loss value is calculated.

[0094] Specifically, for each replica model, the replica is first trained using the first training set. Data from the first training set is then fed into the replica model, which then outputs a predicted position. The error between the predicted position and the position label is used to calculate the error loss. Based on this error loss, the replica model's weights and biases are updated using a backpropagation algorithm. The updated replica model is then tested on the first test set, and the prediction loss is calculated. Using different data sets for the training and test sets allows for better evaluation of the model's learning performance.

[0095] In related technologies, when faced with the diversity and dynamic changes of indoor environments, models trained on a single scene often cannot guarantee the stability of cross-scene performance. This application adjusts the positioning model through the following method to ensure the stability of the model's cross-scene performance.

[0096] In some embodiments, before collecting beamforming information fed back by a fixed terminal connected to a wireless access terminal in a target scenario, the process includes:

[0097] The target scene area is divided into multiple grid cells, a reference object is placed in each grid cell, and beamforming information fed back by the fixed terminal is collected;

[0098] The collected beamforming information and the corresponding grid unit locations of the reference objects are used as sample data sets;

[0099] The localization model is fine-tuned using a sample dataset.

[0100] Specifically, the trained positioning model from the aforementioned embodiments is downloaded from a remote or central server and loaded onto a locally deployed device (such as a smart terminal or edge computing node), serving as the foundation for subsequent fine-tuning. Before applying the positioning model to the target scenario, fine-tuning is performed to ensure rapid adaptation. This involves two steps: sample data collection and local rapid adaptation and fine-tuning of the positioning model.

[0101] In the target scene to be deployed, several representative locations are manually or automatically selected to form grid cells, each containing a single location. Reference objects are placed at each location. For each location, beamforming information and precise location labels are collected as a sample dataset. The collection process should cover the target scene's primary areas of occlusion, reflection, and multipath distribution to ensure representative sample data. Only small batches of sample data are needed to fine-tune the positioning model, significantly reducing the time and labor costs of large-scale data collection.

[0102] The localization model is then fine-tuned using a small-scale gradient update using the sample dataset. After fine-tuning, the updated model parameters are fixed to form a localized localization model suitable for the current target scenario. The fine-tuned localization model is then deployed to edge devices or embedded platforms. If the target scenario changes significantly, the above fine-tuning process is automatically repeated, achieving continuous model adaptation. This fine-tuned localization model enables efficient and accurate online localization of target objects.

[0103] When deploying the positioning model in a new scenario, only a very small amount of sampling and quick fine-tuning are required to complete a smooth transition from a general meta-model to a dedicated model, significantly improving the system's scenario robustness and "plug-and-play" deployment efficiency.

[0104] It should be noted that the method of the embodiment of the present application can be performed by a single device, such as a computer or server. The method of this embodiment can also be applied in a distributed scenario and performed by multiple devices working together. In such a distributed scenario, one of the multiple devices may only perform one or more steps of the method of the embodiment of the present application, and the multiple devices will interact with each other to complete the method.

[0105] It should be noted that the above description is of some embodiments of the present application. Other embodiments are within the scope of the description. In some cases, the actions or steps described in the description can be performed in an order different from that in the above embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0106] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides a perception and positioning device.

[0107] refer to Figure 3 The sensing and positioning device is applied to a wireless access terminal with a beamforming function enabled in a target scene, wherein the target scene includes a target object to be located, and the sensing and positioning device includes:

[0108] The collection module 302 is configured to collect beamforming information fed back by a fixed terminal connected to the wireless access terminal in the target scene;

[0109] The positioning module 304 is configured to input the beamforming information into a trained positioning model and output a positioning position of the target object to be positioned through the positioning model;

[0110] Among them, the positioning model is obtained by training the beamforming information collected in different scenarios.

[0111] In some embodiments, the acquisition module 302 further includes:

[0112] The detection unit is configured to send a downlink data detection packet to the fixed terminal, so that the fixed terminal performs channel estimation after receiving the data detection packet to obtain channel state information, decomposes and compresses the channel state information to obtain beamforming information, and feeds back the beamforming information to the wireless access terminal.

[0113] In some embodiments, the positioning module 304 includes:

[0114] a feature extraction unit configured to input beamforming information into a trained positioning model, fuse the beamforming information through a feature fusion network in the positioning model, and extract a low-dimensional position feature vector;

[0115] a probability regression unit configured to input the low-dimensional position feature vector into a position probability regression network in the positioning model, and output a position probability distribution of the target object in a predefined grid cell in the target scene through the position probability regression network;

[0116] The position determination unit is configured to determine the positioning position of the target object according to the position probability distribution and output the positioning position through the positioning model.

[0117] In some embodiments, a training module is further included, configured to:

[0118] Initialize the lightweight network to obtain the initial model;

[0119] For each of the multiple scenarios, collecting beamforming information fed back by the fixed terminal as a first data set;

[0120] constructing a plurality of combined datasets based on the first dataset of each scenario;

[0121] The initial model is trained using multiple combined data sets, and a positioning model is obtained after the training is completed.

[0122] In some embodiments, the training module includes:

[0123] The first data set acquisition unit is configured to divide the area corresponding to each scene into multiple grid units for each scene, place a reference object in each grid unit, and collect beamforming information fed back by the fixed terminal, and use the beamforming information collected in each scene and the grid unit position of the corresponding reference object as the first data set.

[0124] In some embodiments, the training module includes:

[0125] The combined data set acquisition unit splits and combines different first data sets to obtain multiple combined data sets.

[0126] In some embodiments, a training module is further included, configured to:

[0127] an iterative training unit configured to divide each combined data set into a first training set and a first test set;

[0128] For each first training set, a replica model is constructed based on the initial model, and the replica model has the same model structure as the initial model;

[0129] The initial model is trained for multiple rounds of iterations. The training process for each round of iterations is as follows:

[0130] For each replica model, the replica model is trained using the corresponding first training set and the first test set, and the corresponding loss value is calculated; the model parameters of the initial model are updated according to the loss values ​​of all replica models; in response to determining that the initial model does not meet the preset convergence criteria on at least one first test set, the next round of iterative training is entered; in response to determining that the initial model meets the preset convergence criteria on each first test set, multiple rounds of iterative training are exited.

[0131] In some embodiments, the training module includes:

[0132] The loss value calculation unit is configured to train each replica model through the corresponding first training set to obtain the error loss, and update the model parameters of the replica model based on the error loss using the back propagation algorithm; test the updated replica model through the corresponding first test set to calculate the corresponding loss value.

[0133] In some embodiments, before collecting beamforming information fed back by a fixed terminal connected to the wireless access terminal in the target scene, a fine-tuning module is further included, which is configured to:

[0134] The area corresponding to the target scene is divided into multiple grid cells, a reference object is placed in each grid cell, and beamforming information fed back by a fixed terminal is collected. The collected beamforming information and the corresponding grid cell locations of the reference objects are used as a sample dataset. The sample dataset is used to fine-tune the positioning model.

[0135] For the convenience of description, the above device is described as being divided into various modules according to their functions. Of course, when implementing this application, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0136] The device of the above embodiment is used to implement the corresponding perception positioning method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.

[0137] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein when the processor executes the program, the perception positioning method described in any of the above embodiments is implemented.

[0138] Figure 4A more specific hardware structure diagram of an electronic device provided in this embodiment is shown. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other within the device via the bus 1050.

[0139] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0140] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 1020 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.

[0141] The input / output interface 1030 is used to connect to input / output modules to enable information input and output. The input / output modules can be configured as components within the device (not shown) or externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, and various sensors. Output devices may include a display, speaker, vibrator, indicator light, and the like.

[0142] The communication interface 1040 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, Wi-Fi, Bluetooth, etc.).

[0143] The bus 1050 comprises a pathway for transmitting information between various components of the device, such as the processor 1010 , the memory 1020 , the input / output interface 1030 , and the communication interface 1040 .

[0144] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in a specific implementation, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may only include the components necessary to implement the embodiments of this specification, and does not necessarily include all the components shown in the figure.

[0145] The electronic device of the above embodiment is used to implement the corresponding perception and positioning method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.

[0146] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the perception positioning method described in any of the above embodiments.

[0147] The computer-readable media of this embodiment includes permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.

[0148] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the perception positioning method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0149] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present application is limited to these examples. In line with the present application, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.

[0150] In addition, to simplify the description and discussion, and to avoid obscuring the understanding of the embodiments of the present application, well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided figures. Furthermore, devices may be shown in block diagram form to avoid obscuring the understanding of the embodiments of the present application, and this also takes into account the fact that the implementation details of these block diagram devices are highly dependent on the platform on which the embodiments of the present application will be implemented (i.e., these details should be fully understood by those skilled in the art). Where specific details (e.g., circuits) are set forth to describe the exemplary embodiments of the present application, it will be apparent to those skilled in the art that the embodiments of the present application can be implemented without these specific details or with variations therefrom. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0151] Although the present invention has been described in conjunction with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those skilled in the art based on the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may utilize the discussed embodiments.

[0152] The embodiments of the present application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the specification. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application should be included in the scope of protection of this application.

Claims

1. A perception positioning method, characterized in that: The method is applied to a wireless access terminal with a beamforming function enabled in a target scene, wherein the target scene includes a target object to be located. Collect beamforming information fed back by fixed terminals connected to the wireless access point in the target scenario; The beamforming information is input into the trained positioning model, which then outputs the position of the target object to be positioned, including: The beamforming information is input into the trained positioning model, and the beamforming information is fused through the feature fusion network in the positioning model to extract the low-dimensional position feature vector. The low-dimensional position feature vector is input into the position probability regression network in the positioning model, and the position probability distribution of the target object in the predefined grid cells in the target scene is output through the position probability regression network; Determine the location of the target object based on the location probability distribution, and output the location through the positioning model; Among them, the positioning model is obtained by training the beamforming information collected in different scenarios.

2. The method according to claim 1, characterized in that The collecting of beamforming information fed back by a fixed terminal connected to a wireless access terminal in a target scenario includes: A downlink data detection packet is sent to the fixed terminal, so that the fixed terminal performs channel estimation to obtain channel state information after receiving the data detection packet, decomposes and compresses the channel state information to obtain beamforming information, and feeds the beamforming information back to the wireless access terminal.

3. The method according to claim 1, characterized in that The training method of the positioning model includes: Initialize the lightweight network to obtain the initial model; For each of the multiple scenarios, collecting beamforming information fed back by the fixed terminal as a first data set; constructing a plurality of combined datasets based on the first dataset of each scenario; The initial model is trained using multiple combined data sets, and a positioning model is obtained after the training is completed.

4. The method according to claim 3, characterized in that The collecting, for each of the multiple scenarios, beamforming information fed back by the fixed terminal as a first data set includes: For each of the multiple scenarios, the area corresponding to each scene is divided into multiple grid units, a reference object is placed in each grid unit respectively, and the beamforming information fed back by the fixed terminal is collected. The beamforming information collected in each scenario and the grid unit position of the corresponding reference object are used as the first data set.

5. The method according to claim 3, characterized in that Multiple combined datasets are constructed based on the first dataset of each scenario, including: Different first data sets are split and combined to obtain multiple combined data sets.

6. The method according to claim 3, characterized in that The initial model is trained using multiple combined datasets. After training, a positioning model is obtained, including: Divide each combined dataset into a first training set and a first test set; For each first training set, a replica model is constructed based on the initial model, and the replica model has the same model structure as the initial model; The initial model is trained for multiple rounds of iterations. The training process for each round of iterations is as follows: For each replica model, train the replica model using the corresponding first training set and first test set, and calculate the corresponding loss value; Update the model parameters of the initial model according to the loss values ​​of all replica models; In response to determining that the initial model does not meet a preset convergence criterion on at least one first test set, entering a next round of iterative training; In response to determining that the initial model meets a preset convergence criterion on each first test set, the multiple rounds of iterative training are exited.

7. The method according to claim 6, characterized in that For each replica model, the replica model is trained using the corresponding first training set and first test set, and the corresponding loss value is calculated, including: For each replica model, the replica model is trained using the corresponding first training set to obtain the error loss, and the model parameters of the replica model are updated using the back propagation algorithm based on the error loss; The updated replica model is tested using the corresponding first test set, and the corresponding loss value is calculated.

8. The method according to claim 1, characterized in that Before collecting beamforming information fed back by a fixed terminal connected to a wireless access point in a target scenario, the following steps are included: The target scene area is divided into multiple grid cells, a reference object is placed in each grid cell, and beamforming information fed back by the fixed terminal is collected; The collected beamforming information and the corresponding grid unit locations of the reference objects are used as sample data sets; The localization model is fine-tuned using a sample dataset.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Beam forming device, intelligent antenna and wireless communication equipment

    CN105634578A

  • Beam forming method and device of intelligent metasurface for communication perception integration

    CN115765819A