An Ultra-Wideband Indoor Positioning Method Based on a Multimodal Diffusion Model

Through multimodal diffusion model and feature fusion technology, data loss and signal interference problems of UWB indoor positioning in NLOS environment are solved, and high-precision and robust indoor positioning are achieved, suitable for complex environments.

CN119907098BActive Publication Date: 2025-07-04NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510397494.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-04
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

The existing UWB indoor positioning method faces the problem of degradation of positioning accuracy caused by large-scale data loss and signal interference in the NLOS environment, and fails to make full use of multimodal information and complex environmental modes.

Method used

The ultra-wideband indoor positioning method based on the multimodal diffusion model is adopted. Data is collected through a mobile multi-sensor platform, and after standardization, the multimodal feature extraction module and the improved denoising diffusion implicit model (DDIM) are used for reverse sampling, and feature fusion and position estimation are combined with the U-Net architecture neural network and long-term memory network to establish an end-to-end positioning system.

Benefits of technology

It significantly improves the positioning accuracy and robustness in the NLOS environment, can effectively reconstruct lost data, maintain statistical characteristics and time consistency, adapt to indoor environments of different complexities, and improves computing efficiency and positioning performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119907098B_ABST
    Figure CN119907098B_ABST
Patent Text Reader

Abstract

The present invention discloses an ultra-wideband indoor positioning method based on a multimodal diffusion model, belonging to the field of indoor positioning. The steps include: collecting distance data and RSSI data between each UWB anchor-tag in the indoor environment based on a mobile multi-sensor platform; performing feature extraction and feature fusion on the multimodal input data through a multimodal feature extraction module; using an improved denoising diffusion implicit model for accelerated reverse sampling, and implementing high-quality distance reconstruction by means of a neural network with a U-Net architecture to output a distance estimation value; establishing a position estimation layer, using a long short-term memory network for time series modeling to learn position estimation and output a three-dimensional coordinate positioning value. The method of the present invention can effectively fuse UWB distance data and RSSI data, achieve high-quality reconstruction of missing data, while retaining the statistical characteristics and spatio-temporal correlation of the data, and significantly improve the accuracy and robustness of indoor positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of indoor positioning, and particularly to an ultra-wideband indoor positioning method based on a multimodal diffusion model. Background Art

[0002] Ultra-WideBand (UWB) technology has broad application prospects in the field of indoor positioning due to its high resolution, high precision, and strong anti-interference ability. The indoor measurement data of UWB sensors mainly includes four types: the distance measurement between tags and anchors, the Time of Arrive (TOA) of tag signals arriving at anchors, the Received Signal Strength Indicator (RSSI), and the channel impulse responses (CIR) data between tags and anchors. Among them, distance measurement is the basis of UWB positioning, and the distance between an anchor and a tag is usually calculated using a TOA-based method. RSSI reflects the signal strength of the receiver and is usually used to estimate the distance between devices.

[0003] RSSI measurement consists of two different components: the Remaining Received Signal Strength Indicator (RxRssi) and the First Path Received Signal Strength Indicator (FpRssi). These two RSSI measurements provide complementary information about the signal propagation environment. RxRssi reflects the overall signal strength, while FpRssi indicates the strength of the first-arriving signal path. The relationship between these two measurements is particularly valuable for determining Line-Of-Sight (LOS) or Non-Line-Of-Sight (NLOS) conditions: in an ideal LOS environment, signal attenuation follows a relatively stable propagation law, and the RSSI value maintains a clear relationship with the distance between the tag and the anchor. The path loss model or the log-distance model is usually used for accurate distance estimation.

[0004] However, in the actual indoor environment, UWB signal propagation faces multiple challenges. 1. In the NLOS environment, when the signal encounters obstacles such as walls and furniture, reflection, refraction, and scattering will occur, resulting in the multipath effect, which significantly reduces the ranging accuracy. 2. Environmental factors may further exacerbate the signal propagation problem. When the tag and the base station are blocked or affected by electromagnetic interference, continuous data loss may occur, manifested as the inability to obtain the distance and RSSI measurement values of the tag by multiple base stations simultaneously. 3. Large-scale data loss not only affects the accuracy of individual measurements but also leads to the complete loss of target tracking ability for a long time.

[0005] Existing methods have made important progress in solving precise positioning in complex environments. For example, traditional TOA-based methods mainly rely on Kalman filtering (KF) and various optimization algorithms; deep learning-based methods such as those using Convolutional Neural Networks (CNN), Long Short-Term Memory Networks (LSTM), and Graph Neural Networks (GNN) alleviate the errors of UWB positioning in NLOS environments by learning complex environmental patterns in UWB data, such as spatio-temporal features and location features.

[0006] However, existing methods have the following deficiencies in dealing with positioning problems in complex environments: 1. Traditional TOA-based methods mainly rely on Kalman filtering (KF) and various optimization algorithms. When dealing with data loss, simple preprocessing techniques such as forward filling or linear interpolation are usually used to maintain basic data continuity, but the basic statistical characteristics of UWB measurements cannot be retained. 2. Deep learning-based methods improve performance by learning complex environmental patterns in UWB data, but still rely on basic interpolation techniques when dealing with missing data. Simple data completion strategies cannot retain the statistical distribution of UWB measurements, resulting in distribution shift and potential data contamination, significantly reducing the positioning accuracy during long-term signal occlusion. 3. Most existing methods fail to fully utilize multi-modal information. Although some methods combine distance and RSSI data, these modalities are usually treated as independent inputs or simple concatenations are performed, failing to capture their complex interdependencies and ignoring the supplementary information contained in RSSI signals.

[0007] In recent years, diffusion models have recently demonstrated remarkable capabilities in generative modeling and data reconstruction tasks. These models learn to reverse the gradually noising process, enabling them to generate high-quality samples from complex data distributions and effectively reconstruct missing data while retaining statistical characteristics. The advantage of diffusion models lies in their ability to model complex data distributions through iterative denoising steps, making them particularly suitable for addressing the challenge of data loss in UWB positioning systems.

[0008] The limitations of the existing technologies indicate that there is a need for an innovative UWB indoor positioning method that can effectively handle large-scale data loss, fully integrate multi-modal information, and adapt to complex NLOS environments. Summary of the Invention

[0009] The present invention provides a multi-modal diffusion model-based ultra-wideband indoor positioning method, aiming to solve the problem of decreased positioning accuracy caused by large-scale data loss and signal interference in NLOS environments.

[0010] The present invention adopts the following technical solution: A multi-modal diffusion model-based ultra-wideband indoor positioning method, including the following steps:

[0011] S1. Collect data based on a mobile multi-sensor platform: The mobile multi-sensor platform is provided with a UWB base station, a lidar, and a computing unit, collects the distance data and RSSI signal strength data between each UWB anchor-tag in the indoor environment, obtains the ground truth data through the lidar, and obtains the true value of the movement trajectory of the mobile multi-sensor platform;

[0012] S2. Standardize the collected data, including: distance data processing, RSSI data processing, and data missing processing, to obtain multi-modal input data;

[0013] S3. Through a multi-modal feature extraction module, perform feature extraction and feature fusion on the multi-modal input data;

[0014] S4. Construct a multi-modal diffusion model, based on the forward Markov chain process of the diffusion model, adopt an improved denoising diffusion implicit model (DDIM) for reverse sampling, establish a neural network with a U-Net architecture, input the data after feature extraction and feature fusion, and realize distance reconstruction through multi-step denoising iteration. After training, output the distance estimation value;

[0015] S5. Establish a position estimation layer, accept the input of the distance estimation value, adopt a long short-term memory network for time series modeling, learn position estimation, and output the three-dimensional coordinate positioning value;

[0016] S6. Conduct network training based on supervised learning, introduce a hybrid loss function for iterative optimization, obtain a trained multi-modal diffusion model, which is used to achieve high-precision indoor positioning in a complex non-line-of-sight environment, and effectively solve the problem of decreased positioning accuracy under signal interference and large-scale data loss conditions.

[0017] Preferably, in step S1, the data collection based on the mobile multi-sensor platform includes the following sub-steps:

[0018] S1.1. Several UWB anchors are set in the indoor environment, and each UWB anchor is equipped with a UWB sensor;

[0019] S1.2. Fix the positions of the UWB base station on the mobile multi-sensor platform and each UWB anchor in the indoor environment, and determine the coordinates of the UWB base station and each UWB anchor in a unified coordinate system through a calibration program;

[0020] S1.3. Control the mobile multi-sensor platform to move in the indoor environment according to a predetermined trajectory or randomly, and establish a data collection node through the ROS operating system;

[0021] S1.4. Establish communication by moving the UWB base station on the multi-sensor platform and each UWB anchor point in the indoor environment, obtain the distance values and RSSI signal strength values between the UWB anchor points and tags at each timestamp, and form different data sequences; obtain the reference position as the real-time ground truth through the lidar on the multi-sensor platform.

[0022] Preferably, in step S2, the normalization process includes the following sub-steps:

[0023] S2.1. Distance data processing: Organize the distance data between multiple UWB anchor points and tags in parallel, and aggregate the distance data of consecutive multiple timestamps to form a time window data block, and organize the aggregated data into a four-dimensional tensor structure ;

[0024] S2.2. RSSI data processing:

[0025] The RSSI signal strength data includes two RSSI signal strength values, namely: received signal strength indicator data (RxRssi), first-path received signal strength indicator data (FpRssi), which respectively represent the overall signal strength and the strength of the first-arrival signal path;

[0026] Organize the two RSSI signal strength values through a tensor structure, and add a channel in the first dimension to distinguish the two RSSI signal strength values to obtain a tensor structure ;

[0027] S2.3. Data missing processing: Fill the missing values in the input data using the forward filling method.

[0028] Preferably, in step S3, the multi-modal feature extraction module adopts a parallel feature processing architecture, processes the RSSI signal strength data and distance data respectively, and extracts features from the multi-modal input data. The method is as follows:

[0029] S3.1.1. Preprocess and transform the RSSI signal strength data through the RSSI feature conversion module, and map the two-channel features to a unified feature space;

[0030] S3.1.2. Perform multi-level feature extraction on the distance data and RSSI signal strength data along the distance path and RSSI path respectively. The distance path and RSSI path use the same network structure but do not share parameters; during the feature extraction process, learn the interdependence between different modalities through the convolutional layer;

[0031] The multi-level feature extraction includes a low-level feature layer, a middle-level feature layer, and a high-level feature layer, forming a progressive feature extraction structure. Feature transformation is performed through 3D convolution at each level. The low-level feature layer captures local spatial patterns, the middle-level feature layer synthesizes local features to form an abstract representation, and the high-level feature layer integrates spatial relationships to extract global semantic information.

[0032] Preferably, in step S3, feature fusion is performed on the multi-modal input data as follows:

[0033] S3.2.1. Set fusion points at multiple levels of feature extraction, and splice the features of the RSSI path and the distance path. and perform feature splicing;

[0034] S3.2.2. Perform feature transformation on the spliced features through a convolutional layer to learn the interdependence between different modalities;

[0035] S3.2.3. Use a 3×3×3 convolutional layer with the number of channels from 32 to 16 for feature fusion and dimensionality reduction to form feature vectors at different scales.

[0036] Preferably, in step S4, a multi-modal diffusion model is constructed, and an improved denoising diffusion implicit model is used to construct a Markov chain as follows:

[0037] S4.1. Based on the forward process of the diffusion model, construct a Markov chain, define the noise addition process, and gradually add noise to the original distance data;

[0038] S4.2. Adopt a linear noise scheduling strategy and use the reparameterization method to obtain the noise data at any time step ;

[0039] S4.3. Perform the reverse denoising process based on DDIM, and use the improved denoising diffusion implicit model to convert the conditional probability into a deterministic form to accelerate the sampling process of the diffusion model;

[0040] S4.4. Construct a time step embedding module to map the discrete time step to a high-dimensional space;

[0041] S4.5. Transform the time features through a two-layer fully connected network, map the 32-dimensional time embedding to a 64-dimensional feature vector, and inject it into each layer of the U-Net architecture neural network.

[0042] Preferably, in step S4, a U-Net architecture neural network is constructed as the backbone network of the diffusion model, including an encoder path and a decoder path;

[0043] The encoder path gradually extracts features through multiple downsampling blocks. Each downsampling module contains two convolutional layers and a downsampling operation, and at the same time, a projection adjustment of the time step embedding is added to predict noise;

[0044] The decoder path gradually restores the feature resolution through corresponding multiple upsampling operations, and at the same time, fuses the features of the corresponding layers of the encoder through skip connections. Each upsampling module contains a transposed convolutional layer and two ordinary convolutional layers.

[0045] Skip connections are introduced between the corresponding encoder and decoder layers to solve the problem of vanishing gradients:

[0046] The U-Net architecture neural network outputs the predicted noise component which has the same size as the input data (1×T×A×N) and is used for deterministic distance reconstruction in the DDIM sampling process;

[0047] Through the DDIM sampling strategy, the multi-modal diffusion model can gradually recover high-quality distance measurement values from the noise data.

[0048] Preferably, in step S5, the position estimation layer adopts a CNN-LSTM hybrid architecture to capture spatial relationships and temporal dependencies, and converts the reconstructed distance data into three-dimensional position coordinates for output, including the following sub-steps:

[0049] S5.1. Spatial feature extraction: Construct and design a convolutional neural network module to process multi-dimensional distance data through several 3D convolutional blocks, extract the spatial pattern in the anchor-label distance relationship, and retain the measured temporal structure;

[0050] S5.2. Time series modeling: Use a long short-term memory network to perform time series modeling on the processed spatial features, utilize the information of multiple consecutive time stamps, capture the temporal dependence in the target motion trajectory, and enhance the temporal consistency of the positioning result;

[0051] S5.3. Position coordinate regression: Construct a regression module containing fully connected layers to map the temporal features to three-dimensional position coordinates; output the estimated position coordinates to perform end-to-end learning from distance measurement to position.

[0052] Preferably, in step S6, the hybrid loss function takes into account both the distance reconstruction error and the position estimation error, and balances the optimization objectives through a weighting coefficient.

[0053] Preferably, in step S6, the network training based on supervised learning uses an end-to-end training method to optimize the parameters, including the following sub-steps:

[0054] S6.1. Load the standardized distance data, RSSI signal strength data, and corresponding true value data into the training set and test set;

[0055] S6.2. Sample the true distance data and RSSI signal strength data, add noise at specified time steps according to the forward process of the diffusion model to generate training samples, and perform forward calculation through the multi-modal diffusion model to obtain noise prediction values and corresponding distance and position estimates;

[0056] S6.3. Calculate the total loss, distance loss, and position loss according to the hybrid loss function;

[0057] S6.4. Update the parameters of the multi-modal diffusion model through the backpropagation algorithm according to the gradient of the loss function;

[0058] S6.5. Perform reverse sampling through the denoising diffusion implicit model to generate reconstructed distance measurements from the noise data and predict the target position.

[0059] Compared with the prior art, the present invention adopts the above technical solutions and has the following technical effects:

[0060] 1. Effectively cope with non-line-of-sight environments and improve positioning accuracy: The indoor positioning method of the present invention significantly improves the accuracy of the UWB indoor positioning system through the application of multi-modal data fusion and diffusion models. In complex non-line-of-sight (NLOS) environments, by fusing distance data and RSSI data, it effectively identifies and alleviates the measurement errors caused by the multi-path effect.

[0061] 2. Enhance the robustness in case of data loss: The diffusion model of the present invention can effectively reconstruct the lost measurement data and maintain its statistical characteristics and time consistency. Even in extreme cases of large-scale data loss, it can still maintain an acceptable positioning accuracy, and the increase in positioning error is much smaller than that of traditional methods.

[0062] 3. Improve computational efficiency: The present invention adopts the DDIM sampling strategy, which significantly reduces the number of sampling steps required by the diffusion model, improves the inference efficiency while maintaining the reconstruction quality, and can run in real time on resource-constrained embedded devices.

[0063] 4. Enhance the trajectory time consistency: The position estimation module based on the CNN-LSTM architecture of the present invention can effectively capture the time-dependent relationships in the motion trajectory, generate smoother and more coherent positioning trajectories, and significantly reduce the trajectory jump phenomenon, especially in scenarios where the signal quality temporarily deteriorates.

[0064] 5. Strong adaptability: The ultra-wideband indoor positioning method of the present invention performs excellently in environments with different complexities, and as the environmental complexity increases, its advantages over existing methods become more significant; from a simple living room environment to a complex laboratory environment, this method can maintain stable and reliable positioning performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 is the architecture block diagram of the ultra-wideband indoor positioning method network of the present invention;

[0066] Figure 2 is the layout diagram of the experimental scenario of the embodiment of the present invention;

[0067] Figure 3 is the flowchart of the data processing stage of the ultra-wideband indoor positioning method of the present invention;

[0068] Figure 4 is the schematic diagram of multi-modal feature extraction and multi-scale feature fusion of the ultra-wideband indoor positioning method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0069] In order to make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the application will be further elaborated in detail below with reference to the accompanying drawings. The described embodiments are only a part of the embodiments involved in the present invention. All non-innovative embodiments made by other researchers in the field based on this embodiment fall within the protection scope of the present invention. At the same time, for the step numbers in the embodiments of the present invention, they are only set for the convenience of elaboration and explanation, and no limitation is placed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0070] The present invention proposes an ultra-wideband indoor positioning method based on a multi-modal diffusion model to solve the problems of data loss and signal interference faced by the UWB positioning system in a non-line-of-sight (NLOS) environment. A data processing module, a multi-modal feature extraction module, a U-Net architecture neural network based on a diffusion model, and a position estimation module together constitute an end-to-end system. Through data input, feature fusion, loss calculation, and network training based on supervised learning, high-precision indoor positioning can be achieved using UWB distance data and RSSI signal strength data.

[0071] In an embodiment of the present invention, a complete experimental system is constructed to collect distance measurement data, RSSI signal strength data, and distance ground truth data in multiple environments and perform standardized processing.

[0072] Such as Figure 1As shown, the standardized distance data, RSSI signal strength data, and distance ground truth data are input into the multi-modal denoising module. Through the ultra-wideband indoor positioning method of the present invention, the position estimate and distance estimate are calculated, indoor data estimation is performed, and based on the network training of supervised learning, a hybrid loss function is introduced to iteratively optimize the position ground truth and distance ground truth, achieving high-precision indoor positioning in complex non-line-of-sight environments to verify the effectiveness of the method of the present invention.

[0073] The experimental system of this embodiment includes the following hardware devices: UWB sensors (configured with 6 anchors and 1 UWB base station, the anchors are fixedly installed around the environment, and the UWB base station is installed on the mobile multi-sensor platform), and a mobile multi-sensor platform (a robotic platform with mobility, carrying a UWB base station, a lidar, and an embedded computing unit).

[0074] Among them, the lidar is used to obtain the high-precision reference position as the ground truth, and the embedded computing unit is used to process data and execute the positioning algorithm.

[0075] Before the experiment process, the UWB coordinate system and the lidar coordinate system, as well as the positions of the anchors and the UWB base station in the coordinate system, have been previously transformed and aligned to ensure that all the measurement values provided by the system operate in a reference coordinate system.

[0076] To comprehensively evaluate the performance of this method in different scenarios, this embodiment designs three experimental environments with different complexities, as Figure 2 shown, including parameters such as different anchor position coordinates and environment sizes, specifically as follows:

[0077] Environment 1: Living room environment, as shown in (a) of Figure 2 , an indoor space of about 4m×8m with fewer obstacles, mainly in line-of-sight (LOS) conditions, serving as the benchmark test scenario;

[0078] Environment 2: Apartment environment, as shown in (b) of Figure 2 , a multi-room residential environment of about 8m×12m, containing obstacles such as walls and furniture, resulting in obvious non-line-of-sight conditions;

[0079] Environment 3: Laboratory environment, as shown in (c) of Figure 2 , a large laboratory of about 30m×30m with the most complex spatial layout, having a severe non-line-of-sight environment and multi-path effects, used to test the performance limit of the positioning algorithm under extreme conditions.

[0080] In this embodiment, in step S1, data is collected based on a mobile multi-sensor platform. The distance data and RSSI data between each UWB anchor-tag in the indoor environment are collected. A high-precision reference position is obtained through lidar as the ground truth, and the true value of the motion trajectory of the mobile multi-sensor platform is obtained.

[0081] The specific method is as follows:

[0082] In each experimental environment, the positions of the fixed anchors and UWB base stations are fixed, and their accurate coordinates in a unified coordinate system are determined through a pre-set calibration program;

[0083] The mobile robot platform is controlled to move along a predefined or random trajectory in the test site. A data collection node is established through the ROS operating system, and at the same time, the information of all sensors is recorded, including: UWB distance data, RSSI signal strength data, lidar data, and timestamps for synchronizing each data source.

[0084] Among them, the RSSI signal strength data includes: received signal strength indicator RxRssi and first-path received signal strength indicator FpRssi; the lidar data is used as the true positioning information value.

[0085] The experiment is repeated multiple times in different environments and scenarios to collect sufficient data samples to ensure the diversity and representativeness of the data. Finally, different data sequences are obtained, and the duration, content, and trajectory of each data sequence are different.

[0086] In this embodiment, one of each sequence is selected as the training set, and the rest are used as the validation set, as shown in Table 1 below.

[0087] Table 1 Conditions of each sequence in the experiment

[0088]

[0089] In this embodiment, in step S2, the collected data is standardized.

[0090] Data standardization is the first key step of the method of the present invention, including: distance data processing, RSSI data processing, data missing processing, to obtain multi-modal input data, as Figure 3 shown, specifically as follows:

[0091] First, considering the high sampling frequency (100Hz) of the UWB signal and the input requirements of the deep learning model, the distance data processing aggregates the distance measurement values of 10 consecutive timestamps into a data block. This processing method helps to smooth the instantaneous measurement fluctuations, reduce the calculation amount, and retain the spatial information while reducing the time granularity.

[0092] Organize the aggregated data into a four-dimensional tensor structure:

[0093] ;

[0094] Among them, 1 represents the initial channel dimension, is the time dimension (set to 16 in the experiment), is the number of anchor points (set to 6), is the combined sample size (set to 10). This data organization method effectively retains the time and space structure information of the data.

[0095] Then, for the RSSI data, this embodiment adopts a processing method similar to that of the distance data.

[0096] The RSSI data contains two different components: RxRssi and FpRssi, which respectively reflect the overall signal strength and the strength of the first-arrival signal path.

[0097] The relationship between the two is particularly important for judging line-of-sight or non-line-of-sight conditions:

[0098] When |RxRssi - FpRssi| < 6 dB, it is most likely a line-of-sight environment;

[0099] When |RxRssi - FpRssi| > 10 dB, it is most likely a non-line-of-sight environment;

[0100] When the difference is between 6 - 10 dB, it is an uncertain propagation condition.

[0101] Organize the RSSI data into a tensor structure :

[0102] ;

[0103] Among them, the superscript 2 in the first dimension represents two RSSI values.

[0104] Finally, perform data missing processing. In the actual indoor environment, due to factors such as obstacle occlusion, signal reflection, and multipath effects, UWB distance data often has missing values during acquisition. This embodiment adopts a simple forward filling method to maintain data continuity.

[0105] In this embodiment, step S3 extracts and fuses features from the multi-modal input data through a multi-modal feature extraction module.

[0106] Multi-modal feature extraction is the second core component of the present invention. As Figure 4 shown, a parallel feature processing architecture is designed to process UWB distance data and RSSI signal strength data respectively.

[0107] First, the RSSI feature conversion module preprocesses and converts the RSSI data, mapping the dual-channel features to a unified feature space. The expression is:

[0108] ;

[0109] Among them, is a neural network module containing a convolutional layer, batch normalization, and ReLU activation function. The specific implementation is two consecutive 3D convolutional layers. The first layer maps the 2-channel input to 8 channels, and the second layer maps the 8 channels to 16 channels, both using a 3x3x3 convolutional kernel and a 1×1×1 stride. is the output value after passing through Head RSSI layers.

[0110] Next, the UWB distance data and RSSI signal strength data are subjected to feature extraction along their respective processing paths. The two paths use the same network structure but do not share parameters, allowing them to independently learn the feature representations of specific modalities.

[0111] The convolutional operation in the feature extraction process is designed as:

[0112] ;

[0113] Among them, represents a three-dimensional convolutional operation, represents the convolutional kernel size, represents the stride, , represent the input and output of the convolutional operation respectively.

[0114] The multiple levels include a low-level feature layer, a middle-level feature layer, and a high-level feature layer, forming a progressive feature extraction structure. The low-level feature layer mainly captures local spatial patterns, the middle-level feature layer synthesizes local features to form a more abstract representation, and the high-level feature layer further integrates spatial relationships to extract global semantic information. Each level realizes feature transformation through 3D convolutional operations while maintaining independence in the time dimension.

[0115] It should be noted that although 3D convolutional layers are used in this embodiment, the convolutional kernel size in the depth dimension is set to 1 to ensure that the features of each time step are independently processed and avoid implicit mixing of features between time steps.

[0116] Multi-modal feature fusion is the third core component of the present invention. Fusion points are set at multiple levels of feature extraction. First, the features of the RSSI path and the distance path are concatenated:

[0117] ;

[0118] Among them, represents a splicing function, represents the spliced feature.

[0119] Then, the spliced feature is transformed through a convolutional layer to learn the interdependence between different modalities:

[0120] ;

[0121] Among them, represents a batch normalization operation, is a non-linear activation function, represents the fused feature.

[0122] Specifically, when implementing, a 3×3×3 convolutional layer with the number of channels ranging from 16 + 16 = 32 to 16 is used for feature fusion and dimensionality reduction. This progressive feature fusion method enhances the feature representation ability at multiple abstract levels while ensuring the independence of feature extraction, enabling the model to more effectively utilize multi-modal information.

[0123] In this embodiment, step S4 constructs a multi-modal diffusion model. Based on the forward Markov chain process of the diffusion model, an improved denoising diffusion implicit model (DDIM) is used for reverse sampling. A U-Net architecture neural network is established, and the data after feature extraction and feature fusion is input. Distance reconstruction is achieved through multi-step denoising iteration, and the distance estimation value is output after training.

[0124] The diffusion model is the core technology of the present invention, which is used to reconstruct high-quality distance measurement data that is lost or severely noisy. The forward process of the diffusion model is a process of gradually adding noise to the original data, forming a Markov chain.

[0125] For any time step , this process is defined as:

[0126] ;

[0127] Among them, represents the conditional probability distribution of adding noise from time step to time step , , respectively represent the noise data at time steps , , I represents the identity matrix, represents the Gaussian normal distribution, is a predefined noise intensity scheduling parameter, which is used to control the amount of added noise.

[0128] Using the reparameterization trick, the noise data at any time step t can be directly obtained:

[0129] ;

[0130] wherein, is the cumulative product, , is the original noise-free data, is the random noise of the standard normal distribution.

[0131] In actual training, this embodiment adopts a linear noise scheduling strategy with 30 time steps, gradually growing linearly from a smaller value to a larger value, specifically growing from 0.0001 to 0.02.

[0132] Contrary to the forward process, the reverse process is a process of gradually recovering the original signal from the noisy data. The traditional diffusion model (DDPM) usually requires a large number of sampling steps, reducing the inference efficiency. To solve this problem, this embodiment adopts an improved denoising diffusion implicit model (DDIM) to construct a Markov chain, converting the conditional probability into a deterministic form:

[0133] ;

[0134] wherein, represents the noise prediction network, implemented by the U-Net architecture; represents controlling the sampling variance.

[0135] When setting , the sampling process becomes completely deterministic, significantly reducing the required number of sampling steps while maintaining the reconstruction quality.

[0136] To help the network effectively adapt to different noise levels and adjust the denoising strategy, this embodiment further designs a time step embedding module to map the discrete time steps to a high-dimensional space:

[0137] ;

[0138] wherein, is the input dimension, is the predefined frequency factor, represents the th component of the time feature.

[0139] Specifically, in this embodiment, is set to 32, is set to 10000.

[0140] Subsequently, these time features are further transformed through a two-layer fully connected network, mapping the 32-dimensional time embedding to a 64-dimensional feature vector and injecting it into each layer of the U-Net.

[0141] The U-Net network structure in this embodiment includes an encoder path and a decoder path. The encoder path gradually extracts features through multiple downsampling modules, and each downsampling module contains two convolutional layers and a downsampling operation.

[0142] Taking the first downsampling module as an example, the number of input channels is 16×2 = 32 (the concatenation of distance features and multi-modal features), the number of output channels is 32, the convolutional kernel size is 3×3, the stride is 1×1, and at the same time, a projection of the time step embedding is added to adjust the noise prediction.

[0143] The calculation process of the downsampling module is as follows:

[0144] ;

[0145] Among them, represents the multi-modal features from the auxiliary encoder, is the time step embedding information, represents the feature representation after the -th layer of downsampling, represents the network module containing feature extraction and downsampling operations.

[0146] The decoder path gradually restores the feature resolution through corresponding upsampling operations, and at the same time fuses the features of the corresponding layers of the encoder through skip connections.

[0147] The upsampling module contains a transposed convolutional layer (deconvolution) and two ordinary convolutional layers, and the formula is

[0148] ;

[0149] Among them, represents the feature representation after the -th layer of upsampling, represents the network module containing feature extraction and upsampling operations.

[0150] For example, the first upsampling module reduces the number of feature channels from 32 to 16, uses a 2×2 deconvolutional kernel, with a stride of 2×2, and then connects two 3×3 convolutional layers for feature refinement.

[0151] To solve the problem of gradient disappearance and promote multi-scale feature fusion, this embodiment further introduces skip connections between the corresponding encoder and decoder layers:

[0152] ;

[0153] Among them, represents the feature after the skip connection.

[0154] The network finally outputs the predicted noise component , having the same dimensions as the input (1×16×6×10), is used for deterministic reconstruction during the DDIM sampling process. Through the DDIM sampling strategy, the model can gradually recover high-quality distance measurements from noisy data. This U-Net-based diffusion model architecture can effectively learn the noise distribution, thereby recovering high-quality distance measurement data from severely damaged inputs.

[0155] In this embodiment, step S5 establishes a position estimation layer, which accepts the input of distance estimation values, uses a long short-term memory network for time series modeling, learns position estimation, and outputs three-dimensional coordinate positioning values.

[0156] The position estimation module is the last core component of the present invention, responsible for converting the reconstructed distance data into the final three-dimensional position coordinates. This module adopts a CNN-LSTM hybrid architecture to effectively capture spatial relationships and time dependencies.

[0157] Spatial feature extraction is achieved by a series of three-dimensional convolution operations, which process multi-dimensional distance data and extract spatial patterns in the anchor-label distance relationship.

[0158] In this embodiment, the position estimation module contains 4 consecutive 3D convolution blocks, and the specific structure of each block is as follows:

[0159] The first convolution block: input channel 1, output channel 16, convolution kernel (1,1,5), stride 1, batch normalization, Sigmoid activation function;

[0160] The second convolution block: input channel 16, output channel 32, convolution kernel (1,1,3), stride 1, batch normalization, Sigmoid activation function;

[0161] The third convolution block: input channel 32, output channel 64, convolution kernel (1,1,3), stride 1, batch normalization, Sigmoid activation function;

[0162] The fourth convolution block: input channel 64, output channel 128, convolution kernel (1,1,2), stride 1, batch normalization, Sigmoid activation function.

[0163] These convolution blocks gradually extract features, reduce the spatial dimension and increase the number of channels. In particular, the last fully connected layer compresses the feature dimension of 128×6 to 128 dimensions.

[0164] It should be noted that the size of the convolution kernel in the time dimension is set to 1. This design is to maintain the independence between time steps and avoid mixing features of different time steps.

[0165] To make full use of the information of multiple consecutive timestamps and capture the temporal dependencies in the target motion trajectory, this embodiment uses a Long Short-Term Memory (LSTM) network for time series modeling.

[0166] The LSTM network has the following parameter configurations: an input dimension of 128, a hidden layer dimension of 100, 3 layers of LSTM stacked, and batch_first=True is used to ensure that the first dimension of the input tensor is the batch size. The LSTM network can effectively capture the long-term dependencies in the time series, which is particularly important for smoothing trajectory prediction and improving positioning stability.

[0167] The position coordinate regression module consists of fully connected layers that map the temporal features output by the LSTM to the final three-dimensional position coordinates. The first fully connected layer maps the 100-dimensional LSTM output features to a 60-dimensional intermediate representation, using the ReLU activation function to increase non-linearity; the second fully connected layer maps the 60-dimensional features to the final 3-dimensional output, corresponding to the three-dimensional space coordinates (x, y, z).

[0168] The hybrid architecture design of CNN-LSTM in this embodiment effectively combines the advantages of convolutional neural networks in spatial feature extraction and the capabilities of recurrent neural networks in time series processing, and can accurately infer the position coordinates of the target from the reconstructed distance data.

[0169] To simultaneously optimize the distance reconstruction performance and position estimation accuracy, this embodiment also designs a joint loss function that takes both objectives into account.

[0170] This loss function consists of two parts: a distance loss (using the mean squared error to measure the difference between the predicted distance and the true distance) and a position loss (also using the mean squared error to calculate the difference between the predicted position and the true position). The specific formula is as follows:

[0171] ;

[0172] Among them, represents the hybrid loss function, is the distance loss, specifically the mean squared error between the true distance value and the predicted distance value; is the position loss, specifically the mean squared error between the true positioning value and the predicted positioning value; and are the weight coefficients that balance the two loss terms, and are both set to 1.0 in the experiments of this embodiment, giving the same importance to distance reconstruction and position estimation.

[0173] The model training in this embodiment adopts an end-to-end approach to simultaneously optimize all the parameters of the system, and combines an early stopping strategy to prevent overfitting.

[0174] First, load the preprocessed UWB distance data, RSSI data, and corresponding true position data into the training set and test set, using batch sizes of 256 and 32 respectively; then sample the true distance data, add noise to generate training samples, perform forward calculation through the model to obtain the predicted distance and position; next, calculate the total loss, distance loss, and position loss according to the joint loss function; finally, update the model parameters according to the gradient of the loss function through the backpropagation algorithm.

[0175] During the training process, regularly evaluate the model performance on the test set and monitor key metrics such as the mean squared error of distance and the root mean square error (RMSE) of position, etc. In this embodiment, the Adam optimizer is used for training, the initial learning rate is set to 5e-5, without weight decay (weight_decay = 0), and combined with the early stopping strategy to prevent overfitting.

[0176] In the inference stage, this embodiment efficiently generates reconstructed distance measurements from the noise data through the DDIM sampling strategy and then predicts the target position. The specific process is as follows: First, extract the features of the UWB distance data and RSSI data through a multimodal encoder, then use the DDIM sampling loop to start from random noise and generate reconstructed distance measurement data through a 30-step iterative denoising process, and finally convert the reconstructed distance data into three-dimensional position coordinates through the position estimation module. Compared with the traditional DDPM sampling, DDIM sampling can significantly reduce the sampling steps (from thousands of steps to 30 steps), greatly improve the inference efficiency, and at the same time maintain the sampling quality, making this method suitable for real-time application scenarios.

[0177] To comprehensively evaluate the performance of the method of the present invention, this embodiment further designs a series of comparative experiments and tests its accuracy and robustness in different environments.

[0178] First, the root mean square error (RMSE) is used to measure the average distance error between the predicted position and the true position. Select Transformer-LSTM (a deep learning method based on Transformer and LSTM), STA-GNN-S (a positioning method based on spatio-temporal attention graph neural network), KF-LSTM (a hybrid method combining Kalman filter and LSTM), and CNN-LSTM (a traditional deep learning positioning method using CNN and LSTM) as comparative methods. The experimental results show that the method of the present invention is superior to the comparative methods in all test scenarios.

[0179] More importantly, comparative experiments under different data missing rates (70%, 80%, 90%) show that when the data missing rate increases, the performance of all methods decreases, but the performance of the method of the present invention decreases the least. Even in the case of 90% data missing, the method of the present invention can still maintain a high positioning accuracy, while the performance of the comparative methods drops sharply. This proves the excellent ability of the method of the present invention in dealing with large-scale data loss. Trajectory visualization analysis also shows that in the most challenging laboratory environment, the trajectory reconstructed by the method of the present invention is closest to the real trajectory, and can maintain a high tracking accuracy even in non-line-of-sight areas, while other methods often show large deviations or tracking loss when encountering non-line-of-sight areas.

[0180] In summary, the ultra-wideband indoor positioning method based on multi-modal diffusion model proposed by the present invention effectively solves the problems of large-scale data loss and signal interference in non-line-of-sight environments by combining multi-modal feature extraction, diffusion model and position estimation techniques, and significantly improves the performance and robustness of the UWB indoor positioning system. This method performs excellently in various complex environments, especially when the data quality is poor, providing strong support for high-precision indoor positioning in practical applications.

[0181] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for ultra-wideband indoor positioning based on a multi-modal diffusion model, characterized in that, It includes the following steps: S1. Collect data based on a mobile multi-sensor platform: The mobile multi-sensor platform is equipped with a UWB base station, a lidar, and a computing unit. It collects the distance data and RSSI signal strength data between each UWB anchor-tag in the indoor environment, obtains the ground truth data through the lidar, and obtains the true value of the movement trajectory of the mobile multi-sensor platform; S2. Perform standardized processing on the collected data, including: distance data processing, RSSI data processing, and data missing processing, to obtain input data; S3. Through a multi-modal feature extraction module, perform feature extraction and feature fusion on the input data; S4. Construct a multi-modal diffusion model. Based on the forward Markov chain process of the diffusion model, use an improved denoising diffusion implicit model for reverse sampling. Establish a neural network with a U-Net architecture, input the data after feature extraction and feature fusion, and achieve distance reconstruction through multi-step denoising iteration. After training, output the distance estimation value; S5. Establish a position estimation layer, receive the input of the distance estimation value, use a long short-term memory network for time series modeling, learn position estimation, and output the three-dimensional coordinate positioning value; S6. Load the standardized distance data, RSSI signal strength data, and the corresponding ground truth data into the training set and test set, perform network training based on supervised learning, introduce a hybrid loss function for iterative optimization, and obtain a trained multi-modal diffusion model for achieving high-precision indoor positioning in a complex non-line-of-sight environment.

2. The ultra-wideband indoor positioning method based on a multi-modal diffusion model according to claim 1, wherein In step S1, the data collection based on the mobile multi-sensor platform includes the following sub-steps: S1.

1. Several UWB anchors are set in the indoor environment, and each UWB anchor is equipped with a UWB sensor; S1.

2. Fix the positions of the UWB base station on the mobile multi-sensor platform and each UWB anchor in the indoor environment, and determine the coordinates of the UWB base station and each UWB anchor in a unified coordinate system through a calibration program; S1.

3. Control the mobile multi-sensor platform to move in the indoor environment according to a predetermined trajectory or randomly, and establish a data collection node through the ROS operating system; S1.

4. Establish communication between the UWB base station on the mobile multi-sensor platform and each UWB anchor in the indoor environment, obtain the distance values and RSSI signal strength values between the UWB anchor-tag at each timestamp, and form different data sequences; obtain the reference position as the real-time ground truth through the lidar on the mobile multi-sensor platform.

3. The ultra-wideband indoor positioning method based on a multi-modal diffusion model according to claim 1, wherein, In step S2, the standardized processing includes the following sub-steps: S2.

1. Distance data processing: Parallelize the distance data between multiple UWB anchors and tags, aggregate the distance data of consecutive multiple timestamps to form a time window data block, and organize the aggregated data into a four-dimensional tensor structure , which is expressed as follows: ; Among them, 1 represents the initial channel dimension, T is the time dimension, A is the number of anchors, and N is the combined sample size; S2.

2. RSSI data processing: The RSSI signal strength data contains two types of RSSI signal strength values, namely: received signal strength indicator data, first-path received signal strength indicator data, which respectively represent the overall signal strength and the strength of the first-arrival signal path; Organize the two RSSI signal strength values through a tensor structure, and add a channel in the first dimension to distinguish the two RSSI signal strength values, obtaining a tensor structure , which is expressed as follows: ; Among them, the superscript 2 in the first dimension represents the two RSSI signal strength values; S2.

3. Data missing processing: Use the forward filling method to fill the missing values in the input data.

4. The ultra-wideband indoor positioning method based on a multi-modal diffusion model according to claim 3, characterized in that, In step S3, the multimodal feature extraction module adopts a parallel feature processing architecture to process the distance data and RSSI signal strength data respectively, and extract features from the input data. The method is as follows: S3.1.

1. Preprocess and transform the RSSI signal strength data through the RSSI feature conversion module, and map the dual-channel features to a unified feature space. The expression is: ; Among them, HeadRSSI is a neural network module that includes a convolutional layer, batch normalization, and a ReLU activation function, and includes two consecutive 3D convolutional layers. The first 3D convolutional layer maps the 2-channel input to 8 channels, and the second 3D convolutional layer maps the 8 channels to 16 channels. is the output value after passing through HeadRSSI layers; S3.1.

2. Perform multi-level feature extraction on the distance data and RSSI signal strength data along the distance path and RSSI path respectively. The distance path and RSSI path use the same network structure but do not share parameters; During the feature extraction process, the interdependence between different modalities is learned through the convolutional layer. The convolutional operation is expressed as: ; Among them, represents a three-dimensional convolution operation, represents the convolution kernel size, represents the stride, 、 represent the input and output of the convolution operation respectively; The multi-level feature extraction includes a low-level feature layer, a middle-level feature layer, and a high-level feature layer, forming a progressive feature extraction structure. Each level performs feature transformation through 3D convolution; the low-level feature layer captures local spatial patterns, the middle-level feature layer synthesizes local features to form an abstract representation, and the high-level feature layer integrates spatial relationships to extract global semantic information.

5. The ultra-wideband indoor positioning method based on a multi-modal diffusion model according to claim 4, wherein, In step S3, feature fusion is performed on the multimodal input data. The method is as follows: S3.2.

1. Set fusion points at multiple levels of feature extraction, and splice the features of the RSSI path and the distance path: and perform feature splicing: ; Among them, represents the splicing function, represents the feature after splicing; S3.2.

2. Perform feature transformation on the concatenated features through the convolutional layer to learn the interdependence between different modalities: ; Among them, represents a batch normalization operation, is a non-linear activation function, represents the fused feature; S3.2.

3. Use a 3×3×3 convolutional layer with the number of channels from 32 to 16 for feature fusion and dimensionality reduction to form feature vectors at different scales.

6. The ultra-wideband indoor positioning method based on a multi-modal diffusion model according to claim 5, wherein, In step S4, an improved denoising diffusion implicit model is used to construct a Markov chain. The method is as follows: S4.

1. Based on the forward process of the diffusion model, construct a Markov chain, define the noise addition process, and gradually add noise to the original distance data. For any time step , the forward process is defined as: ; Among them, represents the conditional probability distribution of adding noise from time step to time step and represent the noise data at time steps and respectively. I represents the identity matrix, represents the Gaussian normal distribution, is a predefined noise intensity scheduling parameter used to control the amount of added noise;​ S4.

2. Adopt a linear noise scheduling strategy and use the reparameterization method to obtain the noise data at any time step : ; Among them, is the cumulative product, , is the original noise-free data, is the random noise of the standard normal distribution; S4.

3. Perform a reverse denoising process based on DDIM. Use the improved denoising diffusion implicit model to convert the conditional probability into a deterministic form to accelerate the sampling process of the diffusion model: ; Among them, represents the noise prediction network, which is implemented by the U-Net architecture; represents the control sampling variance. When is set, the sampling process becomes completely deterministic; S4.

4. Construct a time step embedding module to map the discrete time steps to a high-dimensional space: ; Among them, is the input dimension, is a predefined frequency factor, represents the th component of the time feature; S4.

5. Transform the time features through a two-layer fully connected network, map the 32-dimensional time embedding to a 64-dimensional feature vector, and inject it into each layer of the U-Net architecture neural network.

7. The ultra-wideband indoor positioning method based on a multi-modal diffusion model according to claim 6, wherein In step S4, a U-Net architecture neural network is constructed as the backbone network of the diffusion model, including an encoder path and a decoder path; In the encoder path, features are gradually extracted through multiple downsampling blocks. Each downsampling module contains two convolutional layers and a downsampling operation. At the same time, a projection adjustment of the time step embedding is added to predict the noise. The calculation process is: ; Among them, represents the multimodal features from the auxiliary encoder, is the time step embedding information, represents the feature representation after downsampling in the i-th layer, and DownBlock represents a network module that includes feature extraction and downsampling operations; In the decoder path, the feature resolution is gradually restored through corresponding upsampling operations. At the same time, the features of the corresponding layer of the encoder are fused through skip connections. Each upsampling module contains a transposed convolutional layer and two ordinary convolutional layers. The calculation process is: ; Among them, represents the feature representation after upsampling of the i-th layer, and UpBlock represents a network module that includes feature extraction and upsampling operations; A skip connection is introduced between the corresponding encoder and decoder layers to solve the problem of vanishing gradients: ; Among them, represents the feature after skip connection; The noise component predicted by the neural network of the U-Net architecture , the noise component has the same size as the input data and is used for deterministic distance reconstruction in the DDIM sampling process.

8. The ultra-wideband indoor positioning method based on a multi-modal diffusion model according to claim 7, characterized in that, In step S5, the position estimation layer adopts a CNN-LSTM hybrid architecture to capture spatial relationships and temporal dependencies, and convert the reconstructed distance data into a three-dimensional position coordinate output. It includes the following sub-steps: S5.

1. Spatial Feature Extraction: Construct a designed convolutional neural network module, process multi-dimensional distance data through several 3D convolutional blocks, extract the spatial patterns in the anchor-label distance relationship, and retain the measured temporal structure; S5.

2. Time Series Modeling: Use a long short-term memory network to perform time series modeling on the processed spatial features, utilize multiple consecutive timestamp information, capture the temporal dependencies in the target motion trajectory, and enhance the temporal consistency of the positioning results; S5.

3. Position Coordinate Regression: Construct a regression module containing fully connected layers to map the temporal features to three-dimensional position coordinates; output the estimated position coordinates to perform end-to-end learning from distance measurement to position.

9. The ultra-wideband indoor positioning method based on a multi-modal diffusion model according to claim 8, wherein, In step S5.1, the convolutional neural network module includes 4 consecutive 3D convolutional blocks, gradually extracting and refining spatial features, reducing the spatial dimension and increasing the number of channels; the structure of each 3D convolutional block is as follows: The first convolutional block: input channel 1, output channel 16, convolutional kernel (1, 1, 5), stride 1, batch normalization, Sigmoid activation function; The second convolutional block: input channel 16, output channel 32, convolutional kernel (1, 1, 3), stride 1, batch normalization, Sigmoid activation function; The third convolutional block: input channel 32, output channel 64, convolutional kernel (1, 1, 3), stride 1, batch normalization, Sigmoid activation function; The fourth convolutional block: input channel 64, output channel 128, convolutional kernel (1, 1, 2), stride 1, batch normalization, Sigmoid activation function.

10. The ultra-wideband indoor positioning method based on a multi-modal diffusion model according to claim 8, characterized in that, In step S6, the network training based on supervised learning is used to optimize the parameters in an end-to-end training manner, including the following sub-steps: S6.

1. Load the standardized distance data, RSSI signal strength data, and corresponding ground truth data into the training set and test set; S6.

2. Sample the real distance data and RSSI signal strength data, add noise at specified time steps according to the forward process of the diffusion model to generate training samples, perform forward calculation through the multi-modal diffusion model, and obtain the noise prediction values and corresponding distance and position estimates; S6.

3. Calculate the total loss, distance loss, and position loss according to the hybrid loss function; The hybrid loss function takes into account both the distance reconstruction error and the position estimation error, and balances the optimization objectives through a weighting coefficient. The formula is as follows: ; Among them, is the distance loss, using the mean square error of the true distance value and the predicted distance value, is the position loss, using the mean square error of the true positioning value and the predicted positioning value, and is the weight coefficient for balancing the distance loss and the position loss, represents the hybrid loss function; S6.

4. Update the parameters of the multi-modal diffusion model according to the gradient of the loss function through the backpropagation algorithm; S6.

5. Perform reverse sampling through the denoising diffusion implicit model to generate the reconstructed distance measurement values from the noise data and predict the target position.

Citation Information

Patent Citations

  • Ultra-wideband indoor positioning method based on heterogeneous graph and attention mechanism

    CN118984489A

  • User behavior analysis method and device for burglary prevention of mobile equipment in exhibition hall

    CN119580406A