Passive positioning method for passengers in railway passenger transport station based on multi-modal image fusion
Through multimodal image fusion technology, combined with RGB-D cameras and inertial sensors, an adaptive word bag tree and a three-dimensional point cloud model are built, and local sensitive hashing and Kalman filtering are used to achieve efficient and accurate passive positioning of railway passenger stations, solving the problems of high cost, low accuracy and poor real-time performance in the existing technology.
Patent Information
- Application Number
- CN202510459859.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The passive visual positioning method of existing railway passenger stations has problems such as high hardware costs, serious signal interference, insufficient positioning accuracy and poor real-time performance in large and high-complex scenarios, which is difficult to meet the sub-meter positioning needs of passengers.
Multimodal image fusion technology is adopted to collect data through RGB-D cameras and inertial sensors, data augmentation and feature extraction are performed, adaptive word bag tree is built, and three-dimensional point cloud model is optimized using GPU, and local sensitive hash function and Kalman filtering are used for positioning to achieve passive positioning.
It reduces positioning costs, improves positioning accuracy and robustness, meets sub-meter-level accuracy requirements, adapts to the real-time navigation needs of complex scenarios, and optimizes operational efficiency.
Smart Images

Figure CN120339570A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of passenger indoor positioning, and particularly relates to a passive positioning method for passengers in railway passenger stations based on multi-modal image fusion. Background Art
[0002] With the rapid development of the high-speed rail network, as a transportation hub, the service capacity and operation management level of railway passenger stations directly affect the travel efficiency and experience of passengers. In large high-speed railway stations, due to the complex spatial layout and large passenger flow, passengers often face problems such as difficult navigation, getting lost, or being stranded. Therefore, accurate indoor positioning technology has become a key link to improve the service quality and operation efficiency within the station. Currently, the following challenges exist in passenger navigation and passenger flow management within the station:
[0003] Spatial complexity: The multi-layer structure and intensive functional areas (ticket check gates, waiting halls, commercial areas) make it difficult for traditional signage systems to meet the real-time navigation requirements;
[0004] Dynamic interference: During peak hours, the passenger flow density > 5 people / ㎡, and metal structures (steel frames, elevators) cause multipath attenuation of wireless signals;
[0005] Accuracy requirements: Passengers need sub-meter (<1 meter) positioning accuracy to accurately find ticket check gates or shops, while the error of existing active positioning technologies (such as Wi-Fi fingerprinting) often exceeds 3 meters.
[0006] In this context, passive vision positioning technology has become a research hotspot. The core of passive vision positioning is to achieve positioning through environmental visual information (rather than external signals), but the following technical bottlenecks need to be overcome:
[0007] ① Efficiency of large-scale scene reconstruction: The SfM (Structure from Motion) algorithm takes dozens of hours to process 100,000 images;
[0008] ② Robustness of image retrieval: A single feature (such as SIFT (Scale-invariant feature transform)) has a high false matching rate under different lighting conditions and complex scene layouts;
[0009] ③ Real-time performance and resource consumption: The algorithm needs to be adapted to edge devices (such as mobile phones) to achieve lightweight deployment and application in the passenger service scenarios within railway stations.
[0010] Currently, the passive vision positioning methods for railway station passengers in the prior art include:
[0011] ① Bluetooth beacon positioning: The principle is as follows: Bluetooth beacons (such as iBeacon) are deployed inside the station, and the distance is calculated by receiving the signal strength, and the position is estimated by combining the triangulation method. This method has the disadvantages of high hardware cost (one beacon needs to be deployed every 20 meters, and more than 5000 beacons are required for a 100,000-square-meter station, and the deployment cost exceeds 2 million yuan), being easily interfered by environmental signals, and relying on user devices (the Bluetooth of passengers needs to be turned on and a specific APP needs to be installed).
[0012] ② Wi-Fi fingerprint positioning: Its principle is to pre-collect the Wi-Fi signal strength (fingerprint database) at each position inside the station, and determine the position by matching the real-time signal with the fingerprint database. It also has the disadvantages of high maintenance cost and poor adaptability to the dynamic environment with large changes in the passenger flow density in the station space. (Patent number: CN202411691394.X)
[0013] The disadvantages of the above-mentioned existing passive vision positioning methods for railway station passengers include: These methods mainly rely on active positioning schemes, such as deploying Bluetooth beacons, wireless Wi-Fi access points or RFID devices, and using the mobile devices carried by passengers to interact with these infrastructures to achieve positioning. However, such methods have significant disadvantages: First, the device deployment and maintenance costs are high, especially in large stations where a large number of sensors are required to cover a vast area; second, the indoor environment of the station is complex, and metal structures and dense crowds will cause signal interference and attenuation, and the positioning accuracy often cannot meet the sub-meter-level requirements; in addition, active positioning relies on the support of passenger devices, and some passengers may not carry compatible devices or turn off relevant functions, resulting in limited positioning coverage.
[0014] For high-speed railway passenger stations with complex spatial layouts and large scales, these methods still have limitations: They often only target single objects or limited indoor areas for positioning, and it is difficult to directly apply them to large-scale scenarios such as the entire railway station; at the same time, the existence of large-scale image datasets will significantly reduce the efficiency of image matching and retrieval, and the existing vision positioning schemes still have deficiencies in terms of real-time performance and robustness to dynamic interference under massive data.
[0015] Therefore, there is an urgent need for an efficient, accurate and passive positioning method without external devices to meet the needs of large-scale and complex scenarios of high-speed railway passenger stations, provide convenient navigation services for passengers, and improve the station management efficiency at the same time. Summary of the Invention
[0016] An embodiment of the present invention provides a passive positioning method for railway passenger station passengers based on multi-modal image fusion to effectively position railway passenger station passengers.
[0017] To achieve the above object, the present invention adopts the following technical solutions.
[0018] A passive positioning method for passengers in railway passenger stations based on multi-modal image fusion, comprising:
[0019] Collect and construct a multi-modal image dataset through RGB-D cameras and inertial sensor devices in the functional areas of high-speed railway stations, and perform data augmentation on the images in the multi-modal image dataset through random occlusion and affine transformation;
[0020] Extract local features and semantic features from the data-augmented images, and construct an adaptive bag-of-words tree;
[0021] Use a graphics processing unit (GPU) to divide the bag-of-words tree into subsets, select key frames using a key-frame dynamic screening mechanism, and optimize the projection error through sparse bundle adjustment to generate a global three-dimensional point cloud model;
[0022] Generate hash codes using multiple groups of locality-sensitive hashing functions, use the global three-dimensional point cloud model as a physical space mapping benchmark to screen high-similarity candidate images through a voting mechanism, and obtain a preliminary pose estimate from the perspective of the passenger;
[0023] According to the camera internal parameter matrix and rotation matrix from the perspective of the passenger, perform smoothing processing on the preliminary pose estimate from the perspective of the passenger through Kalman filtering, and output the positioning coordinates of the passenger.
[0024] Preferably, the step of collecting and constructing a multi-modal image dataset through RGB-D cameras and inertial sensor devices in the functional areas of high-speed railway stations, and performing data augmentation on the images in the multi-modal image dataset through random occlusion and affine transformation, includes:
[0025] Deploy RGB cameras and inertial measurement units (IMUs) in the functional areas of railway passenger stations, ensure the time stamp alignment of RGB images, depth maps, and IMU data through hardware trigger signals, with the spatial coordinate system having the station center as the origin, the z-axis perpendicular upward, and the x-axis and y-axis parallel to the station ground, and each RGB camera and IMU synchronously capture data to construct a multi-modal image dataset;
[0026] Perform data augmentation on the image data in the multi-modal image dataset. Randomly generate rectangular occlusion areas in the images, simulate dynamic interference with the occluder texture using Gaussian noise, and within the random occlusion areas, the noise value of each pixel follows a Gaussian distribution where the mean μ = 0, and perturb the pixel value I(x, y) of the image within the occlusion area:
[0027]
[0028] Apply a random affine transformation matrix A to each image to obtain the data-augmented image. The random affine transformation matrix A is defined as:
[0029]
[0030] where the scaling factor s ∈ [0.8, 1, 2], the rotation angle θ ∈ [-30°, 30°], and the translation amount t x , t y ∈ [-50, 50] pixels.
[0031] Preferably, extracting local features and semantic features from the data-augmented images and constructing an adaptive bag-of-words tree includes:
[0032] Detecting the FAST corners of the spherical radio telescope for each data-augmented image, calculating the BRIEF descriptor, and generating a 256-dimensional binary feature vector using a rotation-invariant descriptor. The direction angle θ compensation formula of the rotation-invariant descriptor is:
[0033]
[0034] I(x, y) is the image pixel intensity value, and x, y represent the two-dimensional coordinate indices of the pixel in the image or neighborhood;
[0035] Extracting a 128-dimensional semantic feature vector from the original RGB image using a semantic model The 256-dimensional binary feature vectors of N images and the 128-dimensional semantic feature vectors are concatenated into a 384-dimensional hybrid feature vector
[0036] Using the HDBSCAN algorithm for the hybrid feature vector Performing hierarchical clustering, constructing a core distance based on k-nearest neighbors, generating a weighted adjacency graph, then generating a minimum spanning tree through the Prim algorithm, pruning based on the mutual reachability distance, extracting the clustering hierarchy in the tree, calculating the cluster stability λ, calculating the ratio of the number of samples in the cluster to the cluster diameter, removing clusters with λ < 0.7, merging to form clusters with high stability, and obtaining a hierarchical bag-of-words tree. Each leaf node in the bag-of-words tree stores the cluster center feature vector and the spatial coordinates of the corresponding image.
[0037] Preferably, using a graphics processing unit GPU to divide the bag-of-words tree into subsets, selecting key frames using a key-frame dynamic screening mechanism, and optimizing the projection error through a sparse light speed adjustment method to generate a global three-dimensional point cloud model includes:
[0038] The hierarchical cluster structure of the bag-of-words tree, the hybrid feature vector and its corresponding image spatial coordinates, as well as the camera internal parameter matrix K and the initial estimate T of the shooting pose i = [R i | t iInput into the GPU, and the GPU divides the bag-of-words tree into C subsets {a1, a2, …, a c} according to spatial proximity and feature similarity. Each subset contains local features and semantic features;
[0039] The GPU extracts matching feature points {(p i , a j )} for each subset pair (a i , a j ), calculates the essential matrix E ij through the five-point method, and decomposes the essential matrix E ij to obtain the relative rotation matrix Taking the minimization of the rotation matrix difference between all subset pairs as the objective function, the optimization objective function is:
[0040]
[0041] where R i , R j are the rotation matrices of cameras i and j respectively, is the estimated value of the rotation matrix between the two cameras, ||·|| F represents the Frobenius norm, and the optimized global rotation matrix R i is obtained;
[0042] Combine the optimized global rotation matrix R′ i with the initial translation t i to obtain the key-frame pose T i ′ = [R′ i |t i . Based on the key-frame pose T i ′ and the matching feature point pairs {(p i , p j )}, calculate the sparse 3D points X l , the camera internal parameter K, and the feature point matching relationship O = {(k, l, x kl )}, where x kl represents the pixel coordinate observation value of the 3D point X l in the key frame T k ′. Construct the objective function by minimizing the sum of the reprojection errors, and optimize all key-frame poses and 3D points simultaneously:
[0043]
[0044] where π([x, y, z] T ) = [x / z, y / z] T is the perspective projection function, π(KT k ′·X l ) - xkl Define the projection error e between each observation (k, l). kl , where ρ is the Huber robust kernel function, defined as:
[0045]
[0046] where δ is a hyperparameter;
[0047] Divide the 3D point set {X l} and the key frames {T k ′} into independent data blocks and allocate them to different thread blocks of the GPU. Subsequently, update the 3D point coordinates and the key frame poses to the final optimized values, and output the optimized and updated pose T k ′ and the optimized and updated sparse 3D point cloud X l , which includes vertex coordinates, colors, and normal vectors. Use the triangulation positioning method to recover the 3D point cloud of the spatial 3D point set and generate a global 3D point cloud model covering the functional area of the railway passenger station to achieve 3D scene reconstruction.
[0048] Preferably, the method of generating hash codes using multiple groups of locality-sensitive hashing functions, using the global 3D point cloud model as a physical space mapping benchmark to screen candidate images with high similarity to the passenger's perspective image through a voting mechanism, and obtaining a preliminary pose estimate of the passenger's perspective, includes:
[0049] Use 4 groups of locality-sensitive hashing functions Each group generates a 12-bit hash code:
[0050] h i (x) = sign(w i x + b i ) (7)
[0051] where, is a 384-dimensional mixed feature vector of the input image, w i is a random projection vector following the standard normal distribution, with a dimension of 384×1, b i is a uniformly distributed random bias term used to adjust the separation threshold of the hash function, and sign(·) is the sign function that maps real values to binary hash codes of ±1;
[0052] For the image I q captured by the passenger in real time, extract the feature F q through the DenseNet-121 architecture, use the principal component analysis method to reduce the dimension of the feature vector to 384 dimensions, and calculate 4 groups of hash codes {h q i (F(F q )}, and count the number of hash codes in the hash table that are the same as F qReference images for hash code matching, sorted by the number of matching hash tables:
[0053]
[0054] Among them, is an indicator function, which is 1 if the hash codes are equal and 0 otherwise. Retain the candidate image set C = {(F j , T j )} with a retained score S ≥ 3. Calculate the cosine similarity between the candidate image set C and the feature vector of F q . Select the images with the highest similarity and a set number as the final matching result. Align the poses {T j} of the finally matched images and the spatial positions of the passenger's current perspective in the global three-dimensional point cloud model. Using the cosine similarity results of the known multiple high-similarity candidate reference image perspectives and the passenger perspective image poses as weights, generate weights through the Softmax function according to the cosine similarity results and perform weighted averaging to obtain the weighted pose T f :
[0055]
[0056] Take the weighted pose as the preliminary pose estimate T f = [R f |t f , where R f represents the result of converting the rotation matrix of the matching image to a quaternion q j after weighted averaging and then converting back to the rotation matrix, is the weighted result of the translation vector. Screen through a voting mechanism
[0057] Preferably, according to the camera internal parameter matrix and rotation matrix of the passenger perspective, perform smoothing processing on the preliminary pose estimate of the passenger perspective through Kalman filtering and output the positioning coordinates of the passenger, including:
[0058] Take the position p init = [x init , y init , z init T and the initial rotation quaternion q init = [q w , q x , q y , q z T as the initial position and attitude state of the passenger, and combine the velocity component vectors v = [v x , v y , v z T The state vector x = [p, v, q] is formed. T Assuming that the passenger moves uniformly or remains stationary in a short period of time, the state update at the next moment a+1 is as follows:
[0059] p a+1 = p a + v a Δt (9)
[0060] v a+1 = v a (10)
[0061]
[0062] where ω is the angular velocity, and the weighted pose T f is used as the observation value z k . By fusing the Kalman gain, the predicted value and the observation value are obtained through the state update formula:
[0063] K a+1 = P a+1|a + H T (HP a+1|a H T + R a+1 ) -1 (12)
[0064] x a+1 = x a+1|a + K a+1 (z k - Hx a+1|a ) (13)
[0065] P a+1 = (I - K a+1 H)P a+1|a (14)
[0066] where z k - Hx a+1 represents the difference between the observation and the prediction. K f is used to balance the trust levels of the predicted value and the observation value. R a+1 is the observation noise. After multiple iterations, P a+1 gradually converges, and the state estimation tends to be stable. Finally, the state vector x a+1 including the final position p a+1 of the passenger and the attitude is output.
[0067] As can be seen from the technical solutions provided by the embodiments of the present invention described above, the method of the present invention applies passive positioning to reduce costs: without deploying external beacon devices, positioning can be achieved relying on in-station visual information, significantly reducing deployment and maintenance costs, and having great potential for wide promotion. It can provide accurate navigation support for passengers, and at the same time provide passenger flow data for the station to optimize operation efficiency.
[0068] Additional aspects and advantages of the present invention will be given in part in the following description, and these will become apparent from the following description, or can be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0070] Figure 1 It is a processing flow chart of a method for passive positioning of passengers in a railway passenger station based on multi-modal image fusion provided by an embodiment of the present invention.
[0071] Figure 2 It is an overall implementation block diagram of a method for passive positioning of passengers in a railway passenger station based on multi-modal image fusion and parallel three-dimensional reconstruction of the present invention.
[0072] Figure 3 It is a structural diagram of a graph feature encoding network provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0073] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals indicate the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.
[0074] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the description of the present invention means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or coupling. The phrase "and / or" used herein includes any unit and all combinations of one or more of the associated listed items.
[0075] Those skilled in the art can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention pertains. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless defined as here.
[0076] For the convenience of understanding the embodiments of the present invention, the following will further explain with several specific embodiments as examples in conjunction with the accompanying drawings, and each embodiment does not constitute a limitation to the embodiments of the present invention.
[0077] A passive positioning system for passengers in a railway passenger station provided by an embodiment of the present invention is composed of a multi-modal image data acquisition module, an image hybrid feature extraction and adaptive bag-of-words tree construction module, a parallelized three-dimensional reconstruction module, a multi-hash table joint retrieval module, and a dynamic error compensation positioning module. The connection relationships of each module are as follows:
[0078] The multi-modal image data acquisition module is deployed in key areas of the station (such as ticket checking gates, waiting halls), and synchronously acquires RGB images, depth maps, and inertial data through an RGB (Red (R), Green (G), Blue (B))-D camera and an IMU (Inertial Measurement Unit) sensor, and outputs them to the hybrid feature extraction module after spatio-temporal alignment.
[0079] The hybrid feature extraction module extracts local features and semantic features from the input images, constructs an adaptive bag-of-words tree, and outputs the bag-of-words tree to the parallelized three-dimensional reconstruction module.
[0080] The parallelized 3D reconstruction module divides the bag-of-words tree into subsets and uses the GPU (Graphics Processing Unit) to accelerate the generation of a global 3D point cloud model, which is used as a physical space mapping benchmark and input into the multi-hash table joint retrieval module.
[0081] The multi-hash table joint retrieval module receives the real-time images taken by passengers, and based on the physical space mapping benchmark, filters candidate images through hash coding and lightweight matching, and outputs the matching pose to the dynamic error compensation module.
[0082] The dynamic error compensation and positioning module dynamically smooths the pose through Kalman filtering, and finally outputs the coordinates of passengers with sub-meter accuracy.
[0083] The processing flow chart of a passive positioning method for passengers in railway passenger stations based on multi-modal image fusion provided by the embodiments of the present invention is as Figure 1 shown, including the following processing steps;
[0084] Step S1: Collect and construct a multi-modal image data set through RGB-D cameras and inertial sensor devices in the functional areas of high-speed railway stations, enhance data diversity through random occlusion and affine transformation, and record the shooting parameters of each image and the longitude and latitude coordinates in the station coordinate system, providing high-quality input for subsequent feature extraction and 3D reconstruction.
[0085] Deploy multiple hardware required for the method of the present invention in the functional areas (such as the entrance and waiting hall) of railway passenger stations, including multiple groups of RGB-D cameras (such as Intel RealSense D455), with a spacing of 5-10 meters, covering a 360° view; IMUs (such as BMI160) are integrated into the cameras, with a sampling frequency of 100Hz. Each RGB camera and IMU synchronously shoot data to construct a multi-modal image data set.
[0086] Ensure the alignment of the timestamps of RGB images, depth maps, and IMU data through hardware trigger signals. The spatial coordinate system takes the center of the station as the origin, the z-axis is perpendicular upward, and the x-axis and y-axis are parallel to the station ground. Perform data enhancement on the image data in the multi-modal image data set, randomly generate rectangular occlusion areas in the images, with the occlusion area accounting for 10%-30%, and the texture of the occluder is simulated with Gaussian noise to simulate dynamic interference. In the random occlusion area (rectangular area), the noise value of each pixel follows a Gaussian distribution where the mean μ = 0, the standard deviation, and is dynamically adjusted according to the lighting conditions of the station. Perturb the pixel value I(x,y) of the image in the occlusion area:
[0087]
[0088] Apply a random affine transformation matrix \(A\) to each image simultaneously to obtain the data-augmented image, which is defined as:
[0089]
[0090] where the scaling factor \(s\in[0.8,1,2]\), the rotation angle \(\theta\in[-30^{\circ},30^{\circ}]\), and the translation amount \(t\) x ,t y \(\in[-50,50]\) pixels.
[0091] Step S2: Extract local features and semantic features from the data-augmented image and construct an adaptive bag-of-words tree.
[0092] Use the directional FAST (Five-hundred-meter Aperture Spherical radio Telescope) corner detection algorithm and rotation-invariant descriptors to extract the local features of the image, generating two types of binary feature vectors. Combine the self-supervised model to extract the vector semantic features respectively, and perform a concatenation operation to obtain the mixed features. Based on the spatial clustering algorithm, classify and cluster the mixed features, and adaptively eliminate the low-stability branches to reduce redundant calculations and improve the adaptability of the bag-of-words tree to complex scenes.
[0093] Detect FAST (Five-hundred-meter Aperture Spherical radio Telescope) corners for each data-augmented image, calculate BRIEF (Binary Robust Independent Elementary Features) descriptors, and combine rotation-invariant descriptors to generate 256-dimensional binary feature vectors. The direction angle \(\theta\) compensation formula for the above rotation-invariant descriptor is:
[0094]
[0095] In the formula, since the images used in the present invention are color images, \(I(x,y)\) is the image pixel intensity value, and \(x,y\) represent the two-dimensional coordinate indices of the pixels in the image or neighborhood. Use the intensity values of all pixels in the neighborhood to perform a weighted sum on the direction component, and then obtain the main direction angle \(\theta\) of the neighborhood through the binary arctangent function.
[0096] Subsequently, a pre-trained SuperPoint semantic model is used to extract 128-dimensional floating-point vectors by inputting the original RGB image. The decoder of the semantic model contains 4 convolutional blocks (each block contains a convolutional layer, ReLU (Rectified Linear Unit), and pooling). The output channel number of the key point detection head in the decoder is 65, and the key point probability map is generated through Softmax. The descriptor generation head outputs a 128-dimensional floating-point descriptor after Min-Max normalization. Through the above steps, the resolution of the feature map is finally downsampled to 1 / 8 of the original image. For each image, a 128-dimensional semantic feature vector is extracted The above 256-dimensional binary feature vectors of N images are concatenated with the SuperPoint features to form a 384-dimensional hybrid feature vector
[0097]
[0098] In the stage of constructing the adaptive bag-of-words tree, the HDBSCAN algorithm is used for the above hybrid feature vectors for hierarchical clustering. First, the core distance is constructed based on k-nearest neighbors to generate a weighted adjacency graph, then the minimum spanning tree is generated through the Prim algorithm, and pruning is performed based on the mutual reachability distance. The clustering hierarchy in the tree is extracted, the cluster stability λ (determined through cross-validation) is calculated, the ratio of the number of samples within the cluster to the cluster diameter is calculated, the clusters with λ < 0.7 are removed, and the clusters with high stability are merged to balance the clustering purity and coverage. Finally, a hierarchical bag-of-words tree is obtained, and each leaf node in the bag-of-words tree stores the cluster center feature vector and the spatial coordinates of the corresponding image
[0099] Step S3: Use the GPU (Graphics Processing Unit) to divide the bag-of-words tree into subsets to accelerate the calculation of the global rotation matrix. Combine the key frame dynamic screening mechanism to select key frames with high matching quality and wide coverage. Optimize the projection error through the sparse bundle adjustment method to generate a global three-dimensional point cloud model, and input the global three-dimensional point cloud model as a physical space mapping benchmark into the multi-hash table joint retrieval module
[0100] The hierarchical cluster structure of the bag-of-words tree, the hybrid feature vector and the spatial coordinates of its corresponding image, as well as the camera intrinsic matrix K and the initial estimate T of the shooting pose i =[R i |t i are input into the GPU. The GPU divides the bag-of-words tree into C subsets {a1, a2,..., a max} according to spatial proximity (Euclidean distance threshold d c = 10m) and feature similarity (cosine similarity greater than 0.8), and each subset contains local features and semantic features
[0101] In the global rotation matrix optimization stage, first estimate the relative rotation from the local camera pairs. The GPU extracts the matching feature points {(p i ,a j )} for each subset pair (a i ,p j ) obtained by dividing each word bag tree, and calculates the essential matrix E ij through the five-point method, and decomposes it to obtain the relative rotation matrix Taking the minimization of the rotation matrix difference between all subset pairs as the objective function to ensure that the rotation matrices of multiple sets of camera views are consistent in the global coordinate system, the optimization objective function is:
[0102]
[0103] where R i ,R j are the rotation matrices of cameras i and j respectively, is the estimated value of the rotation matrix between the two cameras, and ||·|| F represents the Frobenius norm. In the actual implementation process, the subset pairs are assigned to GPU thread blocks, and each thread block calculates one term to achieve parallel computing. Finally, the optimized global rotation matrix R i is obtained and stored as an NPY format file.
[0104] In the key frame dynamic screening and initialization stage, set the number of matching image points to be no less than 50 to ensure sufficient geometric constraints. The scene coverage rate is not less than 60%, covering the main structures of the current subset. Select the frame with the largest feature distribution entropy to avoid redundancy. Subsequently, combine the optimized global rotation matrix R i ′ obtained previously with the initial translation t i to obtain the key frame pose T i ′ = [R′ i |t i .
[0105] After completing the above process, use the key frame pose T i ′, the sparse 3D points X i calculated from the matching feature point pairs {(p j ,p l )} through multi-view geometric triangulation, the camera internal parameter K, and the feature point matching relationship O = {(k, l, x kl )} as inputs. Where x kl represents the pixel coordinate observation value of the 3D point X l in the key frame T k ′. Construct the objective function by minimizing the sum of the reprojection errors, and optimize all key frame poses and 3D points simultaneously:
[0106]
[0107] where π([x, y, z] T ) = [x / z, y / z] T is the perspective projection function, and π(KT k ′·X l ) - x kl is defined as the projection error e kl between each observation (k, l), and ρ is the Huber robust kernel function, defined as:
[0108]
[0109] where δ is a hyperparameter, and the optimal value determined by cross - validation is 2.
[0110] The 3D point set {X l} and the key frames {T k ′} are divided into independent data blocks and assigned to different thread blocks of the GPU. For example, each thread block processes the Jacobian calculation of 128 3D points, and each thread block processes the pose update of 64 key frames. The specific process is as follows: load the initial pose {T k ′} and the 3D point set {X l}, and pre - calculate the projection error e kl of all observations. Calculate all error terms e kl and their Jacobian matrices in parallel. Accumulate the H and g of each thread block into the global matrix, where H and g are defined as the second - order approximate Hessian matrix and the gradient vector of the objective function. Call the CUDA library to solve the linear equation HΔξ = -g, where Δξ represents a 6 - dimensional vector of the camera pose increment, with the first 3 dimensions representing the rotation increment and the last 3 dimensions representing the translation increment. If the relative error change is less than 1×10 -6 then terminate. Subsequently, update the 3D point coordinates and the key frame poses to the final optimized values, and output the optimized and updated pose T k ′ (CSV file) and the optimized and updated sparse 3D point cloud X l (PLY format), which includes vertex coordinates, colors, and normal vectors. Use the triangulation positioning method to perform 3D point cloud recovery on the spatial 3D point set, generate a global 3D point cloud model covering the functional area of the railway passenger station, and realize 3D scene reconstruction.
[0111] Step S4: Generate hash codes using multiple groups of locality - sensitive hashing functions respectively, screen high - similarity candidate images through a voting mechanism, and calculate the preliminary positioning result based on the weighted result of the candidate image set poses.
[0112] First, use the DenseNet-121 architecture to extract image features from the multimodal image dataset. The DenseNet-121 architecture consists of 3 dense blocks with 3 layers of convolution, and the dense blocks are connected by transition layers composed of 1×1 convolution and 2×2 average pooling layers. The multimodal image dataset collected in step S1 is used as the input, and image feature encoding is performed through DenseNet-121. Each image finally generates a 1024-dimensional feature vector after global average pooling. After that, the 1024-dimensional feature vector extracted based on the DenseNet-121 model is input. and the image dataset D = {(F i , T i )}.
[0113] Use 4 groups of Locality-Sensitive Hashing (LSH) functions Each group generates a 12-bit hash code:
[0114] h i (x) = sign(w i x + b i ) (7)
[0115] where is the 384-dimensional mixed feature vector of the input image, w i is a random projection vector following the standard normal distribution, with a dimension of 384×1, and b i is a uniformly distributed random bias term used to adjust the separation threshold of the hash function. sign(·) is the sign function that maps real values to binary hash codes of ±1.
[0116] Adopt the same method as the image feature extraction in the multimodal image dataset to extract the feature vector F q from the image I q captured by the passenger in real time through DenseNet-121, and use the principal component analysis method to reduce the dimension of the feature vector to 384 dimensions, and calculate 4 groups of hash codes {h q (F i )} of F q . Subsequently, count the reference images in the hash table that match the hash code of F q , and sort them according to the number of matching hash tables:
[0117]
[0118] where is the indicator function, which is 1 if the hash codes are equal, otherwise 0. Retain the candidate image set C = {(F j , T j )} with a score S≥3. Further perform exact matching, and calculate the similarity between the candidate image set C and Fq The cosine similarity of the feature vectors, and select the top-5 images with the highest similarity as the final matching results. For the poses {T j} of the top-5 matching images and the spatial positions of the passenger's current perspective in the global three-dimensional point cloud model, perform unified alignment. Use the cosine similarity results between the known perspectives of multiple highly similar candidate reference images and the passenger perspective image poses as weights, generate weights through the Softmax function according to the cosine similarity results and perform weighted averaging to obtain the weighted pose T f :
[0119]
[0120] Take the above weighted pose as the initial pose estimate T of the passenger's perspective f =[R f |t f , where R f represents the result of converting the rotation matrix of the matching image to a quaternion q j after weighted averaging and then converting back to the rotation matrix, is the weighted result of the translation vector.
[0121] To improve the accuracy and robustness of positioning, introduce the global three-dimensional point cloud model as the mapping benchmark of the physical space, and use the correspondence between the three-dimensional point cloud and the image feature points to optimize the pose. For the top-5 matching images, obtain the correspondence between their two-dimensional feature points and the three-dimensional points. For the passenger image, we use the SIFT algorithm to extract its feature points By matching with the candidate image feature points, indirectly find the corresponding three-dimensional point p l , and obtain the correspondence between the passenger perspective image feature points and the three-dimensional points
[0122] and optimize the pose by minimizing the reprojection error
[0123]
[0124] where π is the Huber kernel function. Use the Levenberg-Marquardt algorithm to solve the non-linear least squares problem to obtain the optimized pose T refined =[R refined |t refined .
[0125] Step S5: According to the camera internal parameter matrix and the rotation matrix of the passenger's perspective, based on the shooting pose of the current perspective obtained by the previous weighting, perform smoothing processing on the initial pose estimate of the passenger's perspective through Kalman filtering, and output the passenger coordinates.
[0126] Use the pose estimate T representing the passenger's initial perspective obtained from the previous stepsrefined is the input. The position p init = [x init , y init , z init T and the initial rotation quaternion q init = [q w , q x , q y , q z T are used as the initial position and attitude state of the passenger. Combining the velocity component vectors v = [v x , v y , v z in three directions T constitutes the state vector x = [p, v, q] T . Assuming that the passenger moves at a constant speed or remains stationary in a short period of time, the state at the next moment a + 1 is updated as follows:
[0127] p a+1 = p a + v a Δt (11)
[0128] v a+1 = v a (12)
[0129]
[0130] where ω is the angular velocity. Subsequently, the weighted pose in step S4 is used as the observation value z k , and the predicted value obtained through the state update formula is fused with the observation value through the Kalman gain:
[0131] K a+1 = P a+1|a + H T (HP a+1|a H T + R a+1 ) -1 (14)
[0132] x a+1 = x a+1|a + K a+1 (z k - Hx a+1|a ) (15)
[0133] P a+1 = (I - K a+1 H)P a+1|a (16)
[0134] where z k - Hx a+1 Indicates the difference between the observation and the prediction, K f Used to balance the confidence between the predicted value and the observed value, R a+1 Is the observation noise. If it is small, the gain K f Increases, and the system trusts the observed value more, quickly correcting the prediction deviation. After multiple iterations, P a+1 Gradually converges, the state estimation tends to be stable, and finally the state vector x including the final position p of the passenger after the state update is output a+1 And the attitude is output a+1 .
[0135] In summary, the passive positioning of the embodiment of the present invention reduces costs: there is no need to deploy external beacon devices, and positioning can be achieved relying on in-station visual information, significantly reducing deployment and maintenance costs, and having great potential for popularization;
[0136] Efficient 3D reconstruction: By parallelizing the hierarchical structure from motion algorithm and GPU acceleration, the efficiency bottleneck of traditional reconstruction methods in large-scale scenes is overcome, the reconstruction time is shortened, and it adapts to the complex environment of high-speed railway stations;
[0137] High precision and robustness: Multi-modal image fusion improves the feature expression ability, and the combination of multi-hash table retrieval and Kalman filtering controls the positioning error within the sub-meter level, improving the matching accuracy;
[0138] Strong real-time performance: The single positioning time of lightweight retrieval is reduced compared with the traditional image retrieval algorithm, meeting the real-time navigation requirements;
[0139] Double improvement in service and management: Provide accurate navigation support for passengers, and at the same time provide passenger flow data for the station to optimize the operation efficiency.
[0140] Those of ordinary skill in the art can understand that the drawings are only schematic diagrams of one embodiment, and the modules or processes in the drawings are not necessarily essential for implementing the present invention.
[0141] From the description of the above embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.
[0142] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device or system embodiments, since they are basically similar to the method embodiments, they are described relatively simply. For the relevant parts, reference can be made to the descriptions in the method embodiments. The device and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution in this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.
[0143] As mentioned above, the above is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A passive positioning method for passengers in railway passenger stations based on multimodal image fusion, characterized in that, Including: Collect and construct a multi-modal image dataset through an RGB-D camera and an inertial sensor device within the functional area of a high-speed railway station, and perform data augmentation processing on the images in the multi-modal image dataset through random occlusion and affine transformation; Extract local features and semantic features from the data-augmented images, and construct an adaptive bag-of-words tree; Use a graphics processing unit (GPU) to divide the bag-of-words tree into subsets, select key frames using a key-frame dynamic screening mechanism, and optimize the projection error through sparse bundle adjustment to generate a global three-dimensional point cloud model; Generate hash codes using multiple groups of locality-sensitive hashing functions, use the global three-dimensional point cloud model as a physical space mapping benchmark, and screen candidate images with high similarity to the passenger's perspective images through a voting mechanism to obtain a preliminary pose estimate of the passenger's perspective; According to the camera internal parameter matrix and rotation matrix of the passenger's perspective, perform smoothing processing on the preliminary pose estimate of the passenger's perspective through Kalman filtering, and output the positioning coordinates of the passenger.
2. The method according to claim 1, characterized in that, The step of collecting and constructing a multi-modal image dataset through an RGB-D camera and an inertial sensor device within the functional area of a high-speed railway station, and performing data augmentation processing on the images in the multi-modal image dataset through random occlusion and affine transformation, includes: Deploy an RGB camera and an inertial measurement unit (IMU) within the functional area of a railway passenger station, ensure the timestamp alignment of RGB images, depth maps, and IMU data through a hardware trigger signal, use the station center as the origin of the spatial coordinate system, with the z-axis perpendicular upward and the x-axis and y-axis parallel to the station ground, and synchronously capture data with each RGB camera and IMU to construct a multi-modal image dataset; Perform data augmentation on the image data in the multi-modal image dataset. Randomly generate rectangular occlusion regions in the image, and simulate dynamic interference for the occluder texture using Gaussian noise. Within the randomly occluded region, the noise value of each pixel follows a Gaussian distribution where the mean μ = 0, and perturb the pixel value I(x, y) of the image within the occluded region: Apply a random affine transformation matrix A to each image to obtain the data-augmented image, and the random affine transformation matrix A is defined as: where the scaling factor s ∈ [0.8, 1, 2], the rotation angle θ ∈ [-30°, 30°], and the translation amount t x , t y ∈ [-50, 50] pixels.
3. The method according to claim 2, characterized in that, The step of extracting local features and semantic features from the data-augmented images, and constructing an adaptive bag-of-words tree, includes: Detect FAST corners of the spherical radio telescope for each data-augmented image, calculate BRIEF descriptors, and generate a 256-dimensional binary feature vector using a rotation-invariant descriptor. The direction angle θ compensation formula of the rotation-invariant descriptor is: I(x,y) is the image pixel intensity value, and x,y represent the two-dimensional coordinate indices of the pixel in the image or neighborhood; Extract a 128-dimensional semantic feature vector from the original RGB image using a semantic model The 256-dimensional binary feature vectors of N images and the 128-dimensional semantic feature vector Concatenate into a 384-dimensional hybrid feature vector The HDBSCAN algorithm is used to perform hierarchical clustering on the mixed feature vectors to construct a core distance based on k-nearest neighbors, generate a weighted adjacency graph, then generate a minimum spanning tree through the Prim algorithm, prune based on the mutual reachability distance, extract the clustering hierarchy in the tree, calculate the cluster stability λ, calculate the ratio of the number of samples within the cluster to the cluster diameter, remove clusters with λ < 0.7, merge them to form clusters with high stability, obtain a hierarchical bag-of-words tree, and each leaf node in the bag-of-words tree stores the cluster center feature vector and the spatial coordinates of the corresponding image.
4. The method according to claim 3, wherein The step of using a graphics processing unit (GPU) to divide the bag-of-words tree into subsets, select key frames using a key-frame dynamic screening mechanism, and optimize the projection error through sparse bundle adjustment to generate a global three-dimensional point cloud model, includes: The hierarchical cluster structure of the bag-of-words tree, the hybrid feature vector and the spatial coordinates of its corresponding image, as well as the camera intrinsic matrix K and the initial estimate T of the shooting pose i = [R i | t i are input into the GPU, and the GPU divides the bag-of-words tree into C subsets {a1, a2,..., a c} according to spatial proximity and feature similarity, and each subset contains local features and semantic features; The GPU extracts the matching feature points \(\{(p i , p j )\}\) for each subset pair \((a i , a j )\), calculates the essential matrix \(E ij \) by the five-point method, and decomposes the essential matrix \(E ij \) to obtain the relative rotation matrix Taking the minimization of the rotation matrix difference between all subset pairs as the objective function, the optimization objective function is: where, R i , R j are the rotation matrices of cameras i and j respectively, is the estimated value of the rotation matrix between the two cameras, ||·|| F represents the Frobenius norm, and the optimized global rotation matrix R i ; Combine the optimized global rotation matrix R′ i with the initial translation t i to obtain the key-frame pose T i ′ = [R′ i | t i . Based on the key-frame pose T i ′ and the matching feature point pairs {(p i , p j )}, calculate the sparse 3D points X l , the camera internal parameter K, and the feature point matching relationship O = {(k, l, x kl )}, where x kl represents the pixel coordinate observation value of the 3D point X l in the key frame T′ k . Construct an objective function to minimize the sum of the reprojection errors, and simultaneously optimize all key-frame poses and 3D points: where π([x,y,z] T ) = [x / z,y / z] T is the perspective projection function, π(KT′ k ·X l ) - x kl is defined as the projection error e between each observation (k, l) kl , and ρ is the Huber robust kernel function, defined as: where δ is a hyperparameter; Divide the three-dimensional point set {X l} and the key frames {T′ k} into independent data blocks and allocate them to different thread blocks of the GPU. Subsequently, update the three-dimensional point coordinates and the key frame poses to the final optimized values, and output the optimized and updated pose T′ k and the optimized and updated sparse three-dimensional point cloud X l , which includes vertex coordinates, colors, and normal vectors. Use the triangulation positioning method to perform three-dimensional point cloud recovery on the three-dimensional point set in space, generate a global three-dimensional point cloud model covering the functional areas of the railway passenger station, and realize three-dimensional scene reconstruction.
5. The method according to claim 4, wherein The step of generating hash codes using multiple groups of locality-sensitive hashing functions, using the global three-dimensional point cloud model as a physical space mapping benchmark, and screening candidate images with high similarity to the passenger's perspective images through a voting mechanism to obtain a preliminary pose estimate of the passenger's perspective, includes: Use 4 sets of locality-sensitive hashing functions Each set generates a 12-bit hash code: h i (x) = sign(w i x + b i ) (7) Among them, is a 384-dimensional mixed feature vector of the input image, and w i is a random projection vector following a standard normal distribution, with a dimension of 384×1. b i is a uniformly distributed random bias term used to adjust the separation threshold of the hash function. sign(·) is the sign function that maps real values to binary hash codes of ±1; The image I captured in real time for passengers q , extract the feature F through the DenseNet-121 architecture q , use the principal component analysis method to reduce the dimensionality of the feature vector to 384 dimensions, and calculate F q 's 4 groups of hash codes {h i (F q )}, count the reference images in the hash table that match the hash code of F q , and sort them according to the number of matching hash tables: Among them, is an indicator function, which is 1 if the hash codes are equal and 0 otherwise. Retain the candidate image set C = {(F j , T j )} with a retention score S ≥ 3. Calculate the cosine similarity between the candidate image set C and the feature vector of F q . Select the images with the highest similarity and a set number as the final matching result. Align the poses {T j} of the finally matched images and the spatial positions of the passenger's current perspective in the global three-dimensional point cloud model. Use the cosine similarity results between the known perspectives of multiple highly similar candidate reference images and the passenger's perspective image poses as weights. Generate weights through the Softmax function according to the cosine similarity results and perform weighted averaging to obtain the weighted pose T f : Use the weighted pose as the initial pose estimate \(T\) from the perspective of the passenger f = [R f | t f , where \(R\) f represents the result of converting the rotation matrix of the matched image to a quaternion \(q\) j after weighted average and then converting back to a rotation matrix, and \(t\) is the weighted result of the translation vector 6. The method according to claim 5, wherein The step of performing smoothing processing on the preliminary pose estimate of the passenger's perspective through Kalman filtering according to the camera internal parameter matrix and rotation matrix of the passenger's perspective, and outputting the positioning coordinates of the passenger, includes: Position p init = [x init , y init , z init T and the initial rotation quaternion q init = [q w , q x , q y , q z T as the initial position and attitude state of the passenger, combined with the velocity component vectors v = [v x , v y , v z T to form the state vector x = [p, v, q] T , assuming that the passenger moves at a constant speed or remains stationary in a short period of time, the state at the next moment a + 1 is updated as: p a+1 = p a + v a Δt (9) v a+1 = v a (10) where ω is the angular velocity, and the weighted pose T f is used as the observation value z k , and the predicted value and the observation value are obtained through the state update formula by fusing with the Kalman gain: K a+1 = P a+1|a + H T (HP a+1|a H T + R a+1 ) -1 (12) x a+1 = x a+1|a + K a+1 (z k - Hx a+1|a ) (13) P a+1 = (I - K a+1 H)P a+1|a (14) where z k -Hx a+1 represents the difference between the observation and the prediction, K f is used to balance the confidence between the predicted value and the observed value, R a+1 is the observation noise. After multiple iterations, P a+1 gradually converges, and the state estimation tends to be stable. Finally, the state vector x that includes the final position p of the passenger after the state update is output a+1 and the attitude a+1 .
Citation Information
Patent Citations
Indoor positioning method and related equipment
CN119584054A
Ciphertext image retrieval method and system under a cloud environment
CN108959478A
Image processing method and device and readable storage medium
CN109993201A
Instant positioning and map construction system and method with semantic perception
CN111968129A
Binocular SLAM three-dimensional reconstruction and positioning based on self-supervised nerve radiation field
CN118918279A