Vision and 5G sensing fused three-dimensional indoor positioning method

By calibrating 5G millimeter-wave base stations and RGB cameras in a unified coordinate system, three-dimensional modeling of indoor reflective surfaces is performed. By optimizing channel state information using visual target detection and the ESPRIT algorithm, and combining direct and reflected path constraints, high-precision three-dimensional indoor dynamic positioning with low cost and low complexity is achieved, solving the problems of beam mismatch and visual system complexity in existing technologies.

CN121865206APending Publication Date: 2026-04-14HANGZHOU INTERNATIONAL INNOVATION INSTITUTE OF BEIHANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-23
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing CSI positioning solutions based on 5G millimeter waves suffer from beam mismatch between beam direction and user target location in dynamic scenarios, making real-time high-precision positioning difficult. Multi-view stereo vision positioning systems are costly and complex, making them impractical for practical applications.

Method used

The 5G millimeter-wave base station and RGB camera were calibrated in a unified coordinate system, and a 3D model of the indoor reflective surface was performed. Visual target detection was used to generate a sensing beam, and the channel state information was optimized through the ESPRIT algorithm. 3D positioning was performed by combining direct and reflected path constraints.

Benefits of technology

It achieves high-precision, real-time 3D indoor dynamic positioning with low cost and low complexity, solves the problems of beam sensitivity and multi-view vision deployment complexity, and improves the measurability and stability of angle and time delay parameters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121865206A_ABST
    Figure CN121865206A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional indoor positioning method based on visual sense and 5G sensing fusion, and relates to the technical field of indoor positioning and space perception. The method comprises the following steps: modeling an indoor reflecting surface to obtain a virtual anchor point; a user azimuth vector detected, identified and tracked by a camera target is projected to a scene reflection surface to construct a candidate reflection path and identify a direct incidence path, and the candidate reflection path and the direct incidence path are mapped to a beam codebook of a base station, so that a beam search range is narrowed, and a higher angle resolution is obtained; the user terminal establishes a communication link through the reference signal, and the base station acquires channel state information; in the channel parameter processing stage, angle and time delay estimation is carried out on channel parameters; an estimation result is matched with a direct or virtual anchor point path, and pseudo paths which do not conform to scene constraints are eliminated, so that a direct path and an effective reflection path are distinguished; the position of the user is jointly solved by using multi-angle constraints formed by the direct path and the reflection path, and continuous estimation of the three-dimensional position of the user is realized by combining a filter.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of indoor positioning and spatial perception technology, and in particular to a three-dimensional indoor positioning method that integrates vision and 5G sensing. Background Technology

[0002] In existing technologies, CSI positioning schemes based on 5G millimeter waves rely on channel state information provided by the base station to extract the angle and delay parameters of multipath and calculate the terminal position based on geometric relationships. In addition, multi-view stereo vision positioning systems can directly solve the three-dimensional coordinates through parallax triangulation of multiple cameras, thus avoiding the beam management problem of the aforementioned wireless positioning.

[0003] However, the perception performance of CSI positioning schemes based on 5G millimeter waves is extremely sensitive to beam pointing. The time misalignment between the training and application phases causes a mismatch between the beam direction and the user's target position in dynamic scenarios, making it impossible to meet the requirements of real-time high-precision dynamic positioning. In addition, in order to achieve wide coverage, multi-view stereo vision positioning systems require the user to always be in the overlapping field of view of multiple cameras. Therefore, a large number of cameras must be densely deployed, which causes the system cost and complexity to increase exponentially, making it impossible to implement in the practical application phase.

[0004] Therefore, there is an urgent need for an indoor positioning method that can guarantee real-time, high-precision dynamic indoor positioning under the premise of low cost and low complexity. Summary of the Invention

[0005] Therefore, it is necessary to provide a three-dimensional indoor positioning method that integrates vision and 5G sensing to address the aforementioned technical problems.

[0006] The present invention adopts the following technical solution: This invention provides a three-dimensional indoor positioning method that integrates vision and 5G sensing, comprising: The location of the 5G millimeter-wave base station and the internal and external parameters of the RGB camera were determined in a unified coordinate system; A 3D model of the reflective surface of the building's indoor environment is performed to obtain the plane equation of the indoor reflective surface; and the virtual anchor point coordinates of the 5G millimeter-wave base station location on the reflective surface of the building's indoor environment are calculated. The system captures images of the user's indoor environment using a calibrated RGB camera, performs target detection on the user's indoor environment images using the YOLOv7 model, obtains the target user image and its pixel representative position, and back-projects it onto the world coordinate system to obtain the camera's line of sight direction. Candidate user location points are selected in the user image along the camera's line of sight. Based on the coordinates of the candidate user location points and virtual anchor points, and combined with the indoor reflective surface plane equation, multiple reflection points of the candidate user location points are calculated to obtain the user perception beam. With the minimum of the Cramer-Rao lower bound of the variance of the position parameter estimation of the reflection points as the optimization objective, the transmission energy in the direction of the user perception beam is optimized using an energy-based adaptive allocation mechanism to obtain the optimized user perception beam. Based on the optimized user-aware beam, channel state information is obtained; and the ESPRIT algorithm is used to perform high-resolution inversion on the channel state information to obtain the multipath arrival angle and propagation delay. Based on the angle of arrival and propagation delay, the direct path and effective reflection path of the user's sensing beam are determined; using the intersection of the directional constraints formed by the direct path and the effective reflection path, the user's millimeter-wave positioning result is determined.

[0007] Preferably, the reflective surfaces of the building's interior environment include: three-dimensional reflective surfaces of walls, floors, and other main structures; the equation of the three-dimensional reflective surfaces is described as follows: ; In the formula, The nth reflective surface of the building's interior environment. Let be the unit normal vector of the reflecting surface. This is the intercept of the reflecting surface. The total number of valid paths involved in the positioning, including one direct path and The reflection path generated by the reflective surface.

[0008] Preferably, calculating the virtual anchor point coordinates of the 5G millimeter-wave base station location on the reflective surface of the indoor environment specifically includes: Substituting the location of the 5G millimeter-wave base station into the equation of the indoor reflective surface, and combining this with the location of the 5G millimeter-wave base station, the coordinates of the virtual anchor point are calculated using the following formula: ; In the formula, The coordinates of the virtual anchor point on the nth reflective surface of the 5G millimeter-wave base station location within the indoor environment are given. The coordinates of the 5G millimeter-wave base station location. This represents the reflective surface corresponding to the location of the nth 5G millimeter-wave base station. Let be the unit normal vector of the reflecting surface. This is the intercept of the reflecting surface.

[0009] Preferably, the indoor environment image of the user is subjected to target detection using the YOLOv7 model to obtain the user image and the user pixel representative position, which is then back-projected onto the world coordinate system to obtain the camera's line of sight direction, specifically including: The user's indoor environment image is input into the trained YOLOv7 model, which outputs the user image and the pixel positions of the vertices of the detection boxes. ; The user pixel position is determined by the center point of the pixel position of the detection box vertex. ;in, ; By back-projecting the user's pixel position, the camera's line of sight relative to the user in the world coordinate system is obtained, using the following formula: ; In the formula, The direction of the camera relative to the user's line of sight in the world coordinate system. For depth parameters, Let be the rotation matrix from the camera coordinate system to the world coordinate system. This is the camera intrinsic parameter matrix.

[0010] Preferably, candidate user location points are selected in the user image along the camera's line of sight. Based on the coordinates of the candidate user location points and virtual anchor points, and combined with the indoor reflective surface plane equation, multiple reflection points of the candidate user location points are calculated to obtain the user perception beam, specifically including: The user candidate location points are selected along the direction of the camera's line of sight, using the following formula: ; In the formula, To capture all candidate user locations along the user's line of sight in the world coordinate system. Let be the position vector of the camera in the world coordinate system. The direction of the camera relative to the user's line of sight in the world coordinate system. For depth parameters; Combining the virtual anchor point coordinates, user candidate location points, and the indoor reflective surface plane equation, the reflection points of multiple user candidate location points are determined using the following formula: , ; In the formula, Let be the coordinates of the reflection point of the nth user candidate location on the nth reflecting surface. Let the coordinates of the nth virtual anchor point be... For the nth candidate user location, , Let be the scale factor along the direction from the virtual anchor point to the user candidate point, and its value corresponds to the position of the intersection point of the ray and the nth reflection plane; The reflection point of the 5G millimeter wave base station location and the user candidate location point, and the beam formed by the user candidate location point are determined as the user perception beam.

[0011] Preferably, the objective function for optimizing the user-sensing beam is the trace of the inverse of the position estimation Fischer matrix. The constraint is that the transmit energy allocated to each user-sensing beam direction is less than or equal to the preset total transmittable energy, as shown in the formula: ; In the formula, To give the first The transmitted energy allocated to the beam direction perceived by each user. It is a set of directions that includes both direct and primary reflection directions. The preset total energy, Here is the Fischer matrix of the reflection point position parameters. This represents taking the trace by inverting the Fischer matrix, i.e., summing the lower bounds of the variances of the position estimates. For position parameters, P={ ,..., } .

[0012] Preferably, channel state information is obtained based on the optimized user-perceived beam, specifically including: The user terminal completes cell identification and time-frequency synchronization by receiving the user-perceived beam, and initiates a random access signal; After completing the access response and link establishment, the 5G millimeter-wave base station obtains channel state information using uplink or downlink reference signals, expressed as: ; In the formula, For the number of multipaths, For the first Complex gain of the user-perceived beam path, For the azimuth and elevation angles of the receiving end, For the azimuth and elevation angles of the transmitting end, For path delay, , These are the steering vectors for the receiving and transmitting arrays, respectively. For subcarrier frequency, This indicates performing a Hermitian transpose on the matrix.

[0013] Preferably, the ESPRIT algorithm is used to perform high-resolution inversion of the channel state information to obtain the multipath arrival angle and propagation delay, specifically including: ESPRIT is constructed by establishing shift-invariant relationships between adjacent subarrays or adjacent subcarriers, forming matrix pairs, as shown in the formula: ; In the formula, Let be the rotation-invariant matrix of the angle. Let be a rotation-invariant matrix in the time delay dimension. For the signal subspace matrix, The selection matrix is ​​the array dimension. The selection matrix is ​​for the frequency domain dimension; Calculate separately and The characteristic values ​​are used to obtain the arrival angle and propagation delay of the user-perceived beam.

[0014] Preferably, based on the angle of arrival and propagation delay, the direct path and effective reflection path of the user-sensing beam are determined; using the intersection of the directional constraints formed by the direct path and the effective reflection path, the millimeter-wave positioning result at the user end is determined, specifically including: Utilizing reflective surfaces Corresponding virtual anchor point Construct an equivalent direct path; Calculation by The distance of propagation along the direction of the departure angle. The obtained point; the propagation distance The path delay is calculated using the following formula: ; In the formula, For the first The propagation distance of the user-perceived beam. At the speed of light, For the first Path delay estimation; If the point is not within the indoor space enclosed by the reflective surface, the path is identified as a pseudo-path and is eliminated. If the path is within the indoor space enclosed by the reflecting surface at that point, then the path is determined to be related to the indoor environment's reflecting surface. The associated primary reflection path; After eliminating false paths, the direct path and effective reflection path of the user-perceived beam are obtained; The intersection of the direct path and the effective reflection path of the user-perceived beam is determined as the millimeter-wave positioning result, using the following formula: ; In the formula, For user location estimation, It is the identity matrix. The number of valid paths involved in positioning, including one direct path and several reflective paths. For the user location variable to be estimated, The starting point of the user-perceived beam is represented by the base station location b when it is the direct path and the virtual anchor point when it is the reflected path. , This is the direction vector obtained by the ESPRIT algorithm.

[0015] Preferably, the method further includes: The millimeter-wave positioning results are optimized using Kalman filtering, including: Define the state vector as follows: ; in, The three-dimensional position of the millimeter-wave positioning result. Given the average velocity, the state transition equation is: ; in The sampling interval is... For process noise, This is the state transition matrix; Based on the aforementioned state transition equation, and utilizing both millimeter-wave point observations and visual projection point observations, the observation equation is as follows: ; In the formula, To observe the noise, For the observation vector, The observation matrix; Visual observation, on the other hand, involves observing the millimeter-wave observation point along the optical center of the camera and the direction vector. Defined rays Projection yields: ; Construct the observation equation: ; in, For visual observation noise, For visual observation vectors, This is the visual observation matrix.

[0016] The above-mentioned at least one technical solution adopted in this invention can achieve the following beneficial effects: In the three-dimensional indoor positioning method that integrates vision and 5G sensing provided by this invention, by fusing 5G millimeter wave and visual perception, the calibration of base stations and cameras and the three-dimensional modeling of indoor reflective surfaces are completed in a unified coordinate system. Virtual anchor points are constructed to convert multipath signals into positioning resources. Then, the user's line of sight is obtained by visual target detection and back projection. Combined with the virtual anchor points, a sensing beam is generated. The beam energy allocation is dynamically optimized with the goal of minimizing the Cramer-Rao lower bound, compressing the full-space beam search range to at least a few geometrically salient directions, significantly shortening the beam training time. The adaptive energy allocation uses CRLB as the optimization criterion to ensure the generation of the optimal sensing beam, avoiding the contradiction between the communication optimal beam and sensing requirements in traditional schemes. This improves the measurability and stability of angle and delay parameters, meets the low-latency and high-robust positioning requirements in high-dynamic scenarios, and improves sensing accuracy and real-time performance while compressing the beam search range. The channel parameters are inverted at high resolution using the ESPRIT algorithm to distinguish between direct and effective reflection paths, and robust positioning is achieved by fusing multi-directional constraints. This method fundamentally solves the problems of beam sensitivity, training lag, and complex and limited coverage of multi-view vision deployment in 5G millimeter wave positioning. It achieves high-precision and high-real-time three-dimensional indoor dynamic positioning based on low cost and low complexity. Attached Figure Description

[0017] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0018] Figure 1 A schematic diagram of a three-dimensional indoor positioning method that integrates vision and 5G sensing provided by the present invention; Figure 2 A schematic diagram of dynamic beam management based on multipath direction estimation for a three-dimensional indoor positioning method that integrates vision and 5G sensing provided by the present invention; Figure 3 The diagram shows the effect of implementing a three-dimensional indoor positioning method that combines vision and 5G sensory fusion, as provided by this invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments in the specification without creative effort are within the scope of protection of this application.

[0020] This invention proposes an indoor 3D positioning method that integrates a 5G millimeter-wave base station and an RGB camera, assisted by 3D reflective surface modeling. The method first completes camera and base station location calibration during system deployment, and models the main reflective surfaces to obtain virtual anchor points. Subsequently, in the vision-assisted beam management stage, the user's azimuth vector, obtained from camera target detection, recognition, and tracking, is projected onto the scene's reflective surfaces to construct candidate reflection paths and identify direct paths. These paths are then mapped to the base station's beamcodebook, thereby narrowing the beam search range and achieving higher angular resolution. In the channel information acquisition stage, the terminal establishes a communication link using reference signals, and the base station acquires channel state information. In the channel parameter processing stage, a high-resolution algorithm is used to estimate the angle and delay of channel parameters. Finally, in the path geometric association stage, the estimated results are matched with direct or virtual anchor point paths based on geometric consistency rules, and pseudo-paths that do not conform to scene constraints are eliminated, thus distinguishing between direct paths and effective reflection paths. Finally, in the positioning and tracking phase, the user's position is jointly calculated using the multi-angle constraints formed by the direct path and the reflected path, and a filter is used to achieve continuous estimation of the user's three-dimensional position.

[0021] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.

[0022] Figure 1 This is a schematic diagram of a three-dimensional indoor positioning method that integrates vision and 5G sensing in this invention, specifically including the following steps: S101: Calibrate the location of the 5G millimeter-wave base station and the internal and external parameters of the RGB camera in a unified coordinate system.

[0023] S102: Perform a 3D model of the reflective surface of the building's indoor environment to obtain the plane equation of the indoor reflective surface; and calculate the virtual anchor point coordinates of the 5G millimeter wave base station location on the reflective surface of the building's indoor environment.

[0024] Optionally, the reflective surfaces of the building's interior environment include: three-dimensional reflective surfaces of walls, floors, and other major structures; the equations of the three-dimensional reflective surfaces are described as follows: ; In the formula, The nth reflective surface of the building's interior environment. Let be the unit normal vector of the reflecting surface. This is the intercept of the reflecting surface. The total number of valid paths involved in the positioning, including one direct path and The reflection path generated by the reflective surface.

[0025] Calculating the virtual anchor point coordinates of the 5G millimeter-wave base station location on the reflective surface of the indoor environment specifically includes: substituting the 5G millimeter-wave base station location into the plane equation of the indoor reflective surface, and combining this with the 5G millimeter-wave base station location to calculate the virtual anchor point coordinates, using the following formula: ; In the formula, The coordinates of the virtual anchor point on the nth reflective surface of the 5G millimeter-wave base station location within the indoor environment are given. The coordinates of the 5G millimeter-wave base station location. This represents the reflective surface corresponding to the location of the nth 5G millimeter-wave base station. Let be the unit normal vector of the reflecting surface. This is the intercept of the reflecting surface.

[0026] S103: Capture the user's indoor environment image using a calibrated RGB camera, perform target detection on the user's indoor environment image using the YOLOv7 model, obtain the target user image and its pixel representative position, and back-project it to the world coordinate system to obtain the camera's line of sight direction.

[0027] Optionally, target detection is performed on the user's indoor environment image using the YOLOv7 model to obtain the user image and the user's pixel representative position, which is then back-projected onto the world coordinate system to obtain the camera's gaze direction. Specifically, this includes: inputting the user's indoor environment image into the trained YOLOv7 model, and outputting the user image and the pixel positions of the vertices of its detection boxes. The user pixel position is determined by taking the center point of the pixel position at the vertex of the detection box. ;in, ; By back-projecting the user's pixel position, the camera's line of sight relative to the user in the world coordinate system is obtained, using the following formula: ; In the formula, The direction of the camera relative to the user's line of sight in the world coordinate system. For depth parameters, Let be the rotation matrix from the camera coordinate system to the world coordinate system. This is the camera intrinsic parameter matrix.

[0028] S104: Select candidate user locations that match the camera's line of sight from multiple pixel locations in the indoor environment image on the user end. Based on the coordinates of the candidate user locations and virtual anchor points, and combined with the indoor reflective surface plane equation, calculate multiple reflection points of the candidate user locations to obtain the user perception beam. With the minimum of the Cramer-Rao lower bound of the variance of the position parameter estimation of the reflection points as the optimization objective, optimize the transmission energy in the direction of the user perception beam using an energy-based adaptive allocation mechanism to obtain the optimized user perception beam.

[0029] Optionally, along the camera's line of sight, candidate user location points are selected. Based on the coordinates of the candidate user location points and virtual anchor points, and combined with the indoor reflective surface plane equation, multiple reflection points of the candidate user location points are calculated to obtain the user perception beam, specifically including: Along the camera's line of sight, candidate user locations are selected in the user image. Based on the coordinates of the candidate user locations and virtual anchor points, and combined with the indoor reflective surface plane equation, multiple reflection points of the candidate user locations are calculated to obtain the user-perceived beam, specifically including: The user candidate location points are selected along the direction of the camera's line of sight, using the following formula: ; In the formula, To capture all candidate user locations along the user's line of sight in the world coordinate system. Let be the position vector of the camera in the world coordinate system. The direction of the camera relative to the user's line of sight in the world coordinate system. For depth parameters; Combining the virtual anchor point coordinates, user candidate location points, and the indoor reflective surface plane equation, the reflection points of multiple user candidate location points are determined using the following formula: , ; In the formula, Let be the coordinates of the reflection point of the nth user candidate location on the nth reflecting surface. Let the coordinates of the nth virtual anchor point be... For the nth candidate user location, , Let be the scale factor along the direction from the virtual anchor point to the user candidate point, and its value corresponds to the position of the intersection point of the ray and the nth reflection plane; The reflection point of the 5G millimeter wave base station location and the user candidate location point, and the beam formed by the user candidate location point are determined as the user perception beam.

[0030] Optionally, the constraint for optimizing the user-perceived beam is that the transmitted energy allocated to each user-perceived beam direction is less than or equal to the preset total transmittable energy, as shown in the formula: ; In the formula, To give the first The transmitted energy allocated to the beam direction perceived by each user. It is a set of directions that includes both direct and primary reflection directions. The preset total energy, Here is the Fischer matrix of the reflection point position parameters. This represents taking the trace by inverting the Fischer matrix, i.e., summing the lower bounds of the variances of the position estimates. For position parameters, P={ ,..., } .

[0031] S105: Obtain channel state information based on the optimized user-aware beam; and use the ESPRIT algorithm to perform high-resolution inversion on the channel state information to obtain the multipath arrival angle and propagation delay.

[0032] Optionally, channel state information is obtained based on the optimized user-aware beam, specifically including: the user terminal completes cell identification and time-frequency synchronization by receiving the user-aware beam, and initiates a random access signal; after the 5G millimeter-wave base station completes the access response and link establishment, it obtains channel state information using the uplink reference signal or downlink reference signal, expressed as: ; In the formula, For the number of multipaths, For the first Complex gain of the user-perceived beam path, For the azimuth and elevation angles of the receiving end, For the azimuth and elevation angles of the transmitting end, For path delay, , These are the steering vectors for the receiving and transmitting arrays, respectively. For subcarrier frequency, This indicates performing a Hermitian transpose on the matrix.

[0033] Optionally, the ESPRIT algorithm is used to perform high-resolution inversion of the channel state information to obtain the multipath arrival angle and propagation delay, specifically including: ESPRIT is constructed by establishing shift-invariant relationships between adjacent subarrays or adjacent subcarriers, forming matrix pairs, as shown in the formula: ; In the formula, in the formula, Let be the rotation-invariant matrix of the angle. Let be a rotation-invariant matrix in the time delay dimension. For the signal subspace matrix, The selection matrix is ​​the array dimension. The selection matrix is ​​for the frequency domain dimension; Calculate separately and The characteristic values ​​are used to obtain the arrival angle and propagation delay of the user-perceived beam.

[0034] S106: Based on the arrival angle and propagation delay, determine the direct path and effective reflection path of the user sensing beam; using the intersection of the directional constraints formed by the direct path and the effective reflection path, determine the millimeter-wave positioning result at the user end.

[0035] The arrival angle and propagation delay are used to determine the direct path and effective reflection path of the user-sensing beam; the intersection of the directional constraints formed by the direct path and the effective reflection path is used to determine the millimeter-wave positioning result at the user end, specifically including: Utilizing reflective surfaces Corresponding virtual anchor point Construct an equivalent direct path; Calculation by The distance of propagation along the direction of the departure angle. The obtained point; the propagation distance The propagation delay is calculated using the following formula: ; In the formula, For the first The propagation distance of the user-perceived beam. At the speed of light, For the first Path delay estimation; If the point is not within the indoor space enclosed by the reflective surface, the path is identified as a pseudo-path and is eliminated. If the path is within the indoor space enclosed by the reflecting surface at that point, then the path is determined to be related to the indoor environment's reflecting surface. The associated primary reflection path. After eliminating false paths, the direct path and effective reflection path of the user-perceived beam are obtained; the intersection of the direct path and effective reflection path of the user-perceived beam is determined as the millimeter-wave positioning result, using the following formula:

[0036] ; In the formula, For user location estimation, It is the identity matrix. The number of valid paths involved in positioning, including one direct path and several reflective paths. For the user location variable to be estimated, The starting point of the user-perceived beam is represented by the base station location b when it is the direct path and the virtual anchor point when it is the reflected path. , This is the direction vector obtained by the ESPRIT algorithm.

[0037] Optionally, the method further includes: optimizing the millimeter-wave positioning results using Kalman filtering, including: Define the state vector as follows: ; in The three-dimensional position of the millimeter-wave positioning result. For the average velocity, the state transition equation is: ; in The sampling interval is... For process noise, This is the state transition matrix; Based on the aforementioned state transition equation, and utilizing both millimeter-wave point observations and visual projection point observations, the observation equation is as follows: ; In the formula, To observe the noise, For the observation vector, The observation matrix; Visual observation, on the other hand, involves observing the millimeter-wave observation point along the optical center of the camera and the direction vector. Defined rays Projection yields: ; Construct the observation equation: ; In the formula, For visual observation noise, For visual observation vectors, This is the visual observation matrix.

[0038] Specifically, this embodiment provides a more detailed implementation plan based on the above method, including: During the system deployment phase, the solution completes the calibration of base station locations and camera parameters, as well as 3D environment modeling, providing information such as camera intrinsic parameter matrices and reflector equations for subsequent steps. During deployment, the 5G millimeter-wave base station and RGB camera are placed within the user's active area, ensuring that the base station's coverage area is within the camera's field of view. Before operation, the base station location and camera extrinsic parameters must first be calibrated in a unified coordinate system. For ease of description, the camera location is set as the origin, and the base station location is denoted as... The camera's rotation matrix and translation vector are denoted as follows: and The camera's intrinsic parameter matrix is ​​denoted as

[0039] ; Furthermore, the main reflective surfaces in the environment, such as walls and floors, are described using the following equations. ; in Let be the unit normal vector of the reflection. The intercept of the reflection is given.

[0040] This scheme utilizes building design drawings and on-site measurement data to comprehensively model the boundaries and structure of the environment. Under a unified coordinate system, it constructs a set of three-dimensional reflective surfaces including walls, floors, and other major structures, thereby accurately depicting the main reflective characteristics of the indoor environment. Based on this, the coordinates of virtual anchor points are obtained using the mirror method.

[0041] ; in, For the base station about the reflective surface The mirror point represents the virtual launch position equivalent to the reflection path.

[0042] In the process of vision-assisted beam management, this scheme combines the direction obtained from image detection with existing model information, and utilizes multipath information to achieve perceptual beamforming, ensuring that the beam meets the design requirement of "optimal perception." First, the image frames acquired by the camera are processed using the YOLOv7 model. YOLO is an end-to-end object detection model; before deployment, it is trained using a labeled user dataset. The trained YOLO model takes the image frames from the camera as input and outputs the pixel positions of the user detection box vertices. The scheme uses the confidence level and the target class ID as the representative position of the user in the pixel plane, taking the center point of the bounding box as the value. ,satisfy:

[0043] ; The line of sight in the world coordinate system is obtained using the back projection method, where Let be the direction vector of the camera's view in the world coordinate system, expressed as follows: ; Sampling is performed along this line of sight within a certain depth range, using the following formula: ; User candidate location points are obtained by combining the reflected surfaces of the already modeled scene. For a single mirror reflection path, the reflection point is derived from the virtual anchor point: , ; The beam management process is periodic. The beam to be emitted in the next cycle can be reduced to a finite number of geometrically significant directions, which correspond to candidate direct points and multiple candidate reflection points, respectively.

[0044] In the beam mapping process, this method further introduces an energy-based adaptive allocation mechanism, which dynamically adjusts the transmission energy allocation ratio in each direction according to the contribution of different directions to positioning observability, thereby improving the theoretical sensing performance of non-direct path beams. The optimization objective is defined as the CRLB of the position parameter estimation variance, thereby minimizing position uncertainty and improving overall positioning accuracy under total energy constraints. In summary, the optimization problem can be expressed as:

[0045] ; Among them, set It is a set of directions that includes both direct and primary reflection directions. To give the first Distribute emission energy in each direction. Total energy that can be emitted , is the Fischer matrix of the position parameters.

[0046] The codebook obtained in this way is significantly smaller in size than the traditional design and has higher accuracy in angular coverage, thus resulting in higher time synchronization and accuracy in the beam management process.

[0047] In the channel information acquisition phase, the scheme extracts channel information based on the "optimal sensing" beam obtained in the previous step to accurately acquire path information. First, the user terminal completes cell identification and time-frequency synchronization by receiving a synchronization signal block (SSB) and initiates random access. After the base station completes the access response and link establishment, it acquires channel state information using the uplink reference signal (SRS) or downlink reference signal (CSI-RS). The millimeter-wave channel can be represented as a sparse multipath superposition model:

[0048] ; in, For the number of multipaths, For the first Complex gain of the path, For the azimuth and elevation angles of the receiving end , For the azimuth and elevation angles of the transmitting end, For path delay, , These are the steering vectors for the receiving and transmitting arrays, respectively.

[0049] When processing CSI, this method employs the ESPRIT algorithm for high-resolution inversion of angle and time delay. ESPRIT forms matrix pairs by constructing shift-invariant relationships between adjacent subarrays or adjacent subcarriers.

[0050] ; in, For the signal subspace matrix, The selection matrix is ​​the array dimension. This is the selection matrix for the frequency domain dimension. (From...) High-resolution eigenvalues ​​can recover two-dimensional angular information. ,Depend on The path delay can be obtained from the eigenvalues. .

[0051] In the path geometry association stage, this method combines existing model information and high-resolution angle information to achieve angle-path correlation and realize triangulation based on 5G millimeter wave. The method first uses the angle and delay estimation results obtained from the ESPRIT algorithm to distinguish between direct and reflected paths. Specifically, based on the estimated propagation delay... The propagation distance can be calculated. Combined with angle parameters Determine the ray originating from the base station. If candidate points are calculated from this angle and distance... If the connection to the base station does not intersect with any reflecting surface along its propagation path, the path is considered a direct path; conversely, if the connection intersects with any reflecting surface... If an intersection point exists, it is further determined whether it satisfies the mirror geometry of the reflecting surface. The specific implementation path is as follows: using the reflecting surface... Corresponding virtual anchor point Construct an equivalent direct path and calculate the path from the given path. Starting point, distance propagated along the estimated angular direction The obtained point. If the point is not within the user's movable area, the path is considered inconsistent with the scene geometry, is a pseudopath, and is discarded; if it is within a reasonable area, the path is classified as being aligned with the reflective surface. The associated primary reflection path.

[0052] After eliminating false paths, the remaining paths are divided into two categories: direct paths and effective reflection paths. Since the starting point of the reflection path can be equivalent to a virtual anchor point... Therefore, the direct path corresponds to the geometric constraint between the base station and the user, while the reflected path corresponds to the geometric constraint between the virtual anchor point and the user. The final positioning stage no longer relies on absolute time delay, but instead utilizes the angular information between these anchor points and the user to determine the user's position through the intersection of multiple angles. In other words, the direct path provides the directional constraint from the base station to the user, the reflected path provides the directional constraint from the virtual anchor point to the user, and the intersection of multiple constraints is the user's millimeter-wave positioning result, i.e., solving...

[0053] ; in, Indicates the starting point of the constraint (the base station location when the direct path is used). b The reflection path is a virtual anchor point. ), This is the direction vector obtained by the ESPRIT algorithm.

[0054] In the multi-source fusion filtering and dynamic tracking stage, this scheme utilizes Kalman filtering to further optimize the millimeter-wave positioning results obtained in the previous step. The state vector is defined as:

[0055] ; in For three-dimensional position, The average velocity is given. The state transition equation is:

[0056] ; in The sampling interval is... This is process noise.

[0057] In the observation equation modeling, the Kalman filter utilizes both millimeter-wave point observations and visual projection point observations. The millimeter-wave module outputs an absolute position measurement. Its observation equation is

[0058] ; in This is noise in millimeter-wave observations. Visual observations, on the other hand, measure the noise by moving the millimeter-wave observation point along the camera's optical center and the direction vector. Defined rays Projection yields:

[0059] ; And construct the observation equation: ; in To address visual observation noise, by simultaneously fusing millimeter-wave point observations and visual projection point observations during the filtering update, the system can balance millimeter-wave positioning with visual geometric constraints, achieving more stable and accurate 3D positioning.

[0060] In summary, this method completes the spatial calibration and modeling of base stations, cameras, and 3D reflective surfaces during system deployment. It leverages visual orientation priors to significantly compress the search space for millimeter-wave beam training. The ESPRIT algorithm then performs high-resolution parameter inversion of the CSI, and combined with geometric consistency and multi-source fusion filtering, ultimately achieving high-precision and stable positioning results. Compared to existing technologies, this method offers significant advantages in terms of speed, positioning accuracy, and system complexity.

[0061] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this invention.

Claims

1. A three-dimensional indoor positioning method integrating vision and 5G sensing, characterized in that, include: The location of the 5G millimeter-wave base station and the internal and external parameters of the RGB camera were determined in a unified coordinate system; A three-dimensional model of the reflective surfaces of the building's interior environment is created to obtain the plane equation of the interior reflective surfaces; And calculate the virtual anchor point coordinates of the 5G millimeter wave base station location on the reflective surface of the building's indoor environment; The system captures images of the user's indoor environment using a calibrated RGB camera, performs target detection on the user's indoor environment images using the YOLOv7 model, obtains the target user image and its pixel representative position, and back-projects it onto the world coordinate system to obtain the camera's line of sight direction. Candidate user location points are selected in the user image along the camera's line of sight. Based on the coordinates of the candidate user location points and virtual anchor points, and combined with the indoor reflective surface plane equation, multiple reflection points of the candidate user location points are calculated to obtain the user perception beam. With the minimum of the Cramer-Rao lower bound of the variance of the position parameter estimation of the reflection points as the optimization objective, the transmission energy in the direction of the user perception beam is optimized using an energy-based adaptive allocation mechanism to obtain the optimized user perception beam. Channel state information is obtained based on the optimized user-perceived beam; The ESPRIT algorithm is then used to perform high-resolution inversion of the channel state information to obtain the multipath arrival angle and propagation delay. Based on the angle of arrival and propagation delay, the direct path and effective reflection path of the user's sensing beam are determined; using the intersection of the directional constraints formed by the direct path and the effective reflection path, the user's millimeter-wave positioning result is determined.

2. The three-dimensional indoor positioning method integrating vision and 5G sensing as described in claim 1, characterized in that, The reflective surfaces of the building's interior environment include: three-dimensional reflective surfaces of walls, floors, and other main structures; the equations of these three-dimensional reflective surfaces are described as follows: ; In the formula, The nth reflective surface of the building's interior environment. Let be the unit normal vector of the reflecting surface. This is the intercept of the reflecting surface. The total number of valid paths involved in the positioning, including one direct path and The reflection path generated by the reflective surface.

3. The three-dimensional indoor positioning method integrating vision and 5G sensing as described in claim 1, characterized in that, The calculation of the virtual anchor point coordinates of the 5G millimeter-wave base station location on the reflective surface of the indoor environment specifically includes: Substituting the location of the 5G millimeter-wave base station into the equation of the indoor reflective surface, and combining this with the location of the 5G millimeter-wave base station, the coordinates of the virtual anchor point are calculated using the following formula: ; In the formula, The coordinates of the virtual anchor point on the nth reflective surface of the 5G millimeter-wave base station location within the indoor environment are given. The coordinates of the 5G millimeter-wave base station location. This represents the reflective surface corresponding to the location of the nth 5G millimeter-wave base station. Let be the unit normal vector of the reflecting surface. This is the intercept of the reflecting surface.

4. The three-dimensional indoor positioning method integrating vision and 5G sensing as described in claim 1, characterized in that, The process of performing target detection on the user's indoor environment image using the YOLOv7 model to obtain the user image and the user's pixel representative position, and then back-projecting it to the world coordinate system to obtain the camera's line of sight, specifically includes: The user's indoor environment image is input into the trained YOLOv7 model, which outputs the user image and the pixel positions of the vertices of the detection boxes. ; The user pixel position is determined by the center point of the pixel position of the detection box vertex. ;in, ; By back-projecting the user's pixel position, the camera's line of sight relative to the user in the world coordinate system is obtained, using the following formula: ; In the formula, The direction of the camera relative to the user's line of sight in the world coordinate system. For depth parameters, Let be the rotation matrix from the camera coordinate system to the world coordinate system. This is the camera intrinsic parameter matrix.

5. The three-dimensional indoor positioning method integrating vision and 5G sensing as described in claim 1, characterized in that, The process of selecting candidate user locations in the user image along the camera's line of sight, calculating multiple reflection points of the candidate user locations based on the coordinates of the candidate user locations and virtual anchor points, and combining this with the indoor reflective surface plane equation, to obtain the user-perceived beam, specifically includes: The user candidate location points are selected along the direction of the camera's line of sight, using the following formula: ; In the formula, To capture all candidate user locations along the user's line of sight in the world coordinate system. Let be the position vector of the camera in the world coordinate system. The direction of the camera relative to the user's line of sight in the world coordinate system. For depth parameters; Combining the virtual anchor point coordinates, user candidate location points, and the indoor reflective surface plane equation, the reflection points of multiple user candidate location points are determined using the following formula: , ; In the formula, Let be the coordinates of the reflection point of the nth user candidate location on the nth reflecting surface. Let the coordinates of the nth virtual anchor point be... For the nth candidate user location, , Let be the scale factor along the direction from the virtual anchor point to the user candidate point, and its value corresponds to the position of the intersection point of the ray and the nth reflection plane; The reflection point of the 5G millimeter wave base station location and the user candidate location point, and the beam formed by the user candidate location point are determined as the user perception beam.

6. The three-dimensional indoor positioning method integrating vision and 5G sensing as described in claim 1, characterized in that, The user-sensing beam is optimized using the trace of the inverse of the position estimation Fischer matrix as the objective function. The constraint is that the transmit energy allocated to each user-sensing beam direction is less than or equal to the preset total transmittable energy, as shown in the formula: ; In the formula, To give the first The transmitted energy allocated to the beam direction perceived by each user. It is a set of directions that includes both direct and primary reflection directions. The preset total energy, Here is the Fischer matrix of the reflection point position parameters. This represents taking the trace by inverting the Fischer matrix, i.e., summing the lower bounds of the variances of the position estimates. For position parameters, P={ ,..., } .

7. The three-dimensional indoor positioning method based on the fusion of vision and 5G sensing as described in claim 1, characterized in that, The step of obtaining channel state information based on the optimized user-perceived beam specifically includes: The user terminal completes cell identification and time-frequency synchronization by receiving the user-perceived beam, and initiates a random access signal; After completing the access response and link establishment, the 5G millimeter-wave base station obtains channel state information using uplink or downlink reference signals, expressed as: ; In the formula, For the number of multipaths, For the first Complex gain of the user-perceived beam path, For the azimuth and elevation angles of the receiving end, For the azimuth and elevation angles of the transmitting end, For path delay, , These are the steering vectors for the receiving and transmitting arrays, respectively. For subcarrier frequency, This indicates performing a Hermitian transpose on the matrix.

8. The three-dimensional indoor positioning method integrating vision and 5G sensing as described in claim 1, characterized in that, The ESPRIT algorithm is used to perform high-resolution inversion of the channel state information to obtain the multipath arrival angle and propagation delay, specifically including: ESPRIT is constructed by establishing shift-invariant relationships between adjacent subarrays or adjacent subcarriers, forming matrix pairs, as shown in the formula: ; In the formula, Let be the rotation-invariant matrix of the angle. Let be a rotation-invariant matrix in the time delay dimension. For the signal subspace matrix, The selection matrix is ​​the array dimension. The selection matrix is ​​for the frequency domain dimension; Calculate separately and The characteristic values ​​are used to obtain the arrival angle and propagation delay of the user-perceived beam.

9. The three-dimensional indoor positioning method integrating vision and 5G sensing as described in claim 1, characterized in that, Based on the angle of arrival and propagation delay, the direct path and effective reflection path of the user-sensing beam are determined; using the intersection of the directional constraints formed by the direct path and the effective reflection path, the millimeter-wave positioning result at the user end is determined, specifically including: Utilizing reflective surfaces Corresponding virtual anchor point Construct an equivalent direct path; Calculation by The distance of propagation along the direction of the departure angle. The obtained point; the propagation distance The path delay is calculated using the following formula: ; In the formula, For the first The propagation distance of the user-perceived beam. At the speed of light, For the first Path delay estimation; If the point is not within the indoor space enclosed by the reflective surface, the path is identified as a pseudo-path and is eliminated. If the path is within the indoor space enclosed by the reflecting surface at that point, then the path is determined to be related to the indoor environment's reflecting surface. The associated primary reflection path; After eliminating false paths, the direct path and effective reflection path of the user-perceived beam are obtained; The intersection of the direct path and the effective reflection path of the user-perceived beam is determined as the millimeter-wave positioning result, using the following formula: ; In the formula, For user location estimation, It is the identity matrix. The number of valid paths involved in positioning, including one direct path and several reflective paths. For the user location variable to be estimated, The starting point of the user-perceived beam is represented by the base station location b when it is the direct path and the virtual anchor point when it is the reflected path. , This is the direction vector obtained by the ESPRIT algorithm.

10. The three-dimensional indoor positioning method integrating vision and 5G sensing as described in claim 1, characterized in that, Also includes: The millimeter-wave positioning results are optimized using Kalman filtering, including: Define the state vector as follows: ; in, The three-dimensional position of the millimeter-wave positioning result. Given the average velocity, the state transition equation is: ; in The sampling interval is... For process noise, This is the state transition matrix; Based on the aforementioned state transition equation, and utilizing both millimeter-wave point observations and visual projection point observations, the observation equation is as follows: ; In the formula, To observe the noise, For the observation vector, The observation matrix; Visual observation, on the other hand, involves observing the millimeter-wave observation point along the optical center of the camera and the direction vector. Defined rays Projection yields: ; Construct the observation equation: ; in, For visual observation noise, For visual observation vectors, This is the visual observation matrix.