Multi-mode indoor positioning method fusing radio frequency and machine vision

By integrating a multimodal indoor positioning method with radio frequency and machine vision, and combining three levels of data fusion and hierarchical decision-making algorithms, the accuracy and adaptability problems of indoor positioning in complex environments in existing technologies are solved, and high-precision real-time positioning is achieved.

CN120593748AActive Publication Date: 2025-09-05YUNTU DATA TECH (ZHENGZHOU) CO LTD
View PDF 14 Cites 0 Cited by

Patent Information

Application Number
CN202510604252.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-09-05
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

Existing indoor positioning technologies have shortcomings in terms of accuracy, real-time performance, deployment cost and adaptability, especially in complex indoor environments, where it is difficult to meet the needs of high-precision real-time positioning.

Method used

A multimodal indoor positioning method that integrates radio frequency and machine vision is adopted. Through monocular ranging combined with three levels of data fusion (original level, feature level, decision level), the advantages of wireless signals and visual features are combined, and a hierarchical decision algorithm is used to adapt to multi-scenario requirements to ensure the adaptability of the positioning method in different environments.

Benefits of technology

It effectively makes up for the shortcomings of single-mode positioning in complex environments, achieves high-precision real-time positioning, adapts to the needs of multiple scenarios, and improves the adaptability and accuracy of positioning methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120593748A_ABST
    Figure CN120593748A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode indoor positioning method fusing radio frequency and machine vision, relates to the technical field of indoor positioning, and solves the problem that the existing indoor positioning technology is difficult to meet high-precision real-time positioning in a complex indoor environment. The multi-mode indoor positioning method comprises the following steps: calibrating and starting a camera device, establishing an internal and external coordinate mapping matrix in combination with the height of the camera device, and simultaneously acquiring and processing indoor wireless signals to obtain wireless signal characteristics of all equipment; detecting and segmenting a target area in a video stream of the camera device, measuring and calculating the average depth of the target area in combination with the depth estimation model, and obtaining visual positioning coordinates of a target person according to the coordinate mapping matrix; measuring and calculating wireless positioning coordinates of the target equipment according to the target area in combination with wireless signal characteristics; according to the data level, a hierarchical decision algorithm is adopted, the visual positioning coordinates and the wireless positioning coordinates are fused hierarchically, fused positioning coordinates of the target person are obtained, and then indoor positioning of the person is completed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of indoor positioning technology, and in particular to a multimodal indoor positioning method integrating radio frequency and machine vision. Background Art

[0002] With the rapid development of the Internet of Things (IoT) and intelligent technologies, indoor positioning technology has broad application prospects in many fields, including intelligent building management, security monitoring, logistics management, and people tracking. However, existing indoor positioning technologies still face many challenges in terms of accuracy, real-time performance, deployment costs, and adaptability. New solutions are urgently needed to meet this growing demand.

[0003] Using methods such as mobile phone positioning and tablet positioning, the target's location is determined by measuring wireless signal strength, arrival time, or phase difference. Advantages include mature technology, low equipment cost, and the ability to achieve positioning over a wide range. However, its disadvantages are that positioning accuracy is significantly affected by environmental factors such as signal obstruction and multipath effects, and high-precision positioning is difficult to achieve in complex indoor environments. Sensors such as accelerometers and gyroscopes are used to measure the target's trajectory. This method does not rely on external signals and has good anti-interference capabilities. However, due to sensor noise and accumulated errors, positioning accuracy degrades significantly after long-term use, making it difficult to meet high-precision positioning requirements. Camera-captured images are used for target recognition and position estimation. Visual positioning methods offer high positioning accuracy and rich environmental information, but they rely on good lighting conditions and camera coverage. Furthermore, the computational complexity of image processing and analysis is high, making real-time performance and deployment costs major challenges.

[0004] Patent No. CN202210381782.2 discloses a CSI indoor positioning method based on multimodal GAN. The method aims to achieve low-cost, high-precision indoor positioning, and the steps are: obtaining CSI data, extracting three types of data features from it: average amplitude, phase difference, and CIR amplitude distribution center moment; fusing multidimensional data into an image through the KCCA algorithm; using the GAN network to perform image expansion and training on the data set consisting of images and category labels; and realizing position estimation through a multi-image positioning algorithm based on spectral clustering. The feature of the above invention is that indoor positioning is achieved by constructing a CSI multidimensional image with low acquisition cost and high fingerprint discrimination and using the GAN network for image expansion and training, which improves the stability of the positioning performance, effectively reduces the positioning error caused by noise and information loss, and can meet the high-precision and low-cost requirements in indoor positioning application scenarios.

[0005] Patent No. CN202010486081.6 discloses a multimodal indoor positioning and navigation method and system, wherein the method includes the following steps: collecting location information of indoor structures, drawing indoor maps, collecting indoor real-scene image information, correspondingly integrating reference image information with reference location information, building an indoor location information database, establishing a central area, and storing it on a server; when the mobile terminal is indoors, turning on the camera module, obtaining real-time image information, and feeding it back to the server; the server matches the real-time image information with the reference image in the location information database, outputs the matched reference location information, determines whether the matched reference location information is in the central area, and feeds back the judgment result to the mobile terminal; the mobile terminal selects a positioning scheme based on the judgment result, and if the judgment result is yes, enables the first mode for positioning and navigation, and if the judgment result is no, enables the second mode for positioning and navigation. The above invention can achieve precise indoor positioning and navigation.

[0006] Although the above patents can accomplish indoor positioning, existing indoor positioning technologies still have significant deficiencies in terms of accuracy, real-time performance, deployment costs, and adaptability. Especially in complex indoor environments, it is difficult to achieve high-precision real-time positioning. Summary of the Invention

[0007] The purpose of the present invention is to provide a multimodal indoor positioning method that integrates radio frequency and machine vision. It can combine three levels of data fusion (original level, feature level, and decision level) through a monocular ranging method, fully combining the advantages of wireless signals and visual features, and effectively compensating for the shortcomings of single-modal positioning in complex environments. Through a hierarchical decision algorithm, it adapts to the needs of multiple scenarios and ensures the adaptability of the positioning method in different environments.

[0008] The present invention utilizes the following technical solutions:

[0009] A multimodal indoor positioning method integrating radio frequency and machine vision comprises the following steps:

[0010] S1: Calibrate and start the camera device, and establish an internal and external coordinate mapping matrix based on the camera device height. At the same time, collect and process the wireless signals in the room to obtain the wireless signal characteristics of all devices;

[0011] S2: Detect and segment the target area in the video stream of the camera device, combine the depth estimation model to measure the average depth of the target area, and obtain the visual positioning coordinates of the target person according to the coordinate mapping matrix;

[0012] S3: Calculate the wireless positioning coordinates of the target device based on the target area and wireless signal characteristics;

[0013] S4: Based on the data level, a hierarchical decision algorithm is used to hierarchically fuse the visual positioning coordinates and wireless positioning coordinates to obtain the fused positioning coordinates of the target person, thereby completing the indoor positioning of the person.

[0014] Preferably, step S1 includes the following steps:

[0015] S11: Start the camera device to take several calibration plate images;

[0016] S12: using mathematical tools to calculate the intrinsic parameter matrix of the camera device according to the calibration plate image;

[0017] S13: Obtaining an extrinsic parameter matrix of the camera device according to the camera device coordinate system and the world coordinate system;

[0018] S14: Establishing an internal and external coordinate mapping matrix based on the camera device height by combining the internal parameter matrix and the external parameter matrix;

[0019] S15: Capture the radio frequencies emitted by all devices in the room through a wireless data collector to obtain several wireless signals;

[0020] S16: aligning multiple wireless signals in the time domain through a phase-locked loop to obtain a wireless synchronization signal;

[0021] S17: Denoising and clustering the wireless synchronization signal according to the Laida criterion combined with the density clustering algorithm to obtain a noise-free synchronization signal;

[0022] S18: performing data enhancement on the noise-free synchronization signal by using Gaussian white noise to obtain a multi-mode synchronization signal;

[0023] S19: Extract features from the time domain, frequency domain, and spatial domain of the multi-mode synchronous signal to obtain wireless signal features of all devices.

[0024] Preferably, step S2 includes the following steps:

[0025] S201: receiving a video stream captured by a camera device with zero frame loss through a double buffer queue, and converting the video stream into a plurality of frame images using a frame-by-frame extraction algorithm;

[0026] S202: Eliminate color difference of each frame image by using an adaptive white balance algorithm to obtain a color-free image;

[0027] S203: Using a multi-scale Laplacian pyramid algorithm to compress and fuse the color-free images to obtain a fused image;

[0028] S204: The fused image is detected by combining the detection branch of the hybrid perception model with the human body key point heat map to obtain the posture confidence, and then compared with the preset confidence threshold for judgment:

[0029] If the posture confidence is greater than or equal to the confidence threshold, the target person in the current frame fusion image is marked using the human body bounding box. At the same time, the human body bounding box is combined with the body texture attributes to decompose into several part element points and match them with the body parts in all frame fusion images:

[0030] If any part element point has a matching body part, the body part is marked and the part bounding box is obtained until there is no body part of the target person in the fused image;

[0031] If any part element point has no matching body part, it means that the target person has left the current video stream;

[0032] If the posture confidence is less than the confidence threshold, it means that the suspicious person in the current fused image has a similar posture to the target person. At the same time, the fused image is marked to obtain a suspicious bounding box.

[0033] S205: Segmenting the human body bounding box, part bounding box, and suspicious bounding box using the segmentation branch of the hybrid perception model combined with the occlusion coefficient matrix to obtain instance-level masks;

[0034] S206: Generate a confidence assessment depth map based on the fused images of all frames using a depth estimation model combined with an attention gating mechanism;

[0035] S207: Segment the target person in each fused image frame according to the instance-level mask to obtain a target region;

[0036] S208: Calculating the depth of the target area according to the confidence evaluation depth map to obtain an average depth of the target area;

[0037] S209: Establishing a world coordinate system with the vertical projection of the camera device on the ground as the origin, and obtaining the height of the camera device and the height of the target person;

[0038] S210: Calculating the height of the camera device and the height of the target person according to the coordinate mapping matrix and the average depth of the target area to obtain the visual positioning coordinates of the target person.

[0039] Preferably, step S3 includes the following steps:

[0040] S31: Based on the wireless signal characteristics of all devices, a dynamic Doppler frequency shift spectrum and phase difference matrix are constructed, and the spatiotemporal feature tensor is constructed in combination with the timestamp;

[0041] S32: Using an improved particle filter algorithm to perform interference compensation on the spatiotemporal feature tensor to obtain a spatiotemporal compensated feature tensor;

[0042] S33: Based on the dynamic Doppler frequency shift spectrum and the visual positioning coordinates of the target person, the synchronized device of the target person is selected from all devices as the target device;

[0043] S34: performing data fusion on the spatiotemporal compensation feature tensor based on the signal quality index of the wireless signal characteristics of the target device, and calculating the positioning fusion data of the target device;

[0044] S35: Use a deep metric learning algorithm combined with a hybrid loss function and a multi-lateral positioning function to convert the target device's positioning fusion data into wireless positioning data;

[0045] S36: Using the heterogeneous coordinate system conversion matrix combined with the world coordinate system, the wireless positioning data is converted to obtain the wireless positioning coordinates of the target device.

[0046] Preferably, step S4 includes the following steps:

[0047] Based on the relationship between the target device and the target person, the visual positioning coordinates and the wireless positioning coordinates are hierarchically fused:

[0048] If the target device is bound to the target person, the wireless positioning coordinates and visual positioning coordinates are fused at the original level and feature level according to the positioning accuracy;

[0049] The process of primitive level fusion is:

[0050] The wireless signal characteristics of the target device are calculated using the least squares method to obtain the original wireless distance between the target device and the wireless beacon;

[0051] Calculate the original visual distance between the target person and the wireless beacon based on the coordinate mapping matrix and the average depth of the target area;

[0052] According to the positioning accuracy, the weights corresponding to the wireless original distance and the visual original distance are estimated respectively to obtain the wireless dynamic weight and the visual dynamic weight;

[0053] Calculate the wireless dynamic weight, visual dynamic weight, wireless positioning coordinates and visual positioning coordinates to obtain the original level fusion positioning coordinates of the target person;

[0054] The process of feature-level fusion is as follows:

[0055] The weighted centroid algorithm is used to analyze the wireless signal characteristics of the target device in combination with the wireless positioning coordinates to obtain wireless depth positioning data;

[0056] The feature expansion algorithm is used to expand the visual positioning coordinates of the target person to obtain visual depth positioning data;

[0057] The wireless depth positioning data and the visual depth positioning data are weightedly fused to obtain the feature-level fused positioning coordinates of the target person.

[0058] If the target device is not bound to the target person, the wireless positioning coordinates and visual positioning coordinates are fused at the decision level according to the scene environment;

[0059] The process of decision-level fusion is as follows:

[0060] Calculate the absolute positioning difference between the wireless positioning coordinates and the visual positioning coordinates, and compare it with the positioning threshold:

[0061] If the absolute positioning difference is less than the positioning threshold, the wireless positioning coordinates and the visual positioning coordinates are weightedly fused to obtain the first-level decision-making fusion positioning coordinates of the target person, and the target device is determined to belong to the target person;

[0062] If the absolute positioning difference is greater than or equal to the positioning threshold, it is determined that the target device does not belong to the target person;

[0063] If all N absolute positioning differences are less than the positioning threshold, all absolute positioning differences are sorted from small to large, and the visual positioning coordinates corresponding to the smallest absolute positioning difference are selected for weighted fusion with the wireless positioning coordinates to obtain the first-level decision-making fusion positioning coordinates of the target person, and the target device is determined to belong to the target person;

[0064] Based on the human body bounding box and part bounding box in all frame fusion images, combined with the timestamp, a person trajectory sequence is generated; based on the dynamic Doppler frequency shift spectrum combined with the wireless positioning coordinates of the target device and the timestamp, a device trajectory sequence is generated; the trajectory matching degree between the person trajectory sequence and the device trajectory sequence is calculated and compared with the matching threshold:

[0065] If the trajectory matching degree is greater than the matching threshold, the wireless positioning coordinates and the visual positioning coordinates are weightedly fused to obtain the second-level decision-making fusion positioning coordinates of the target person, and the target device is determined to belong to the target person;

[0066] If the trajectory matching degree is less than or equal to the matching threshold, it is determined that the target device does not belong to the target person;

[0067] The distance between the target person and the camera device is measured using the confidence assessment depth map to obtain the person distance; the distance between the target device and the camera device is measured based on the geometric relationship between the wireless positioning coordinates and the camera device to obtain the device distance;

[0068] Calculate the absolute distance difference between the person's distance and the device's distance, and compare it with the distance threshold range based on physical constraints:

[0069] If the absolute distance difference is within the distance threshold range, it is determined that the wireless positioning coordinates and the visual positioning coordinates are weightedly fused to obtain the third-level decision-making fused positioning coordinates of the target person, and the target device is determined to belong to the target person;

[0070] If the absolute distance difference is not within the distance threshold range, it is determined that the target device does not belong to the target person;

[0071] Define the binding status of each device location and the target person location;

[0072] Calculate the posterior probability of the binding state based on the wireless positioning coordinates and visual positioning coordinates combined with the dynamic Doppler frequency shift spectrum and phase difference matrix;

[0073] The maximum a posteriori probability is combined with the hierarchical decision algorithm to determine the binding status between the target device and the target person;

[0074] According to the trajectory sequence of wireless positioning coordinates and visual positioning coordinates, the binding status is dynamically updated to obtain the fourth decision-level fusion positioning coordinates of the target person, and it is determined that the target device belongs to the target person, thereby completing the indoor positioning of the person.

[0075] The present invention combines a monocular ranging method with three levels of data fusion (original level, feature level, and decision level), fully integrating the advantages of wireless signals and visual features, effectively compensating for the shortcomings of single-modality positioning in complex environments; through a hierarchical decision algorithm, it adapts to the needs of multiple scenarios and ensures the adaptability of the positioning method in different environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technical personnel in this field, other drawings can be obtained based on the provided drawings without any creative work.

[0077] Figure 1 This is a principle block diagram of the multimodal indoor positioning method;

[0078] Figure 2 Schematic diagram of coordinate transformation. DETAILED DESCRIPTION

[0079] The present invention is described in detail below with reference to the accompanying drawings and embodiments:

[0080] like Figure 1-Figure 2 As shown, the multimodal indoor positioning method integrating radio frequency and machine vision described in the present invention includes the following steps in sequence:

[0081] S1: Calibrate and start the camera device, and establish an internal and external coordinate mapping matrix based on the camera device height. At the same time, collect and process the wireless signals in the room to obtain the wireless signal characteristics of all devices;

[0082] In this embodiment, the signal characteristics include signal strength (RSSI), time of flight (ToF), angle of arrival (AoA), etc.

[0083] S2: Detect and segment the target area in the video stream of the camera device, combine the depth estimation model to measure the average depth of the target area, and obtain the visual positioning coordinates of the target person according to the coordinate mapping matrix;

[0084] S3: Calculate the wireless positioning coordinates of the target device based on the target area and wireless signal characteristics;

[0085] S4: Based on the data level, a hierarchical decision algorithm is used to hierarchically fuse the visual positioning coordinates and wireless positioning coordinates to obtain the fused positioning coordinates of the target person, thereby completing the indoor positioning of the person.

[0086] In this embodiment, the data level includes the original level, the feature level and the decision level.

[0087] Raw data refers to the initial data collected directly from sensors, devices, or systems without any processing. This data usually exists in the most basic format and may contain noise, redundant information, or invalid values.

[0088] The feature level is information extracted from raw data that can characterize the key attributes of the data. These features are generated through mathematical transformations or domain knowledge to simplify the data and enhance its expressiveness.

[0089] The decision-making level is the final operational instruction or conclusion generated by algorithms or rules based on feature-level data. Its goal is to convert features into specific actions or judgments.

[0090] In the present invention, step S1 includes the following steps:

[0091] S11: Start the camera device to take several calibration plate images;

[0092] S12: using mathematical tools to calculate the intrinsic parameter matrix of the camera device according to the calibration plate image;

[0093] S13: Obtaining an extrinsic parameter matrix of the camera device according to the camera device coordinate system and the world coordinate system;

[0094] S14: Establishing an internal and external coordinate mapping matrix by combining the internal parameter matrix and the external parameter matrix according to the height of the camera device.

[0095] In this embodiment, the intrinsic parameter matrix is ​​the core parameter that associates the image coordinate system with the camera coordinate system; the extrinsic parameter matrix establishes the mapping relationship between the camera coordinate system and the world coordinate system;

[0096] Use a camera to take multiple images of the calibration plate, use Open Vision and other visual tools to calibrate, and calculate the intrinsic parameter matrix K of the camera, including the focal length (f x ,f y ), optical center coordinates (C x ,C y ). The internal parameter matrix is ​​as follows:

[0097]

[0098] where f x 、f y is the focal length of the camera in the horizontal and vertical directions, in pixels, C x 、C y is the projection position of the optical axis of the camera device on the image plane, located at the center of the image.

[0099] S15: Capture the radio frequencies emitted by all devices in the room through a wireless data collector to obtain several wireless signals;

[0100] S16: aligning multiple wireless signals in the time domain through a phase-locked loop to obtain a wireless synchronization signal;

[0101] S17: Denoising and clustering the wireless synchronization signal according to the Laida criterion combined with the density clustering algorithm to obtain a noise-free synchronization signal;

[0102] S18: performing data enhancement on the noise-free synchronization signal by using Gaussian white noise to obtain a multi-mode synchronization signal;

[0103] S19: Extract features from the time domain, frequency domain, and spatial domain of the multi-mode synchronous signal to obtain wireless signal features of all devices.

[0104] In the present invention, step S2 includes the following steps:

[0105] S201: receiving a video stream captured by a camera device with zero frame loss through a double buffer queue, and converting the video stream into a plurality of frame images using a frame-by-frame extraction algorithm;

[0106] S202: Eliminate color difference of each frame image by using an adaptive white balance algorithm to obtain a color-free image;

[0107] S203: Using a multi-scale Laplacian pyramid algorithm to compress and fuse the color-free images to obtain a fused image;

[0108] S204: The fused image is detected by combining the detection branch of the hybrid perception model with the human body key point heat map to obtain the posture confidence, and then compared with the preset confidence threshold for judgment:

[0109] In this embodiment, a human keypoint heatmap is a probability distribution map used in computer vision to represent the locations of key points in human poses (such as joints and facial features). Its core principle is to use the intensity value of each pixel in a two-dimensional matrix (usually the same size as the input image) to reflect the probability that the location belongs to a keypoint. The generation and interpretation of heatmaps is one of the core steps in the task of Human Pose Estimation.

[0110] Human body key points refer to the significant anatomical landmarks of human posture, such as the head, shoulders, elbows, wrists, hips, knees, ankles, etc., and usually 17 to 25 key points are defined (such as the COCO dataset standard is 17).

[0111] The heat map generates a two-dimensional probability map for each keypoint, where each pixel value represents the probability that the location belongs to the keypoint (ranging from 0 to 1). Regions with higher probabilities (usually represented by warm colors, such as red) are more likely to be the true location of the keypoint.

[0112] If the posture confidence is greater than or equal to the confidence threshold, the target person in the current frame fusion image is marked using the human body bounding box. At the same time, the human body bounding box is combined with the body texture attributes to decompose into several part element points and match them with the body parts in all frame fusion images:

[0113] In this embodiment, the volume texture attribute is a type of three-dimensional texture, also known as a three-dimensional texture. It is a logical extension of the traditional two-dimensional texture. A two-dimensional texture is usually a simple bitmap that provides color values ​​for the surface points of a three-dimensional model. A volume texture can be viewed as a three-dimensional data structure composed of multiple two-dimensional textures stacked together to describe data in three-dimensional space. Volume textures are accessed using three-dimensional texture coordinates, typically using (u, v, w) coordinates to specify a location in the texture, where u and v represent plane coordinates and w represents a depth coordinate.

[0114] If any part element point has a matching body part, the body part is marked and the part bounding box is obtained until there is no body part of the target person in the fused image;

[0115] If any part element point has no matching body part, it means that the target person has left the current video stream;

[0116] If the posture confidence is less than the confidence threshold, it means that the suspicious person in the current fused image has a similar posture to the target person. At the same time, the fused image is marked to obtain a suspicious bounding box.

[0117] S205: Segment the human body bounding box, part bounding box, and suspicious bounding box through the segmentation branch of the hybrid perception model combined with the occlusion coefficient matrix to obtain an instance-level mask.

[0118] S206: Generate a confidence assessment depth map based on the fused images of all frames using a depth estimation model combined with an attention gating mechanism;

[0119] S207: Segment the target person in each fused image frame according to the instance-level mask to obtain a target region;

[0120] S208: Calculating the depth of the target area according to the confidence evaluation depth map to obtain an average depth of the target area;

[0121] S209: Establishing a world coordinate system with the vertical projection of the camera device on the ground as the origin, and obtaining the height of the camera device and the height of the target person;

[0122] S210: According to the coordinate mapping matrix and the average depth of the target area, the camera height and the height of the target person are calculated to obtain the visual positioning coordinates (x P ,y P ).

[0123] In this embodiment, the world coordinate system is established with the vertical projection of the camera device on the ground as the origin, and the height of the camera device is h C The intersection of the optical axis of the camera device and the ground is C, the position of the person from the ground to the top of the head is P0P1, and the height of the person is h P , the midpoint is P, then the height of P is h P / 2. The distance from the camera to point P is the depth d estimated by the algorithm. The horizontal distance d between the person and the origin P for:

[0124]

[0125] Point P can be approximately regarded as the center point of the character segmentation area on the image, and its horizontal pixel distance d from the optical center x It can be used to calculate the horizontal angle θ between the person and the optical axis:

[0126]

[0127] where f x is the horizontal focal length in the intrinsic parameter matrix.

[0128] From this, we can get the coordinates of the character in the world coordinate system:

[0129]

[0130] The height of the characters in the above process can be approximately replaced by the average height of the population.

[0131] In the present invention, step S3 includes the following steps:

[0132] S31: Based on the wireless signal characteristics of all devices, a dynamic Doppler frequency shift spectrum and phase difference matrix are constructed, and the spatiotemporal feature tensor is constructed in combination with the timestamp;

[0133] In this embodiment, the Doppler shift spectrum refers to the difference between the received signal frequency and the transmitted frequency when the signal source and receiver are in relative motion. This frequency change is caused by the relative motion between the transmitting source and the receiver. According to the Doppler effect, the frequency of the signal changes when the signal source (or receiver) moves relative to the receiver. For the receiver, the frequency increases as the signal source approaches and decreases as it moves away. This frequency change is called the Doppler shift spectrum.

[0134] S32: Using an improved particle filter algorithm to perform interference compensation on the spatiotemporal feature tensor to obtain a spatiotemporal compensated feature tensor;

[0135] S33: Based on the dynamic Doppler frequency shift spectrum and the visual positioning coordinates of the target person, the synchronized device of the target person is selected from all devices as the target device;

[0136] S34: performing data fusion on the spatiotemporal compensation feature tensor based on the signal quality index of the wireless signal characteristics of the target device, and calculating the positioning fusion data of the target device;

[0137] S35: Use a deep metric learning algorithm combined with a hybrid loss function and a multi-lateral positioning function to convert the target device's positioning fusion data into wireless positioning data;

[0138] S36: Using the heterogeneous coordinate system conversion matrix combined with the world coordinate system, the wireless positioning data is converted to obtain the wireless positioning coordinates of the target device.

[0139] In the present invention, step S4 includes the following steps:

[0140] Based on the relationship between the target device and the target person, the visual positioning coordinates and the wireless positioning coordinates are hierarchically fused:

[0141] If the target device is bound to the target person, the wireless positioning coordinates and visual positioning coordinates are fused at the original level and feature level according to the positioning accuracy;

[0142] The process of primitive level fusion is:

[0143] The wireless signal characteristics of the target device are calculated using the least squares method to obtain the original wireless distance between the target device and the wireless beacon;

[0144] Calculate the original visual distance between the target person and the wireless beacon based on the coordinate mapping matrix and the average depth of the target area;

[0145] According to the positioning accuracy, the weights corresponding to the wireless original distance and the visual original distance are estimated respectively to obtain the wireless dynamic weight and the visual dynamic weight;

[0146] Calculate the wireless dynamic weight, visual dynamic weight, wireless positioning coordinates and visual positioning coordinates to obtain the original level fusion positioning coordinates of the target person;

[0147] The process of feature-level fusion is as follows:

[0148] The weighted centroid algorithm is used to analyze the wireless signal characteristics of the target device in combination with the wireless positioning coordinates to obtain wireless depth positioning data;

[0149] The feature expansion algorithm is used to expand the visual positioning coordinates of the target person to obtain visual depth positioning data;

[0150] The wireless depth positioning data and the visual depth positioning data are weightedly fused to obtain the feature-level fused positioning coordinates of the target person.

[0151] If the target device is not bound to the target person, the wireless positioning coordinates and visual positioning coordinates are fused at the decision level according to the scene environment;

[0152] The process of decision-level fusion is as follows:

[0153] Calculate the absolute positioning difference between the wireless positioning coordinates and the visual positioning coordinates, and compare it with the positioning threshold:

[0154] If the absolute positioning difference is less than the positioning threshold, the wireless positioning coordinates and the visual positioning coordinates are weightedly fused to obtain the first-level decision-making fusion positioning coordinates of the target person, and the target device is determined to belong to the target person;

[0155] If the absolute positioning difference is greater than or equal to the positioning threshold, it is determined that the target device does not belong to the target person;

[0156] If all N absolute positioning differences are less than the positioning threshold, all absolute positioning differences are sorted from small to large, and the visual positioning coordinates corresponding to the smallest absolute positioning difference are selected for weighted fusion with the wireless positioning coordinates to obtain the first-level decision-making fusion positioning coordinates of the target person, and the target device is determined to belong to the target person;

[0157] Based on the human body bounding box and part bounding box in all frame fusion images, combined with the timestamp, a person trajectory sequence is generated; based on the dynamic Doppler frequency shift spectrum combined with the wireless positioning coordinates of the target device and the timestamp, a device trajectory sequence is generated; the trajectory matching degree between the person trajectory sequence and the device trajectory sequence is calculated and compared with the matching threshold:

[0158] If the trajectory matching degree is greater than the matching threshold, the wireless positioning coordinates and the visual positioning coordinates are weightedly fused to obtain the second-level decision-making fusion positioning coordinates of the target person, and the target device is determined to belong to the target person;

[0159] If the trajectory matching degree is less than or equal to the matching threshold, it is determined that the target device does not belong to the target person;

[0160] The distance between the target person and the camera device is measured using the confidence assessment depth map to obtain the person distance; the distance between the target device and the camera device is measured based on the geometric relationship between the wireless positioning coordinates and the camera device to obtain the device distance;

[0161] Calculate the absolute distance difference between the person's distance and the device's distance, and compare it with the distance threshold range based on physical constraints:

[0162] If the absolute distance difference is within the distance threshold range, it is determined that the wireless positioning coordinates and the visual positioning coordinates are weightedly fused to obtain the third-level decision-making fused positioning coordinates of the target person, and the target device is determined to belong to the target person;

[0163] If the absolute distance difference is not within the distance threshold range, it is determined that the target device does not belong to the target person;

[0164] Define the binding status of each device location and the target person location;

[0165] Calculate the posterior probability of the binding state based on the wireless positioning coordinates and visual positioning coordinates combined with the dynamic Doppler frequency shift spectrum and phase difference matrix;

[0166] The maximum a posteriori probability is combined with the hierarchical decision algorithm to determine the binding status between the target device and the target person;

[0167] According to the trajectory sequence of wireless positioning coordinates and visual positioning coordinates, the binding status is dynamically updated to obtain the fourth decision-level fusion positioning coordinates of the target person, and it is determined that the target device belongs to the target person, thereby completing the indoor positioning of the person.

[0168] Example:

[0169] The key to monocular ranging is to accurately know the intrinsic and extrinsic parameters of the camera device. Therefore, the camera device needs to be calibrated first, mainly including intrinsic and extrinsic parameter calibration.

[0170] Intrinsic calibration is the core parameter that relates the image coordinate system to the camera coordinate system, and is also a necessary parameter for converting relative depth to absolute depth.

[0171] Use a camera to take multiple images of the calibration plate, use Open Vision and other visual tools to calibrate, and calculate the intrinsic parameter matrix K of the camera, including the focal length (f x ,f y ), optical center coordinates (C x ,C y ).

[0172] External parameter calibration establishes the mapping relationship between the camera device coordinate system and the world coordinate system. Indoor positioning is usually two-dimensional positioning, and the height relationship between the camera device and the person is fixed. Therefore, the optical center coordinates of the intrinsic parameter matrix and the actual height of the camera device can be obtained by laser ranging, and the coordinate mapping can be established in combination with the height of the camera device.

[0173] A lightweight object detection model is used to perform regular, low-frequency detection on the camera's video stream to conserve computing power. When a person is detected, the sampling frame rate is increased for continuous detection, while an image segmentation model is used to determine the target's segmented region. This phase clearly identifies the target region within the image that requires ranging.

[0174] Using the depth estimation model, a depth map corresponding to the video frame is generated, and the average depth of the target segmented area is obtained. The depth estimation algorithm used in the present invention is pre-trained based on a large-scale, multi-scene dataset, and can directly derive high-precision pixel-level depth predictions from images under different lighting, material and geometric conditions. Most traditional monocular ranging models require estimating the actual depth based on the size of the target through geometric relationships. This algorithm greatly simplifies the calculation process and avoids cumulative errors. To further improve accuracy, indoor scenes can be combined to collect small-scale datasets and perform low-cost fine-tuning on the model.

[0175] The world coordinate system is established with the vertical projection of the camera device on the ground as the origin, and the height of the camera device is h C The intersection of the optical axis of the camera device and the ground is C, the position of the person from the ground to the top of the head is P0P1, and the height of the person is h P , the midpoint is P, then the height of P is h P / 2. The distance from the camera to point P is the depth d estimated by the algorithm. To further improve accuracy, facial recognition technology can be used to obtain the person's estimated gender and age, adopt a smoother height estimate, or use the precise height data associated with Face ID.

[0176] Without changing the existing wireless positioning system layout, a three-level data fusion approach can fully utilize existing wireless networks and surveillance cameras. Wireless positioning technology infers the location of the target device by collecting wireless signal characteristics. Common signal characteristics include signal strength (RSSI), time of flight (ToF), and angle of arrival (AoA). RSSI and ToF can both estimate or calculate the distance between the device and the wireless beacon. For algorithms that use distance for positioning (such as trilateration and multilateration, least squares method), visual distance estimation can be directly included in the calculation. The accuracy of visual distance estimation is generally much better than that of wireless distance measurement and is given greater weight.

[0177] For algorithms that directly use signal features for positioning (weighted centroid algorithms, Kalman filters, particle filters, Bayesian filters, deep neural networks, etc.), visually estimated depth information can also be incorporated into the calculation as a visual feature. The visual features involved in fusion have expanded from the average depth of the target area (a single value) to depth information of arbitrary granularity (multi-dimensional features including image horizontal coordinates, image vertical coordinates, and depth). This fully utilizes the high resolution, richness, and local detail capabilities of visual features, combined with the global positioning capabilities of wireless signals, to achieve higher-precision multimodal positioning.

[0178] The first two levels of data fusion rely on binding wireless receiving devices (mobile phones, tablets, Bluetooth badges, etc., whose location is obtained wirelessly) with people (whose location is obtained visually). Currently, most lightweight visual models support facial recognition, and people can be bound to devices through methods such as gate recognition and camera capture.

[0179] In the following cases, the binding relationship between the character and the device cannot be directly obtained:

[0180] a. Limited face recognition: The camera cannot recognize or capture faces due to factors such as insufficient resolution, occlusion, poor lighting, or shooting angle.

[0181] b. The device is not explicitly bound to the user: for example, a shared device (tablet, work badge) or a device shared by multiple people.

[0182] c. Dynamic scenes: When a person moves quickly or enters an obstructed area, the binding between the person and the device they carry may be temporarily lost.

[0183] Through decision-level fusion, the independent positioning information of wireless and visual is combined, and the possible binding relationship between the device and the person is inferred through algorithms, thereby achieving personnel positioning.

[0184] Distance-based matching: If the wireless and visual positioning results are close (i.e., the distance in physical space is small), the device can be inferred to belong to the person. If the distance between a device's wireless positioning result and multiple visual positioning results is less than a set threshold, the closest visual positioning result is selected. This algorithm is suitable for simple environments with low device and person density (such as offices or small meeting rooms). When wireless and visual positioning accuracy are high, it can quickly and efficiently complete binding.

[0185] Motion trajectory-based matching: In dynamic scenarios, the binding relationship can be inferred by matching the trajectories of devices and people. If the time-series trajectories of wireless and visual positioning results are similar, the device can be determined to belong to the person. This algorithm is suitable for scenarios with frequent human movement and where the binding between devices and people needs to be dynamically updated (such as shopping malls and factory floors). It can also handle short-term binding loss due to occlusion or signal fluctuations.

[0186] Determination based on depth information and scene constraints: In complex scenes, the depth map information and scene constraints are combined to further confirm the binding relationship through an inference algorithm. When a person is identified by visual positioning, the depth map is used to obtain the distance between the person and the camera device. At the same time, the distance between the device and the camera device is estimated by combining the wireless positioning results with the geometric relationship of the camera device. If the wireless depth information matches the visual depth information (the error is within the threshold range), binding is performed. Based on the physical constraints in the scene (such as room layout, personnel density, etc.), unreasonable binding results are eliminated. This algorithm is suitable for indoor environments with rich depth map data and known scene structure (such as smart home or factory monitoring).

[0187] Fusion based on probabilistic graphical models: Using probabilistic graphical models such as Bayesian inference or hidden Markov models (HMMs), wireless and visual positioning results are jointly modeled to dynamically infer the binding relationship between devices and people. The binding status of each device and person's position is first defined. The posterior probability of the binding status is calculated based on the wireless and visual positioning results. Finally, the maximum a posteriori probability (MAP) is used to determine the binding status. Over time, the binding status is dynamically updated based on changes in the wireless and visual positioning results. This algorithm is suitable for complex multi-target scenarios (such as shopping malls and parking lots) and can handle binding conflicts between multiple devices.

[0188] Fingerprinting, a key branch of wireless positioning technology, introduces the concept of decision-level fusion, enabling deep integration of visual features into the positioning process. The basic principle of the fingerprinting method is to pre-collect signal characteristics (such as RSSI and CSI) from various reference points in an indoor environment and store them in a fingerprint database. During positioning, the target device's real-time signal characteristics are compared with the data in the fingerprint database to find the best match.

[0189] After the introduction of visual features, the fingerprint library of each reference point not only contains wireless signal characteristics, but also adds background depth information and its statistical characteristics (such as average depth, depth variance, depth histogram, etc.) collected under different lighting conditions. In the online stage, the appearance and movement of people will significantly change the real-time depth map information. By comparing it with the background depth characteristics in the fingerprint library, the person's position can be independently determined, thereby completing the indoor positioning of the target person.

Claims

1. A multimodal indoor positioning method integrating radio frequency and machine vision, characterized by: The method comprises the following steps in sequence: S1: Calibrate and start the camera device, and establish an internal and external coordinate mapping matrix based on the camera device height. At the same time, collect and process the wireless signals in the room to obtain the wireless signal characteristics of all devices; S2: Detect and segment the target area in the video stream of the camera device, combine the depth estimation model to measure the average depth of the target area, and obtain the visual positioning coordinates of the target person according to the coordinate mapping matrix; S3: Calculate the wireless positioning coordinates of the target device based on the target area and wireless signal characteristics; S4: Based on the data level, a hierarchical decision algorithm is used to hierarchically fuse the visual positioning coordinates and wireless positioning coordinates to obtain the fused positioning coordinates of the target person, thereby completing the indoor positioning of the person.

2. The multimodal indoor positioning method integrating radio frequency and machine vision according to claim 1, characterized in that: The step S1 includes the following steps: S11: Start the camera device to take several calibration plate images; S12: Calculate the intrinsic parameter matrix of the camera device according to the calibration plate image; S13: Obtaining an extrinsic parameter matrix of the camera device according to the camera device coordinate system and the world coordinate system; S14: Establishing an internal and external coordinate mapping matrix by combining the internal parameter matrix and the external parameter matrix according to the height of the camera device.

3. The multimodal indoor positioning method integrating radio frequency and machine vision according to claim 1, characterized in that: The step S1 further includes the following steps: S15: Capture the radio frequencies emitted by all devices in the room and obtain several wireless signals; S16: aligning the multiple wireless signals in the time domain to obtain a wireless synchronization signal; S17: De-noising and clustering the wireless synchronization signal to obtain a noise-free synchronization signal; S18: performing data enhancement on the noise-free synchronization signal to obtain a multi-mode synchronization signal; S19: Extract features from the time domain, frequency domain, and spatial domain of the multi-mode synchronous signal to obtain wireless signal features of all devices.

4. The multimodal indoor positioning method integrating radio frequency and machine vision according to claim 1, characterized in that: The step S2 includes the following steps: S201: Receive a video stream captured by a camera device with zero frame loss, and convert the video stream into a plurality of frame images; S202: Obtaining a color-free image by eliminating color difference from each frame of image; S203: compressing and fusing the non-color-biased image to obtain a fused image; S204: The fused image is detected by combining the detection branch of the hybrid perception model with the human body key point heat map to obtain the posture confidence, and then compared with the preset confidence threshold for judgment: If the posture confidence is greater than or equal to the confidence threshold, the target person in the current frame fusion image is marked using the human body bounding box. At the same time, the human body bounding box is combined with the body texture attributes to decompose into several part element points and match them with the body parts in all frame fusion images: If any part element point has a matching body part, the body part is marked and the part bounding box is obtained until there is no body part of the target person in the fused image; If any part element point has no matching body part, it means that the target person has left the current video stream; If the posture confidence is less than the confidence threshold, it means that the suspicious person in the current fused image has a similar posture to the target person. At the same time, the fused image is marked to obtain a suspicious bounding box. S205: Segment the human body bounding box, part bounding box, and suspicious bounding box through the segmentation branch of the hybrid perception model combined with the occlusion coefficient matrix to obtain an instance-level mask.

5. The multimodal indoor positioning method integrating radio frequency and machine vision according to claim 1, characterized in that: The step S2 further includes the following steps: S206: Generate a confidence assessment depth map based on the fused images of all frames using a depth estimation model combined with an attention gating mechanism; S207: Segment the target person in each fused image frame according to the instance-level mask to obtain a target region; S208: Calculating the depth of the target area according to the confidence evaluation depth map to obtain an average depth of the target area; S209: Establishing a world coordinate system with the vertical projection of the camera device on the ground as the origin, and obtaining the height of the camera device and the height of the target person; S210: Calculating the height of the camera device and the height of the target person according to the coordinate mapping matrix and the average depth of the target area to obtain the visual positioning coordinates of the target person.

6. The multimodal indoor positioning method integrating radio frequency and machine vision according to claim 1, characterized in that: The step S3 includes the following steps: S31: Based on the wireless signal characteristics of all devices, a dynamic Doppler frequency shift spectrum and phase difference matrix are constructed, and the spatiotemporal feature tensor is constructed in combination with the timestamp; S32: Using an improved particle filter algorithm to perform interference compensation on the spatiotemporal feature tensor to obtain a spatiotemporal compensated feature tensor; S33: Based on the dynamic Doppler frequency shift spectrum and the visual positioning coordinates of the target person, the synchronized device of the target person is selected from all devices as the target device; S34: performing data fusion on the spatiotemporal compensation feature tensor based on the signal quality index of the wireless signal characteristics of the target device, and calculating the positioning fusion data of the target device; S35: Use a deep metric learning algorithm combined with a hybrid loss function and a multi-lateral positioning function to convert the target device's positioning fusion data into wireless positioning data; S36: Using the heterogeneous coordinate system conversion matrix combined with the world coordinate system, the wireless positioning data is converted to obtain the wireless positioning coordinates of the target device.

7. The multimodal indoor positioning method integrating radio frequency and machine vision according to claim 1, characterized in that: The step S4 includes the following steps: Based on the relationship between the target device and the target person, the visual positioning coordinates and the wireless positioning coordinates are hierarchically fused: If the target device is bound to the target person, the wireless positioning coordinates and visual positioning coordinates are fused at the original level and feature level according to the positioning accuracy; If the target device is not bound to the target person, the wireless positioning coordinates and visual positioning coordinates are fused at the decision level according to the scene environment.

8. The multimodal indoor positioning method integrating radio frequency and machine vision according to claim 7, characterized in that: The process of the original level fusion is as follows: The wireless signal characteristics of the target device are calculated using the least squares method to obtain the original wireless distance between the target device and the wireless beacon; Calculate the original visual distance between the target person and the wireless beacon based on the coordinate mapping matrix and the average depth of the target area; According to the positioning accuracy, the weights corresponding to the wireless original distance and the visual original distance are estimated respectively to obtain the wireless dynamic weight and the visual dynamic weight; The wireless dynamic weight, visual dynamic weight, wireless positioning coordinates and visual positioning coordinates are calculated to obtain the original level fusion positioning coordinates of the target person.

9. The multimodal indoor positioning method integrating radio frequency and machine vision according to claim 7, characterized in that: The process of feature-level fusion is as follows: Analyze the wireless signal characteristics of the target device in combination with the wireless positioning coordinates to obtain wireless deep positioning data; Expand the visual positioning coordinates of the target person to obtain visual depth positioning data; The wireless depth positioning data and the visual depth positioning data are weightedly fused to obtain the feature-level fused positioning coordinates of the target person.

10. The multimodal indoor positioning method integrating radio frequency and machine vision according to claim 7, characterized in that: The process of decision-level fusion is as follows: Calculate the absolute positioning difference between the wireless positioning coordinates and the visual positioning coordinates, and compare it with the positioning threshold: If the absolute positioning difference is less than the positioning threshold, the wireless positioning coordinates and the visual positioning coordinates are weightedly fused to obtain the first-level decision-making fusion positioning coordinates of the target person, and the target device is determined to belong to the target person; If the absolute positioning difference is greater than or equal to the positioning threshold, it is determined that the target device does not belong to the target person; If all N absolute positioning differences are less than the positioning threshold, all absolute positioning differences are sorted from small to large, and the visual positioning coordinates corresponding to the smallest absolute positioning difference are selected for weighted fusion with the wireless positioning coordinates to obtain the first-level decision-making fusion positioning coordinates of the target person, and the target device is determined to belong to the target person; Based on the human body bounding box and part bounding box in all frame fusion images, combined with the timestamp, a person trajectory sequence is generated; based on the dynamic Doppler frequency shift spectrum combined with the wireless positioning coordinates of the target device and the timestamp, a device trajectory sequence is generated; the trajectory matching degree between the person trajectory sequence and the device trajectory sequence is calculated and compared with the matching threshold: If the trajectory matching degree is greater than the matching threshold, the wireless positioning coordinates and the visual positioning coordinates are weightedly fused to obtain the second-level decision-making fusion positioning coordinates of the target person, and the target device is determined to belong to the target person; If the trajectory matching degree is less than or equal to the matching threshold, it is determined that the target device does not belong to the target person; The distance between the target person and the camera device is measured using the confidence assessment depth map to obtain the person distance; the distance between the target device and the camera device is measured based on the geometric relationship between the wireless positioning coordinates and the camera device to obtain the device distance; Calculate the absolute distance difference between the person's distance and the device's distance, and compare it with the distance threshold range based on physical constraints: If the absolute distance difference is within the distance threshold range, it is determined that the wireless positioning coordinates and the visual positioning coordinates are weightedly fused to obtain the third-level decision-making fused positioning coordinates of the target person, and the target device is determined to belong to the target person; If the absolute distance difference is not within the distance threshold range, it is determined that the target device does not belong to the target person; Define the binding status of each device location and the target person location; Calculate the posterior probability of the binding state based on the wireless positioning coordinates and visual positioning coordinates combined with the dynamic Doppler frequency shift spectrum and phase difference matrix; The maximum a posteriori probability is combined with the hierarchical decision algorithm to determine the binding status between the target device and the target person; According to the trajectory sequence of wireless positioning coordinates and visual positioning coordinates, the binding status is dynamically updated to obtain the fourth decision-level fusion positioning coordinates of the target person, and it is determined that the target device belongs to the target person, thereby completing the indoor positioning of the person.

Citation Information

Patent Citations

  • Multi-mode indoor positioning and navigation method and system

    CN111664848A

  • CSI indoor positioning method based on multi-mode GAN

    CN114745684A

  • Semi-physical simulation method of inertial navigation / laser velocimeter combined navigation

    CN104596515A

  • Indoor fusion positioning method based on multi-features

    CN106123897A

  • Systems and methods for indoor positioning using wireless positioning nodes

    CN113573268A