Indoor mobile terminal visual navigation positioning method and system based on video stream transmission
Through the lightweight video streaming of YOLO and visual SLAM algorithms, combined with monocular camera depth estimation and PnP feature matching, the high cost and deployment problems of traditional indoor positioning technology are solved, and efficient, real-time navigation positioning and repositioning of indoor mobile terminals is achieved, and rescue efficiency and security are improved.
Patent Information
- Application Number
- CN202510269452.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-18
AI Technical Summary
In complex and dangerous indoor emergency rescue sites, traditional positioning technology relies on external sensors and base stations, which are costly and difficult to deploy quickly, and cannot locate rescue personnel in real time, resulting in inefficient rescue efficiency and susceptible to environmental damage in emergencies.
The lightweight YOLO object detection algorithm and visual SLAM algorithm are used to establish communication through video streaming, combining monocular camera depth estimation and PnP feature matching to realize high-precision navigation and positioning of indoor mobile terminals, and restore positioning through repositioning function when communication is interrupted.
It realizes fast and real-time high-precision navigation and positioning in an indoor environment, reduces equipment computing needs, has environmental information and location information sharing functions, and improves rescue efficiency and personnel safety.
Smart Images

Figure CN120333411A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of positioning technologies, and particularly relates to an indoor mobile visual navigation and positioning method and system based on video stream transmission. Background Art
[0002] In recent years, with the advancement of the urbanization process, the scale of building complexes has been continuously expanding. The living and production demands of using fire and electricity indoors have also led to frequent accidental accidents such as fires and explosions, causing heavy casualties and property losses. Such accidents occur suddenly, with complex and changeable scenes, endangering the lives of rescue workers and trapped people. In the face of complex and dangerous indoor emergency rescue, individual combat, fire fighting and disaster relief and other scenes, how to ensure the safety of personnel during the operation and effectively exert the individual combat ability, and improve the rescue efficiency is a difficult problem to be solved in the relevant fields.
[0003] The operating environment is intricate and complex. It is difficult for on-site operators and backstage commanders to quickly understand the situation of the rescue site, and it is impossible to confirm their own position information and the environmental structure, resulting in low rescue efficiency. If the backstage command center can real-time locate the operators and formulate the best plan according to the actual situation, it can significantly reduce the ineffective search time, improve the overall operation efficiency, and at the same time further ensure the safety of rescue workers. However, some traditional indoor positioning technologies rely on external sensors, infrastructure, and pre-deployed positioning base stations, with high large-scale costs and difficult to deploy in a timely manner in sudden tasks, and are often unable to be used due to environmental damage (such as building internal power system failures, etc.). Summary of the Invention
[0004] Object of the Invention: In order to solve the problems existing in the above-mentioned prior art, the present invention provides an indoor mobile visual navigation and positioning method and system based on video stream transmission.
[0005] Technical Solution: The present invention provides an indoor mobile visual navigation and positioning method based on video stream transmission, specifically including the following steps:
[0006] Step 1: Perform lightweight processing on the YOLO module to obtain a lightweight YOLO module;
[0007] Step 2: Establish communication between the SLAM module and the lightweight YOLO module;
[0008] Step 3: The SLAM module and the lightweight YOLO module simultaneously acquire the image information of the visual front end; at time t, if the lightweight YOLO module recognizes a target in the image, the lightweight YOLO module sends the size of the target and the pixel coordinate information of the target in the image to the SLAM module; record the timestamp corresponding to this image as x;
[0009] Step 4: The SLAM module constructs a depth scale factor based on the received information, and the SLAM module sends the timestamp of the image obtained at time t to the lightweight YOLO module;
[0010] Step 5: The lightweight YOLO module compares the received timestamp with the timestamp x. If the two are consistent, the initialization of the SLAM module is completed, and the SLAM module corrects the parameters in the mapping process based on the depth scale factor. Otherwise, the lightweight YOLO module continues to wait for the timestamp sent by the SLAM module.
[0011] Furthermore, the specific content of step 1 is as follows: Apply L1 regularization to the scaling factors in the BN layer of the CNN structure of the YOLO module to obtain the regularized scaling factors corresponding to each channel in the BN layer. Delete the channels corresponding to the regularized scaling factors that are less than the preset threshold, and retrain the YOLO module after deleting the channels to adjust the module accuracy to obtain the lightweight YOLO module.
[0012] Furthermore, the specific content of step 2 is as follows: Use the camera as a publishing node to publish ROS-format visual image messages to topic 1. Use the SLAM module and the lightweight YOLO module as subscribers, both subscribing to topic 1. After the lightweight YOLO module detects a target, it publishes topic 2, and the SLAM module is used as a subscriber to topic 2.
[0013] Furthermore, in step 2, the SLAM module and the lightweight YOLO module use the Socket socket technology to identify communication endpoints in the form of file paths and establish a connection.
[0014] Furthermore, the expression of the depth scale factor σ in step 4 is:
[0015]
[0016] where D1 represents the relative depth estimated by triangulation, and D2 is the true depth calculated by the lightweight YOLO module.
[0017] Furthermore, if the SLAM tracking is lost, a relocalization method based on PnP feature matching and key frame matching is adopted.
[0018] A system for an indoor mobile visual navigation and positioning method based on video stream transmission includes an application layer, a hardware layer, and an intermediate layer;
[0019] The application layer includes a Master system, a camera node, a lightweight YOLO module, and a SLAM module; the Master system is responsible for the communication between various modules or nodes within the application layer;
[0020] The function of the middle layer is to connect the application layer and the hardware layer, and coordinate and manage computer resources and network communication;
[0021] The hardware layer includes mobile terminals, computers, and wireless routers. The mobile terminal calls a fixed-focus monocular camera to obtain front-end visual information. The PC runs lightweight YOLO algorithms and visual SLAM algorithms for navigation and positioning, and the wireless router provides a local area network.
[0022] Beneficial effects:
[0023] 1. The present invention performs lightweight processing such as network pruning on the YOLO object detection algorithm based on deep learning, which can better retain the detection accuracy of the network model, improve the inference speed, reduce the model size and the number of parameters, reduce the computing requirements for edge devices or mobile devices, and further estimate the scene depth information in the monocular image by quickly identifying cooperative targets in the indoor environment.
[0024] 2. The present invention proposes to establish data communication between the object detection model and the visual SLAM positioning algorithm using ROS or Socket socket technology, which can achieve high-precision scene depth estimation in the indoor space, meet the requirements of rapid initialization of visual SLAM, and provide a true scale for subsequent visual odometry mapping and pose positioning.
[0025] 3. The indoor mobile visual positioning technology based on video stream transmission proposed by the present invention makes full use of the convenience of mobile and wireless communication technologies, overcomes the limitations of traditional positioning methods, can perform better remote and real-time navigation and positioning of mobile terminals in weak-texture indoor environments with wireless local area network signals, can achieve repositioning through PnP feature matching and key frame matching after the tracking and positioning is interrupted, and has the function of sharing environmental information and location information. Description of the drawings
[0026] Figure 1 It is a process diagram of program initialization based on ROS communication;
[0027] Figure 2 It is a schematic flow diagram of the network pruning process;
[0028] Figure 3 It is a framework diagram of the ROS system based on the mobile video stream;
[0029] Figure 4 It is a diagram of the local area network communication method based on ROS video stream transmission. Detailed implementation manners
[0030] The drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.
[0031] The flowchart of this embodiment is as Figure 1 shown, specifically as follows:
[0032] 1. First, establish communication between the lightweight YOLO algorithm and the monocular vision SLAM algorithm in the Ubuntu system based on the Linux operating system through the ROS topic publishing and subscribing method, or identify the communication endpoints in the form of file paths through the Socket socket technology and establish a connection.
[0033] 2. Initialize and start the monocular SLAM program and start the lightweight YOLO algorithm to jointly obtain the image information of the visual front end; the target detection algorithm (lightweight YOLO algorithm) identifies the cooperative target and outputs the pixel coordinate information of the detected target in the image; in this embodiment, the door in the image is used as the detection target.
[0034] 3. The SLAM program reads the pixel coordinates, combines the prior size information, and calculates the actual depth information using the monocular ranging principle, and feeds back the timestamp of the current image frame to the lightweight YOLO program.
[0035] 4. After receiving the timestamp information, the lightweight YOLO program will compare it with the timestamp of the detection frame. If they are the same, it will perform the next frame of detection; otherwise, it will wait continuously.
[0036] 5. While feeding back the timestamp information, the SLAM program will use the calculated depth information to construct a scale factor, which is used to correct the relative scale obtained by triangulation and complete the SLAM initialization process.
[0037] In this embodiment, the lightweight YOLO algorithm is obtained through network pruning. Specifically, a cooperative target training dataset is constructed, which contains 3000 indoor "door" pictures in total. Network pruning and other lightweight processing are performed on the YOLOv5s target detection algorithm based on deep learning to reduce its model parameters, weights, and computational overhead, and training is carried out under this training set to improve the detection efficiency of the detection model for the indoor cooperative target "door".
[0038] The training dataset constructed in this embodiment contains N cooperative target photos in total. The network pruning process is as Figure 2 shown. This process is directly executed in the CNN structure. By implementing channel-level sparsification for pruning operations, the computing power requirements for devices can be effectively reduced. Using the method of channel reduction, the network model becomes more compact and concise, and the detection accuracy after pruning can be comparable to that of the original model. Network pruning does not need to change the structure of the CNN. Only L1 regularization needs to be implemented on the scaling factor of the BN layer. Through this method, the scaling factor of the BN layer can be reduced to zero, so as to realize the identification and elimination of irrelevant channels in the network, as shown in the following formula:
[0039]
[0040] Among them, L represents the loss function, f(x, W) is the output of the model, x is the input sample data, W represents the model weight parameters, y is the true label of each input sample x, and l(f(x, W), y) is the loss between the model input and the true label y. In the second term, λ is the weight coefficient of regularization; is the regularization term, which is used to control the complexity of the model parameters; Γ represents a set that contains all the scaling factors γ to be regularized; g(·) is the sparsity penalty generated by the scaling factor. During training, g(s) = |s|, and the L1 norm method is used for sparsification. The scaling factor γ is the scaling coefficient in BN. Changing the initial value of γ results in very different sparsity effects. Therefore, in practical applications, the scaling factor of the BN layer is used as the initial value of the scaling factor γ.
[0041] BN has excellent generalization ability and is widely used in most convolutional neural networks to promote fast convergence. Let Z in and Z out respectively represent the normalized values of the input and output of the BN layer. The expression of the BN layer is as follows:
[0042]
[0043]
[0044] Among them, is the normalized value, μ B represents the input average value, σ B represents the input standard deviation, B represents a set of data in the current batch, ∈ is a small constant used to avoid division-by-zero errors during the normalization process, β represents the offset parameter, and γ and β change during training, indicating that the scale of the normalized activation can undergo arbitrary linear changes.
[0045] Adding an additional regularization term for γ in the network usually does not have a negative impact on the original performance. By directly using the parameters in the BN layer as the channel pruning scale factors, effective channel pruning scale factors can be obtained conveniently and efficiently.
[0046] In a convolutional neural network, each convolutional layer usually has multiple channels, and each channel corresponds to a specific convolutional kernel. During the pruning process, when implementing global thresholding by introducing a scaling factor for all convolutional layers in the neural network, if the scaling factor of a channel is observed to be less than the preset threshold, all the input, output connections, and corresponding weights of that channel are removed. The pruned network model becomes more compact and lightweight in terms of the number of parameters, memory occupancy, and computational efficiency. The resulting accuracy loss can be compensated through a fine-tuning step.
[0047] There are three steps in total for model pruning, namely sparse training, channel pruning, and model fine-tuning. Considering the gentle sparse change curve of the BN layer, selecting a sparsity rate of p for sparse training and then pruning and fine-tuning the channels at a ratio of q can better retain the model accuracy and improve the inference speed. In this embodiment, the sparsity rate p is selected as 0.028 for sparse training, and then the channels are pruned and fine-tuned at a ratio of q = 30%.
[0048] In step 1, the ROS system architecture based on the mobile video stream is as Figure 3 shown. The application layer mainly includes application programs on the mobile side and the PC side, such as the mobile camera node, the PC SLAM program node, the lightweight YOLO program node, etc., which divides and modularizes the functions of the navigation system designed in this embodiment. The Master is responsible for the mutual search and communication of each functional node in the system.
[0049] The middle layer located between the application layer and the operating system layer is mainly responsible for coordinating and managing computer resources and network communication. As a bridge, it connects independent application programs or even software of different operating systems, and it contains the client libraries encapsulated by ROS, such as roscpp, rospy, rosjava, etc., that is, the C++ library, Python library, and Java language library of ROS. The mobile camera node, SLAM program node, and lightweight YOLO program node involved in this system are based on the JAVA, C++, and Python languages respectively, and involve the Android system and the Linux system.
[0050] The hardware layer mainly includes a mobile device, a desktop computer or a laptop, and a wireless router that provides a local area network WiFi network. The mobile device calls a fixed-focus monocular camera to obtain front-end visual information. The PC runs the lightweight YOLO algorithm and the visual SLAM algorithm for navigation and positioning, and the wireless router provides a local area network WiFi network for both.
[0051] The PC and the mobile device are required to be connected to the same local area network. The mobile device connects to the IP address of the PC, and publishes the video stream data obtained by the camera through the ROS topic. The ROS node on the PC subscribes to this topic to obtain the video stream data in real time and synchronously runs the target detection program and the visual SLAM program, realizing the indoor mobile visual positioning function based on video stream transmission. Attached Figure 4 shows the local area network communication method based on ROS video stream transmission.
[0052] First, the mobile camera node, as a publishing node, publishes visual image messages in ROS format to "Topic 1". Secondly, the lightweight target detection program and the visual SLAM program on the PC respectively act as subscribers and subscribe to "Topic 1". The ROS system converts the ROS image message into an OpenCV image message through the built-in cv_Bridge library, which is convenient for the PC-side algorithm program to read.
[0053] In this process, the lightweight YOLO algorithm will detect the image frames in real time. If a cooperative target is found, it will act as a publishing node to publish the message and store it in "Topic 2". The SLAM node still acts as a subscriber and subscribes to "Topic 2".
[0054] In step 3, the real scene depth estimation of the monocular camera is realized by combining the target detection result and the target prior size information. In the pinhole imaging process of the camera, D is the real distance from the three-dimensional object cooperative target to the camera, that is, the scene depth; F is the focal length of the fixed-focus monocular camera; W is the actual height or width of the cooperative target; H is the number of pixels occupied by the cooperative target projected on the pixel plane. According to the mathematical relationship of similar triangles, there is:
[0055]
[0056] The fixed focal length F of the monocular camera can be obtained through camera calibration methods such as Zhang Zhengyou et al., and the number of pixels H can also be obtained by solving the image program. According to the above formula, when the prior information W of the cooperative target is known, the real depth information D can be obtained through the monocular ranging principle.
[0057] Assume that the relative depth estimated by triangulation is D1, and the real depth calculated by the monocular ranging principle is D2, then the depth scale factor σ can be constructed:
[0058]
[0059] The depth scale factor σ can be used to correct the relative scale obtained in the local map building process of SLAM, correct the coordinates and translations of the map points, and obtain more accurate navigation and positioning information.
[0060] In this embodiment, if the SLAM positioning process is suddenly interrupted, its relocalization function is implemented by PnP feature matching and the matching principle between the current frame and the key frames.
[0061] The goal of PnP feature matching is to find the correspondence between the observation points of the current frame and the known feature points in the map. Let the feature points extracted from the current frame be p i =(u i , v i ), representing the coordinates in the image coordinate system; the 3D point in the map is P i =(X i , Y i , Z i ), representing the coordinates in the world coordinate system; the projection model of the camera can be described by the following formula:
[0062] p i =K·[R|t]·P i
[0063]
[0064] where K represents the camera intrinsic matrix; R∈SO(3) and t∈R 3 are the rotation matrix and translation vector of the camera, which define the pose of the camera in the world coordinate system; p i is the pixel coordinate of the feature point, and P i is the corresponding world coordinate; the goal of relocalization is to estimate R and t.
[0065] Through a set of matched point pairs {(p i , P i )}, relocalization can be regarded as a PnP problem:
[0066]
[0067] where π(K, R, t, P i ) represents the pixel coordinates obtained by projecting the 3D point Pi onto the image plane, and the goal of optimization is to minimize the reprojection error. The finally solved R and t are the pose of the current frame relative to the map.
[0068] The matching principle between the current frame and the key frames is to find the similarity between the current frame and some key frames through the key frame database stored in the map, and select candidate key frames from them. First, extract the feature points from the current frame and calculate their descriptors. Second, use the bag-of-words model or direct descriptor matching to quickly retrieve the most similar key frames by vector similarity or nearest neighbor search. Finally, for the candidate key frames, use geometric constraints (such as the fundamental matrix F or the essential matrix E) to verify the relative relationship between the current frame and the key frames. After eliminating the false matches, the most reliable candidate key frames are retained.
[0069] By combining the PnP feature matching and the matching between the current frame and the key frame, the fast and accurate relocalization function of visual navigation can be achieved.
[0070] In addition, it should be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the present invention will not separately describe various possible combination methods.
Claims
1. An indoor mobile visual navigation and positioning method based on video stream transmission, characterized in that, Specifically, it includes the following steps: Step 1: Lightweight the YOLO module to obtain a lightweight YOLO module; Step 2: Establish communication between the SLAM module and the lightweight YOLO module; Step 3: The SLAM module and the lightweight YOLO module simultaneously obtain the image information of the visual front end; at time t, if the lightweight YOLO module recognizes a target in the image, the lightweight YOLO module sends the size of the target and the pixel coordinate information of the target in the image to the SLAM module; record the timestamp corresponding to this image as x; Step 4: The SLAM module constructs a depth scale factor based on the received information, and the SLAM module sends the timestamp of the image obtained at time t to the lightweight YOLO module; Step 5: The lightweight YOLO module compares the received timestamp with the timestamp x. If the two are consistent, the initialization of the SLAM module is completed, and the SLAM module corrects the parameters in the mapping process based on the depth scale factor. Otherwise, the lightweight YOLO module continues to wait for the timestamp sent by the SLAM module.
2. The indoor mobile visual navigation and positioning method based on video stream transmission according to claim 1, characterized in that, The specific content of Step 1 is as follows: Implement L1 regularization on the scaling factors in the BN layer of the CNN structure of the YOLO module to obtain the regularized scaling factors corresponding to each channel in the BN layer. Delete the channels corresponding to the regularized scaling factors that are less than the preset threshold, and retrain the YOLO module after deleting the channels to adjust the module accuracy to obtain a lightweight YOLO module.
3. The indoor mobile visual navigation and positioning method based on video stream transmission according to claim 1, wherein The specific content of Step 2 is as follows: Use the camera as a publishing node to publish ROS-formatted visual image messages to Topic 1. Use the SLAM module and the lightweight YOLO module as subscribers to subscribe to Topic 1. After the lightweight YOLO module detects a target, it publishes Topic 2, and use the SLAM module as a subscriber to Topic 2.
4. A method for indoor mobile visual navigation and positioning based on video stream transmission according to claim 1, characterized in that, In Step 2, the SLAM module and the lightweight YOLO module identify the communication endpoints in the form of file paths through the Socket socket technology and establish a connection.
5. A method for indoor mobile visual navigation and positioning based on video stream transmission according to claim 1, characterized in that The expression of the depth scale factor σ in Step 4 is: where D1 represents the relative depth estimated by triangulation, and D2 is the true depth calculated by the lightweight YOLO module.
6. A method for indoor mobile visual navigation and positioning based on video stream transmission according to claim 1, characterized in that If the SLAM tracking is lost, then use the PnP feature matching and key frame matching method for relocalization.
7. A system for an indoor mobile visual navigation and positioning method based on video stream transmission according to claim 1, characterized in that It includes an application layer, a hardware layer, and an intermediate layer; The application layer includes the Master system, the camera node, the lightweight YOLO module, and the SLAM module; the Master system is responsible for the communication between various modules or nodes within the application layer; The role of the intermediate layer is to connect the application layer and the hardware layer, and coordinate and manage computer resources and network communication; The hardware layer includes a mobile terminal, a computer, and a wireless router. The mobile terminal calls a fixed-focus monocular camera to obtain the front-end visual information. The PC runs the lightweight YOLO algorithm and the visual SLAM algorithm for navigation and positioning, and the wireless router provides a local area network.