A pose display method, device and system
By enabling terminal devices and servers of the cloud-edge management system to work together, a fused positioning trajectory that generates a 3D visual map using self-positioning trajectory and image data is generated. This solves the positioning problem of terminal devices in indoor environments, achieving high frame rate and high-precision indoor positioning, which is suitable for the positioning needs of the energy industry.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-15
- Publication Date
- 2026-03-24
AI Technical Summary
In indoor environments, due to poor GPS or BeiDou signals, terminal devices cannot achieve accurate positioning, especially in energy industries such as coal, electricity, and petrochemicals where indoor positioning needs cannot be met.
Through the cloud-edge management system, the terminal device acquires target images and motion data during movement, performs self-localization trajectory calculation, and selects some images to send to the server. The server generates a fused positioning trajectory in a 3D visual map based on this data, achieving high frame rate and high accuracy indoor positioning.
It achieves high frame rate and high accuracy indoor positioning, reduces the amount of data transmitted over the network, and lowers the computing and storage resource consumption of terminal devices, making it suitable for indoor positioning needs in energy industries such as coal, electricity, and petrochemicals.
Smart Images

Figure CN114185073B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision, and in particular to a pose display method, apparatus and system. Background Technology
[0002] GPS (Global Positioning System) is a high-precision radio navigation and positioning system based on artificial Earth satellites. GPS can provide accurate geographical location, vehicle speed, and precise time information anywhere in the world and in near-Earth space. The BeiDou Navigation Satellite System consists of three parts: space segment, ground segment, and user segment. It can provide users with high-precision, high-reliability positioning, navigation, and timing services globally, 24 / 7, and has regional navigation, positioning, and timing capabilities.
[0003] Because terminal devices are equipped with GPS or BeiDou satellite navigation systems, they can be used for positioning when needed. Outdoors, where GPS or BeiDou signals are relatively strong, accurate positioning is possible. However, indoors, weaker GPS or BeiDou signals prevent accurate positioning. For example, in energy industries such as coal, electricity, and petrochemicals, the demand for positioning is increasing, often occurring indoors where signal obstruction and other issues prevent accurate device location. Summary of the Invention
[0004] This application provides a pose display method applied to a cloud-edge management system, the cloud-edge management system including a terminal device and a server, the server including a 3D visual map of the target scene, the method including:
[0005] During the movement of the target scene, the terminal device acquires the target image of the target scene and the motion data of the terminal device, and determines the self-localization trajectory of the terminal device based on the target image and the motion data; if the target image includes multiple frames, a portion of the images are selected from the multiple frames as the test image, and the test image and the self-localization trajectory are sent to the server.
[0006] The server generates a fused positioning trajectory of the terminal device in the 3D visual map based on the image to be tested and the self-localization trajectory, and the fused positioning trajectory includes multiple fused positioning poses;
[0007] For each fused positioning pose in the fused positioning trajectory, the server determines the target positioning pose corresponding to the fused positioning pose and displays the target positioning pose.
[0008] This application provides a cloud-edge management system, which includes terminal devices and a server. The server includes a 3D visual map of the target scene, wherein:
[0009] The terminal device is used to acquire a target image of the target scene and motion data of the terminal device during the movement of the target scene, and determine the self-localization trajectory of the terminal device based on the target image and the motion data; if the target image includes multiple frames, a portion of the images are selected from the multiple frames as the image to be tested, and the image to be tested and the self-localization trajectory are sent to the server.
[0010] The server is configured to generate a fused positioning trajectory of the terminal device in the 3D visual map based on the image to be tested and the self-localization trajectory, the fused positioning trajectory including multiple fused positioning poses; for each fused positioning pose in the fused positioning trajectory, a target positioning pose corresponding to the fused positioning pose is determined and the target positioning pose is displayed.
[0011] This application provides a pose display device applied to a server in a cloud-edge management system. The server includes a 3D visual map of the target scene. The device includes:
[0012] An acquisition module is used to acquire the image to be tested and the self-localization trajectory; wherein, the self-localization trajectory is determined by the terminal device based on the target image of the target scene and the motion data of the terminal device, and the image to be tested is a portion of the multiple frames included in the target image;
[0013] The generation module is used to generate a fused positioning trajectory of the terminal device in the three-dimensional visual map based on the image to be tested and the self-localization trajectory, wherein the fused positioning trajectory includes multiple fused positioning poses;
[0014] The display module is used to determine the target positioning pose corresponding to each fused positioning pose in the fused positioning trajectory, and to display the target positioning pose.
[0015] As can be seen from the above technical solutions, this application proposes a cloud-edge combined positioning and display method. The edge terminal device collects target images and motion data, and performs high-frame-rate self-positioning based on the target images and motion data to obtain a high-frame-rate self-positioning trajectory. The cloud server receives the test image and self-positioning trajectory sent by the terminal device, and obtains a high-frame-rate fused positioning trajectory based on the test image and self-positioning trajectory, i.e., a high-frame-rate fused positioning trajectory in a 3D visual map. This achieves high-frame-rate and high-precision positioning, realizing high-precision, low-cost, and easily deployable indoor positioning. It is a vision-based indoor positioning method that can display the fused positioning trajectory. In this method, the terminal device calculates the high-frame-rate self-positioning trajectory and only sends the self-positioning trajectory and a small amount of test images, reducing the amount of data transmitted over the network. Global positioning is performed on the server, thereby reducing the computational and storage resource consumption of the terminal device. This method can be applied in energy industries such as coal, electricity, and petrochemicals to achieve indoor positioning of personnel (such as workers and inspection personnel), quickly obtaining personnel location information and ensuring personnel safety. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating a pose display method in one embodiment of this application;
[0017] Figure 2 This is a schematic diagram of the cloud-edge management system in one embodiment of this application;
[0018] Figure 3 This is a flowchart illustrating the process of determining a self-localization trajectory in one embodiment of this application;
[0019] Figure 4 This is a flowchart illustrating the determination of a global positioning trajectory in one embodiment of this application;
[0020] Figure 5 This is a schematic diagram of self-localization trajectory, global localization trajectory, and fused localization trajectory;
[0021] Figure 6 This is a flowchart illustrating the process of determining the fused positioning trajectory in one embodiment of this application;
[0022] Figure 7 This is a schematic diagram of the pose display device in one embodiment of this application. Detailed Implementation
[0023] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the application. The singular forms “a,” “the,” and “the” as used in this application and claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to any and all possible combinations comprising one or more of the associated listed items.
[0024] It should be understood that although the terms first, second, third, etc., may be used to describe various information in embodiments of this application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" may also be interpreted as "when," "when," or "in response to a determination."
[0025] This application proposes a pose display method, which can be applied to a cloud-edge management system. The cloud-edge management system may include terminal devices (i.e., edge-end terminal devices) and servers (i.e., cloud-based servers). The server may include a 3D visual map of the target scene (e.g., indoor environment, outdoor environment, etc.). See [link to relevant documentation]. Figure 1 The diagram shown is a flowchart of the pose display method, which may include:
[0026] Step 101: During the movement of the terminal device in the target scene, the terminal device acquires the target image of the target scene and the motion data of the terminal device, and determines the self-localization trajectory of the terminal device based on the target image and motion data.
[0027] For example, if the target image includes multiple frames, the terminal device traverses the multiple frames to find the current frame; determines the self-localization pose corresponding to the current frame based on the self-localization pose corresponding to the K frames preceding the current frame, the map position of the terminal device in the self-localization coordinate system, and the motion data; and generates the self-localization trajectory of the terminal device in the self-localization coordinate system based on the self-localization poses corresponding to the multiple frames.
[0028] For example, if the current frame image is a key image, the map position in the self-localization coordinate system can be generated based on the current position of the terminal device (i.e., the position corresponding to the current frame image). If the current frame image is not a key image, it is not necessary to generate the map position in the self-localization coordinate system based on the current position of the terminal device.
[0029] If the number of matching feature points between the current frame and the previous frame does not reach a preset threshold, the current frame is determined to be a key image. If the number of matching feature points between the current frame and the previous frame reaches a preset threshold, the current frame is determined to be a non-key image.
[0030] Step 102: If the target image includes multiple frames, the terminal device selects a portion of the images from the multiple frames as the image to be tested, and sends the image to be tested and the self-localization trajectory to the server.
[0031] For example, a terminal device can select M frames from multiple frames as the image to be tested, where M can be a positive integer, such as 1, 2, 3, etc. Clearly, the terminal device sends a portion of the image to be tested from the multiple frames to the server, thereby reducing the amount of data transmitted over the network and saving network bandwidth resources.
[0032] Step 103: The server generates a fused positioning trajectory of the terminal device in the 3D visual map based on the image to be tested and the self-localization trajectory. The fused positioning trajectory may include multiple fused positioning poses.
[0033] For example, the server can determine the target map points corresponding to the image under test from the 3D visual map of the target scene, and determine the global positioning trajectory of the terminal device in the 3D visual map based on the target map points. Then, the server generates the fused positioning trajectory of the terminal device in the 3D visual map based on the self-localization trajectory and the global positioning trajectory. For example, the frame rate of the fused positioning poses included in the fused positioning trajectory can be greater than the frame rate of the global positioning poses included in the global positioning trajectory, that is, the frame rate of the fused positioning trajectory is higher than the frame rate of the global positioning trajectory. The fused positioning trajectory can be a high frame rate pose in the 3D visual map, and the global positioning trajectory can be a low frame rate pose in the 3D visual map. The higher frame rate of the fused positioning trajectory indicates that the number of fused positioning poses is greater than the number of global positioning poses. In addition, the frame rate of the fused positioning poses included in the fused positioning trajectory can be equal to the frame rate of the self-localization poses included in the self-localization trajectory, that is, the frame rate of the fused positioning trajectory is equal to the frame rate of the self-localization trajectory, that is, the self-localization trajectory can be a high frame rate pose. The frame rate of the fused localization trajectory is equal to the frame rate of the self-localization trajectory, indicating that the number of fused localization poses is equal to the number of self-localization poses.
[0034] In one possible implementation, the 3D visual map may include, but is not limited to, at least one of the following: a pose matrix corresponding to a sample image, a global descriptor corresponding to a sample image, a local descriptor corresponding to a feature point in the sample image, and map point information. Specifically, the server determines target map points corresponding to the image under test from the 3D visual map of the target scene, and determines the global positioning trajectory of the terminal device in the 3D visual map based on the target map points. This may include, but is not limited to: for each frame of the image under test, selecting candidate sample images from multiple sample images based on the similarity between the image under test and multiple frames of sample images corresponding to the 3D visual map; obtaining multiple feature points from the image under test; for each feature point, determining the target map point corresponding to the feature point from multiple map points corresponding to the candidate sample images; determining the global positioning pose in the 3D visual map corresponding to the image under test based on the multiple feature points and the target map points corresponding to the multiple feature points; and generating the global positioning trajectory of the terminal device in the 3D visual map based on the global positioning poses corresponding to all images under test.
[0035] The server selects candidate sample images from the multiple sample images based on the similarity between the image to be tested and the corresponding multi-frame sample images of the 3D visual map. This can include: determining the global descriptor to be tested corresponding to the image to be tested, and determining the distance between the global descriptor to be tested and the sample global descriptor corresponding to each frame of the 3D visual map; wherein the 3D visual map includes at least the sample global descriptor corresponding to each frame of the sample map. Based on the distance between the global descriptor to be tested and each sample global descriptor, candidate sample images are selected from the multiple sample images; wherein the distance between the global descriptor to be tested and the sample global descriptor corresponding to the candidate sample image is the minimum distance; or, the distance between the global descriptor to be tested and the sample global descriptor corresponding to the candidate sample image is less than a distance threshold.
[0036] The server determines the global descriptor corresponding to the image under test, which may include, but is not limited to: determining the bag-of-words vector corresponding to the image under test based on a trained dictionary model, and using the bag-of-words vector as the global descriptor corresponding to the image under test; or, inputting the image under test into a trained deep learning model to obtain the target vector corresponding to the image under test, and using the target vector as the global descriptor corresponding to the image under test. Of course, the above are just examples of determining the global descriptor under test, and there are no limitations on this.
[0037] The server determines the target map point corresponding to the feature point from multiple map points corresponding to the candidate sample image. This may include, but is not limited to: determining a test local descriptor corresponding to the feature point, where the test local descriptor represents the feature vector of the image patch where the feature point is located, and the image patch may be located in the test image; determining the distance between the test local descriptor and the sample local descriptor corresponding to each map point corresponding to the candidate sample image; wherein the 3D visual map includes at least the sample local descriptor corresponding to each map point corresponding to the candidate sample image. Then, based on the distance between the test local descriptor and each sample local descriptor, the target map point can be selected from the multiple map points corresponding to the candidate sample image; wherein the distance between the test local descriptor and the sample local descriptor corresponding to the target map point can be the minimum distance, and the minimum distance is less than a distance threshold.
[0038] The server generates a fused positioning trajectory for the terminal device in the 3D visual map based on the self-localization trajectory and the global positioning trajectory. This fusion trajectory may include, but is not limited to: selecting N self-localization poses corresponding to the target time period from all self-localization poses included in the self-localization trajectory, and selecting P global positioning poses corresponding to the target time period from all global positioning poses included in the global positioning trajectory; where N is greater than P. Based on the N self-localization poses and P global positioning poses, N fused positioning poses are determined, with a one-to-one correspondence between the N self-localization poses and the N fused positioning poses. The fused positioning trajectory for the terminal device in the 3D visual map is then generated based on the N fused positioning poses.
[0039] After generating the fused positioning trajectory of the terminal device in the 3D visual map based on the self-localization trajectory and the global positioning trajectory, the server can further select an initial fused positioning pose from the fused positioning trajectory and an initial self-localization pose corresponding to the initial fused positioning pose from the self-localization trajectory. A target self-localization pose is then selected from the self-localization trajectory, and a target fused positioning pose is determined based on the initial fused positioning pose, the initial self-localization pose, and the target self-localization pose. Finally, a new fused positioning trajectory is generated based on the target fused positioning pose and the fused positioning trajectory to replace the original fused positioning trajectory.
[0040] Step 104: For each fused positioning pose in the fused positioning trajectory, the server determines the target positioning pose corresponding to the fused positioning pose and displays the target positioning pose.
[0041] For example, the server can determine the fused localization pose as the target localization pose and display it on a 3D visual map. Alternatively, the server can convert the fused localization pose into a target localization pose on a 3D visual map based on the target transformation matrix between the 3D visual map and the 3D visualization map, and then display the target localization pose on the 3D visualization map.
[0042] For example, the method for determining the target transformation matrix between a 3D visual map and a 3D visualized map may include, but is not limited to: for each of multiple calibration points, determining the coordinate pair corresponding to that calibration point, which may include the position coordinates of the calibration point in the 3D visual map and the position coordinates of the calibration point in the 3D visualized map; and determining the target transformation matrix based on the coordinate pairs corresponding to the multiple calibration points. Alternatively, obtaining an initial transformation matrix, mapping the position coordinates in the 3D visual map to the mapped coordinates in the 3D visualized map based on the initial transformation matrix, and determining whether the initial transformation matrix has converged based on the relationship between the mapped coordinates and the actual coordinates in the 3D visualized map; if yes, then determining the initial transformation matrix as the target transformation matrix; if not, then adjusting the initial transformation matrix, using the adjusted transformation matrix as the initial transformation matrix, and returning to perform the operation of mapping the position coordinates in the 3D visual map to the mapped coordinates in the 3D visualized map based on the initial transformation matrix, and so on, until the target transformation matrix is obtained. Alternatively, the 3D visualization map is sampled to obtain a first point cloud corresponding to the 3D visualization map; and the 3D visual map is sampled to obtain a second point cloud corresponding to the 3D visual map; the ICP algorithm is used to register the first point cloud and the second point cloud to obtain the target transformation matrix between the 3D visual map and the 3D visualization map.
[0043] As can be seen from the above technical solutions, this application proposes a cloud-edge combined positioning and display method. The edge terminal device collects target images and motion data, and performs high-frame-rate self-positioning based on the target images and motion data to obtain a high-frame-rate self-positioning trajectory. The cloud server receives the test image and self-positioning trajectory sent by the terminal device, and obtains a high-frame-rate fused positioning trajectory based on the test image and self-positioning trajectory, i.e., a high-frame-rate fused positioning trajectory in a 3D visual map. This achieves high-frame-rate and high-precision positioning, realizing high-precision, low-cost, and easily deployable indoor positioning. It is a vision-based indoor positioning method and can display the fused positioning trajectory in a 3D visualization map. In the above method, the terminal device calculates the high-frame-rate self-positioning trajectory and only sends the self-positioning trajectory and a small amount of test images, reducing the amount of data transmitted over the network. Global positioning is performed on the server, thereby reducing the computational and storage resource consumption of the terminal device. This method can be applied in energy industries such as coal, electricity, and petrochemicals to achieve indoor positioning of personnel (such as workers, inspectors, etc.), quickly obtaining personnel location information and ensuring personnel safety.
[0044] The pose display method of this application will be described below with reference to specific embodiments.
[0045] This application proposes a cloud-edge combined visual positioning and display method. During the movement of a terminal device within a target scene, the server determines the fused positioning trajectory of the terminal device in a 3D visual map and displays the fused positioning trajectory. The target scene can be an indoor environment; that is, when the terminal device moves within an indoor environment, the server determines the fused positioning trajectory of the terminal device in the 3D visual map, thus proposing a vision-based indoor positioning method. Of course, the target scene can also be an outdoor environment; there is no limitation on this.
[0046] See Figure 2 The diagram illustrates the structure of a cloud-edge management system. This system can include terminal devices (i.e., edge devices) and servers (i.e., cloud servers). Of course, it can also include other devices, such as wireless base stations and routers, without limitation. The server can include a 3D visual map of the target scene and a corresponding 3D visualization map. The server can generate a fused positioning trajectory for the terminal device within this 3D visual map and display this trajectory (which needs to be converted to a format suitable for display on the 3D visualization map) on the 3D visualization map. This allows administrators to view the fused positioning trajectory on the 3D visualization map via a web interface.
[0047] Terminal devices can include visual sensors and motion sensors. Visual sensors, such as cameras, are used to acquire images of the target scene as the terminal device moves. For ease of distinction, this image is designated as the target image, which includes multiple frames (i.e., multiple real-time images during the terminal device's movement). Motion sensors, such as IMUs (Inertial Measurement Units), are measurement devices containing gyroscopes and accelerometers. Motion sensors are used to acquire motion data of the terminal device, such as acceleration and angular velocity, as it moves.
[0048] For example, the terminal device can be a wearable device (such as a video safety helmet, smartwatch, smart glasses, etc.), with the visual sensor and motion sensor deployed on the wearable device; or, the terminal device can be a recorder (such as a device carried by staff during work, integrating real-time audio and video acquisition, photography, recording, intercom, positioning, etc.), with the visual sensor and motion sensor deployed on the recorder; or, the terminal device can be a camera (such as a split-type camera), with the visual sensor and motion sensor deployed on the camera. Of course, the above are just examples, and the type of terminal device is not limited; it can also be a smartphone, as long as the visual sensor and motion sensor are deployed.
[0049] For example, the terminal device can acquire target images and motion data, and perform high frame rate self-localization based on the target images and motion data to obtain a high frame rate self-localization trajectory (such as a 6DOF (six degrees of freedom) self-localization trajectory). The self-localization trajectory can include multiple self-localization poses. Since the self-localization trajectory is a high frame rate self-localization trajectory, the number of self-localization poses in the self-localization trajectory is relatively large.
[0050] The terminal device can select a portion of images from multiple frames of the target image as the test image, and send the high-frame-rate self-localization trajectory and the test image to the server. The server can obtain the self-localization trajectory and the test image, and perform low-frame-rate global localization based on the test image and the 3D visual map of the target scene to obtain a low-frame-rate global localization trajectory (i.e., the global localization trajectory of the test image in the 3D visual map). The global localization trajectory can include multiple global localization poses. Since the global localization trajectory is a low-frame-rate global localization trajectory, the number of global localization poses in the global localization trajectory is relatively small.
[0051] Based on a high-frame-rate self-localization trajectory and a low-frame-rate global localization trajectory, the server can fuse the high-frame-rate self-localization trajectory and the low-frame-rate global localization trajectory to obtain a high-frame-rate fused localization trajectory, which is the high-frame-rate fused localization trajectory in the 3D visual map, thus obtaining a high-frame-rate global localization result. The fused localization trajectory can include multiple fused localization poses. Since the fused localization trajectory is a high-frame-rate fused localization trajectory, the number of fused localization poses in the fused localization trajectory is relatively large.
[0052] In the above embodiments, the pose (such as self-localization pose, global localization pose, fused localization pose, etc.) can be position and orientation, and is generally represented by rotation matrix and translation vector, without limitation.
[0053] In summary, this embodiment achieves a globally unified high frame rate visual positioning function based on target images and motion data, obtaining a high frame rate fused positioning trajectory (such as 6DOF pose) in a 3D visual map. This is a high frame rate globally consistent positioning method that enables high frame rate, high precision, low cost, and easy deployment of indoor positioning functions for terminal devices, achieving a globally consistent high frame rate positioning function indoors.
[0054] The above process of the embodiments of this application will be described in detail below with reference to specific application scenarios.
[0055] I. Self-localization of terminal devices. Terminal devices are electronic devices equipped with visual sensors and motion sensors. They can acquire target images of the target scene (such as continuous video images) and motion data of the terminal device (such as IMU data), and determine the self-localization trajectory of the terminal device based on the target images and motion data.
[0056] The target image may include multiple frames. For each frame, the terminal device determines the self-localization pose corresponding to that image. That is, multiple frames correspond to multiple self-localization poses. The self-localization trajectory of the terminal device may include multiple self-localization poses. It can be understood that the self-localization trajectory is a set of multiple self-localization poses.
[0057] For the first frame of a multi-frame image, the terminal device determines the self-localization pose corresponding to the first frame. For the second frame, the terminal device determines the self-localization pose corresponding to the second frame, and so on. The self-localization pose corresponding to the first frame can be the origin of the reference coordinate system (i.e., the self-localization coordinate system). The self-localization pose corresponding to the second frame is a pose point in the reference coordinate system, that is, a pose point relative to the origin (i.e., the self-localization pose corresponding to the first frame). The self-localization pose corresponding to the third frame is a pose point in the reference coordinate system, that is, a pose point relative to the origin, and so on. The self-localization pose corresponding to each frame is a pose point in the reference coordinate system.
[0058] In summary, after obtaining the self-localization pose corresponding to each frame of the image, these self-localization poses can be combined into a self-localization trajectory in the reference coordinate system, which includes these self-localization poses.
[0059] In one possible implementation, see Figure 3 As shown, the self-localization trajectory is determined using the following steps:
[0060] Step 301: Obtain the target image of the target scene and the motion data of the terminal device.
[0061] Step 302: If the target image includes multiple frames, then traverse the multiple frames to find the current frame image.
[0062] When selecting the first frame from multiple frames as the current frame, the self-localization pose corresponding to the first frame can be the origin of the reference coordinate system (i.e., the self-localization coordinate system), meaning the self-localization pose coincides with this origin. When selecting the second frame from multiple frames as the current frame, subsequent steps can be used to determine the self-localization pose corresponding to the second frame. When selecting the third frame from multiple frames as the current frame, subsequent steps can be used to determine the self-localization pose corresponding to the third frame, and so on, allowing each frame to be selected as the current frame.
[0063] Step 303: Calculate the feature point correlation between the current frame image and the previous frame image using the optical flow algorithm. The optical flow algorithm utilizes the temporal changes of pixels in the current frame image and the correlation between the previous frame image to find the correspondence between the current and previous frames, thereby calculating the motion information of objects between the current and previous frames.
[0064] Step 304: Determine whether the current frame image is a key image based on the number of matching feature points between the current frame image and the previous frame image. For example, if the number of matching feature points between the current frame image and the previous frame image does not reach a preset threshold, it indicates that the change between the current frame image and the previous frame image is large, resulting in a relatively small number of matching feature points between the two frames. In this case, the current frame image is determined to be a key image, and step 305 is executed. If the number of matching feature points between the current frame image and the previous frame image reaches a preset threshold, it indicates that the change between the current frame image and the previous frame image is small, resulting in a relatively large number of matching feature points between the two frames. In this case, the current frame image is determined to be a non-key image, and step 306 is executed.
[0065] For example, the matching ratio between the current frame image and the previous frame image can be calculated based on the number of matching feature points between the current frame image and the previous frame image. For instance, it can be the ratio of the number of matching feature points to the total number of feature points. If the matching ratio does not reach a preset ratio, the current frame image is determined to be a key image; if the matching ratio reaches the preset ratio, the current frame image is determined to be a non-key image.
[0066] Step 305: If the current frame image is a critical image, then generate a map position in the self-localization coordinate system (i.e., the reference coordinate system) based on the current position of the terminal device (i.e., the position the terminal device was in when acquiring the current frame image), thus generating a new 3D map position. If the current frame image is not a critical image, then it is not necessary to generate a map position in the self-localization coordinate system based on the current position of the terminal device.
[0067] Step 306: Based on the self-localization poses of the K preceding frames, the map position of the terminal device in the self-localization coordinate system, and the motion data of the terminal device, determine the self-localization pose corresponding to the current frame. K can be a positive integer or a value configured based on experience, without any restrictions.
[0068] For example, pre-integration can be performed on all motion data between the current frame and the previous frame to obtain the inertial measurement constraints between these two frames. Based on the self-localization pose and motion data (e.g., velocity, acceleration, angular velocity, etc.) of the K preceding frames (e.g., sliding window), the map position in the self-localization coordinate system, and the inertial measurement constraints (velocity, acceleration, angular velocity, etc. between the previous and current frames), bundle optimization can be used for joint optimization and update to obtain the self-localization pose corresponding to the current frame. There are no restrictions on this bundle optimization process.
[0069] For example, in order to maintain the size of the variables to be optimized, a frame and part of the map location within the sliding window can be marginalized, and this constraint information can be preserved in a priori form.
[0070] For example, the terminal device can use the VIO (Visual Inertial Odometry) algorithm to determine its self-localization pose. That is, the input data for the VIO algorithm is the target image and motion data, and the output data is the self-localization pose. For instance, based on the target image and motion data, the VIO algorithm can obtain the self-localization pose. For example, by executing steps 301-306 using the VIO algorithm, the self-localization pose can be obtained. This VIO algorithm can include, but is not limited to, VINS (Visual Inertial Navigation Systems), SVO (Semi-direct Visual Odometry), MSCKF (Multi-State Constraint Kalman Filter), etc., as long as it can obtain the self-localization pose.
[0071] Step 307: Generate the self-localization trajectory of the terminal device in the self-localization coordinate system based on the self-localization poses corresponding to multiple frames of images. The self-localization trajectory includes multiple self-localization poses in the self-localization coordinate system.
[0072] At this point, the terminal device can obtain the self-localization trajectory in the self-localization coordinate system. This self-localization trajectory can include the self-localization poses corresponding to multiple frames of images. Obviously, since the visual sensor can collect a large number of images, the terminal device can obtain the self-localization poses corresponding to these images. That is, the self-localization trajectory can include a large number of self-localization poses. In other words, the terminal device can obtain a high frame rate self-localization trajectory.
[0073] II. Data Transmission. If the target image consists of multiple frames, the terminal device can select a portion of the images from the multiple frames as the image to be tested, and send the image to be tested and the self-localization trajectory to the server. For example, the terminal device can send the self-localization trajectory and the image to be tested to the server via a wireless network (such as 4G, 5G, Wi-Fi, etc.). Since the frame rate of the image to be tested is low, the network bandwidth occupied is small.
[0074] III. 3D Visual Map of the Target Scene. A 3D visual map of the target scene needs to be pre-constructed and stored on a server. This allows the server to perform global localization based on the 3D visual map. The 3D visual map is a method of storing image information of the target scene. It can be constructed by acquiring multiple frames of sample images of the target scene and building a 3D visual map based on these images. For example, based on multiple frames of sample images of the target scene, visual mapping algorithms such as SFM (Structure From Motion) or SLAM (Simultaneous Localization and Mapping) can be used to construct a 3D visual map of the target scene. There are no restrictions on the construction method.
[0075] After obtaining a 3D visual map of the target scene, the 3D visual map can include the following information:
[0076] Sample image pose: A sample image is a representative image used to construct a 3D visual map. In other words, a 3D visual map can be constructed based on sample images. The pose matrix of the sample images (which can be simply referred to as the sample image pose) can be stored in the 3D visual map. The 3D visual map can include the sample image pose.
[0077] Sample Global Descriptor: For each frame of sample image, there is a corresponding image global descriptor, which is denoted as the sample global descriptor. The sample global descriptor is a high-dimensional vector representing the sample image and is used to distinguish the image features of different sample images.
[0078] Specifically, for each frame of a sample image, a bag-of-words vector corresponding to that sample image can be determined based on a trained dictionary model, and this bag-of-words vector can be used as the global descriptor for that sample image. For example, the visual bag-of-words method is one way to determine global descriptors. In the visual bag-of-words method, a bag-of-words vector can be constructed, which is a vector representation method used for image similarity detection. This bag-of-words vector can be used as the global descriptor for the sample image.
[0079] In the bag-of-words method, a "dictionary", also known as a dictionary model, needs to be trained in advance. It is generally trained by clustering feature point descriptors in a large number of images to obtain a classification tree. Each classification tree can represent a visual "word", and these visual "words" form the dictionary model.
[0080] For a sample image, all feature point descriptors in the sample image can be classified as "words" and the frequency of each word can be counted. In this way, the frequency of each word in the dictionary can form a vector, which is the bag-of-words vector corresponding to the sample image. This bag-of-words vector can be used to measure the similarity between two images and is used as the global descriptor of the sample image.
[0081] For each sample image, the image can be input into a trained deep learning model to obtain a target vector corresponding to that image. This target vector is then used as the global descriptor for that sample image. For example, deep learning is one method for determining global descriptors. In deep learning, a deep learning model can perform multiple convolutions on the sample image to obtain a high-dimensional target vector, which is then used as the global descriptor for the sample image.
[0082] In deep learning methods, it is necessary to pre-train deep learning models, such as CNN (Convolutional Neural Networks) models. These models are typically trained using a large number of images, and there are no restrictions on the training method. For a sample image, the sample image can be input into the deep learning model, which processes it to obtain a high-dimensional target vector. This target vector is then used as the global descriptor for that sample image.
[0083] The sample local descriptor corresponding to the feature point of the sample image: For each frame of the sample image, the sample image may include multiple feature points. A feature point can be a specific pixel location in the sample image, and each feature point can correspond to an image local descriptor, denoted as the sample local descriptor. The sample local descriptor is a vector that describes the features of an image patch within the vicinity of the feature point (i.e., the pixel location). This vector can also be called the descriptor of the feature point. In summary, the sample local descriptor is a feature vector used to represent the image patch where the feature point is located, and this image patch can be located within the sample image. It is important to note that for a feature point (i.e., a two-dimensional feature point) in the sample image, this feature point can correspond to a map point in a three-dimensional visual map (i.e., a three-dimensional map point). Therefore, the sample local descriptor corresponding to a feature point can also be the sample local descriptor corresponding to the map point of that feature point.
[0084] Algorithms such as ORB (Oriented Fast and Rotated BRIEF), SIFT (Scale-Invariant Feature Transform), and SURF (Speeded Up Robust Features) can be used to extract feature points from sample images and determine the corresponding local descriptors. Alternatively, deep learning algorithms (such as SuperPoint, DELF, and D2-Net) can be used to extract feature points from sample images and determine the corresponding local descriptors. There are no restrictions on this, as long as the feature points can be obtained and the local descriptors can be determined.
[0085] Map point information: Map point information may include, but is not limited to: the 3D spatial location of the map point, all observed sample images, and the corresponding 2D feature point (i.e., the feature point corresponding to the map point) number.
[0086] IV. Global Positioning of the Server. Based on the acquired 3D visual map of the target scene, after obtaining the image to be tested, the server determines the target map point corresponding to the image to be tested from the 3D visual map of the target scene, and determines the global positioning trajectory of the terminal device in the 3D visual map based on the target map point.
[0087] For each frame of the image to be tested, the server can determine the global localization pose corresponding to that frame. Assuming there are M frames of images to be tested, then these M frames correspond to M global localization poses. The global localization trajectory of the terminal device in the 3D visual map can include these M global localization poses; it can be understood that the global localization trajectory is a set of these M global localization poses. For the first frame of the M frames, the global localization pose corresponding to the first frame is determined; for the second frame, the corresponding global localization pose is determined, and so on. Each global localization pose is a pose point in the 3D visual map, i.e., a pose point in the 3D visual map coordinate system. In summary, after obtaining the global localization poses corresponding to the M frames of images to be tested, these global localization poses are combined to form the global localization trajectory in the 3D visual map, and this global localization trajectory includes these global localization poses.
[0088] A 3D visual map based on the target scene, in one possible implementation, see [link to relevant documentation]. Figure 4 As shown, the server can determine the global positioning trajectory of the terminal device in the 3D visual map using the following steps:
[0089] Step 401: The server obtains the image of the target scene from the terminal device.
[0090] For example, a terminal device can acquire a target image, which includes multiple frames. The terminal device can select M frames from the multiple frames as the test images and send the M test images to the server. For instance, the multiple frames may include key images and non-key images. Based on this, the terminal device can select key images from the multiple frames as the test images, while excluding non-key images. Alternatively, the terminal device can select the test images from the multiple frames at fixed intervals. Assuming the fixed interval is 5 (of course, the fixed interval can be arbitrarily configured based on experience, and there are no restrictions), then the 1st frame can be selected as the test image, the 6th (1+5)th frame as the test image, the 11th (6+5)th frame as the test image, and so on, selecting one test image every 5 frames.
[0091] Step 402: For each frame of the image to be tested, determine the global descriptor to be tested corresponding to that image.
[0092] For example, for each frame of the image to be tested, the image to be tested can correspond to an image global descriptor, which can be denoted as the global descriptor to be tested. The global descriptor to be tested is a high-dimensional vector representing the image to be tested. The global descriptor to be tested is used to distinguish the image features of different images to be tested.
[0093] Specifically, for each frame of the image to be tested, a bag-of-words vector corresponding to the image is determined based on a trained dictionary model, and this bag-of-words vector is used as the global descriptor to be tested for that image. Alternatively, for each frame of the image to be tested, the image is input into a trained deep learning model to obtain a target vector corresponding to the image, and this target vector is used as the global descriptor to be tested for that image.
[0094] In summary, the global descriptor corresponding to the image under test can be determined based on the bag-of-words method or deep learning method. The determination method is the same as that for the determination of the global descriptor of the sample, and will not be repeated here.
[0095] Step 403: For each frame of the image to be tested, determine the similarity between the global descriptor of the image to be tested and the global descriptor of each frame of the sample image corresponding to the 3D visual map.
[0096] Referring to the above embodiments, a three-dimensional visual map may include a global descriptor corresponding to each frame of sample image. Therefore, the similarity between the global descriptor to be tested and each sample global descriptor can be determined. Taking the similarity as "distance similarity" as an example, the distance between the global descriptor to be tested and each sample global descriptor can be determined, such as Euclidean distance, that is, the Euclidean distance between two feature vectors is calculated.
[0097] Step 404: Based on the distance between the global descriptor to be tested and each sample global descriptor, select candidate sample images from the multi-frame sample images corresponding to the 3D visual map; wherein, the distance between the global descriptor to be tested and the sample global descriptor corresponding to the candidate sample image is the minimum distance; or, the distance between the global descriptor to be tested and the sample global descriptor corresponding to the candidate sample image is less than the distance threshold.
[0098] For example, assuming that the 3D visual map corresponds to sample image 1, sample image 2 and sample image 3, then the distance 1 between the global descriptor to be tested and the global descriptor corresponding to sample image 1 can be calculated, the distance 2 between the global descriptor to be tested and the global descriptor corresponding to sample image 2 can be calculated, and the distance 3 between the global descriptor to be tested and the global descriptor corresponding to sample image 3 can be calculated.
[0099] In one possible implementation, if distance 1 is the minimum distance, then sample image 1 is selected as a candidate sample image. Alternatively, if distance 1 is less than a distance threshold (which can be configured empirically), and distance 2 is less than a distance threshold, but distance 3 is not less than a distance threshold, then both sample image 1 and sample image 2 are selected as candidate sample images. Alternatively, if distance 1 is the minimum distance and distance 1 is less than a distance threshold, then sample image 1 is selected as a candidate sample image; however, if distance 1 is the minimum distance and distance 1 is not less than a distance threshold, then no candidate sample image can be selected, i.e., relocalization fails.
[0100] In summary, for each frame of the image to be tested, candidate sample images corresponding to the image to be tested can be selected from multiple sample images corresponding to the 3D visual map, and the number of candidate sample images is at least one.
[0101] Step 405: For each frame of the image to be tested, obtain multiple feature points from the image to be tested. For each feature point, determine the local descriptor to be tested corresponding to the feature point. The local descriptor to be tested is used to represent the feature vector of the image block where the feature point is located, and the image block can be located in the image to be tested.
[0102] For example, the image to be tested may include multiple feature points. A feature point can be a specific pixel location within the image, and each feature point can correspond to a local image descriptor. This local image descriptor is denoted as the local descriptor to be tested. The local descriptor to be tested is a vector that describes the features of an image patch within the vicinity of the feature point (i.e., the pixel location). This vector can also be called the descriptor of the feature point. In summary, the local descriptor to be tested is a feature vector used to represent the image patch where the feature point is located.
[0103] Algorithms such as ORB, SIFT, and SURF can be used to extract feature points from the image under test and determine the corresponding local descriptors. Alternatively, deep learning algorithms (such as SuperPoint, DELF, and D2-Net) can be used to extract feature points from the image under test and determine the corresponding local descriptors. There are no restrictions on this, as long as the feature points can be obtained and the local descriptors can be determined.
[0104] Step 406: For each feature point in the image to be tested, determine the distance between the local descriptor of the feature point and the local descriptor of each map point in the candidate sample image (i.e., the local descriptor of each feature point in the candidate sample image), such as Euclidean distance, i.e., calculate the Euclidean distance between two feature vectors.
[0105] Referring to the above embodiments, for each frame of sample image, the 3D visual map includes a sample local descriptor corresponding to each map point of the sample image. Therefore, after obtaining the candidate sample image corresponding to the image to be tested, the sample local descriptor corresponding to each map point of the candidate sample image is obtained from the 3D visual map. After obtaining each feature point corresponding to the image to be tested, the distance between the test local descriptor corresponding to the feature point and the sample local descriptor corresponding to each map point of the candidate sample image is determined.
[0106] Step 407: For each feature point, based on the distance between the local descriptor to be tested corresponding to the feature point and the sample local descriptor corresponding to each map point in the candidate sample image, select a target map point from multiple map points corresponding to the candidate sample image; wherein, the distance between the local descriptor to be tested and the sample local descriptor corresponding to the target map point is the minimum distance, and the minimum distance is less than the distance threshold.
[0107] For example, assuming the candidate sample image corresponds to map point 1, map point 2, and map point 3, we can calculate the distance 1 between the local descriptor to be tested corresponding to the feature point and the sample local descriptor corresponding to map point 1, the distance 2 between the local descriptor to be tested and the sample local descriptor corresponding to map point 2, and the distance 3 between the local descriptor to be tested and the sample local descriptor corresponding to map point 3.
[0108] In one possible implementation, if distance 1 is the minimum distance, then map point 1 can be selected as the target map point. Alternatively, if distance 1 is less than a distance threshold (which can be configured empirically), and distance 2 is less than the distance threshold, but distance 3 is not less than the distance threshold, then both map point 1 and map point 2 can be selected as target map points. Alternatively, if distance 1 is the minimum distance and distance 1 is less than the distance threshold, then map point 1 can be selected as the target map point; however, if distance 1 is the minimum distance and distance 1 is not less than the distance threshold, then a target map point cannot be selected, i.e., relocation fails.
[0109] In summary, for each feature point in the image to be tested, a target map point corresponding to that feature point is selected from the candidate sample image corresponding to the image to be tested, thus obtaining the matching relationship between the feature point and the target map point.
[0110] Step 408: Based on multiple feature points corresponding to the image under test and the target map points corresponding to the multiple feature points, determine the global localization pose in the 3D visual map corresponding to the image under test.
[0111] For a single frame of the image to be tested, the image can correspond to multiple feature points, and each feature point corresponds to a target map point. For example, the target map point corresponding to feature point 1 is map point 1, the target map point corresponding to feature point 2 is map point 2, and so on, thus obtaining multiple matching pairs. Each matching pair includes a feature point (i.e., a two-dimensional feature point) and a map point (i.e., a three-dimensional map point in the three-dimensional visual map). The feature point represents the two-dimensional position in the image to be tested, and the map point represents the three-dimensional position in the three-dimensional visual map. That is, the matching pair includes the mapping relationship from two-dimensional position to three-dimensional position, i.e., the mapping relationship from two-dimensional position in the image to three-dimensional position in the three-dimensional visual map.
[0112] If the total number of matching pairs does not meet the requirement, it means that the global localization pose in the 3D visual map corresponding to the image under test cannot be determined based on the matching pairs. If the total number of matching pairs meets the requirement (i.e., the total number reaches the preset value), it means that the global localization pose in the 3D visual map corresponding to the image under test can be determined based on the matching pairs.
[0113] For example, the PnP (Perspective NPoint) algorithm can be used to calculate the global localization pose of the image under test in a 3D visual map, and there are no restrictions on the calculation method. For instance, the input data of the PnP algorithm consists of multiple matching pairs. For each matching pair, the matching pair includes the 2D position in the image under test and the 3D position in the 3D visual map. Based on multiple matching pairs, the PnP algorithm can be used to calculate the pose of the image under test in the 3D visual map, i.e., the global localization pose.
[0114] In summary, for each frame of the image to be tested, the global localization pose in the corresponding 3D visual map is obtained, that is, the global localization pose of the image to be tested in the 3D visual map coordinate system is obtained.
[0115] In one possible implementation, after obtaining multiple matching pairs, valid matching pairs can be found from among them. Based on these valid matching pairs, the global localization pose of the image under test in the 3D visual map can be calculated using the PnP algorithm. For example, the RANSAC (RANdom Sample Consensus) detection algorithm can be used to find valid matching pairs from all matching pairs; this process is not limited.
[0116] Step 409: Generate the global positioning trajectory of the terminal device in the 3D visual map based on the global positioning poses corresponding to the M frames of images to be tested. This global positioning trajectory includes multiple global positioning poses in the 3D visual map. At this point, the server can obtain the global positioning trajectory in the 3D visual map, that is, the global positioning trajectory in the 3D visual map coordinate system. This global positioning trajectory can include the global positioning poses corresponding to the M frames of images to be tested, i.e., the global positioning trajectory can include M global positioning poses. Since the M frames of images to be tested are selected from all images, the global positioning trajectory can include the global positioning poses corresponding to a small number of images to be tested; that is, the server can obtain a low frame rate global positioning trajectory.
[0117] V. Server-side Fusion Localization. After obtaining the high-frame-rate self-localization trajectory and the low-frame-rate global localization trajectory, the server fuses the high-frame-rate self-localization trajectory with the low-frame-rate global localization trajectory to obtain the high-frame-rate fused localization trajectory in the 3D visual map coordinate system, which is the fused localization trajectory of the terminal device in the 3D visual map. The fused localization trajectory is the high-frame-rate pose in the 3D visual map, while the global localization trajectory is the low-frame-rate pose in the 3D visual map. That is, the frame rate of the fused localization trajectory is higher than that of the global localization trajectory, and the number of fused localization poses is greater than the number of global localization poses.
[0118] See Figure 5As shown, the white solid circle represents the self-localization pose. The trajectory composed of multiple self-localization poses is called the self-localization trajectory, meaning the self-localization trajectory includes multiple self-localization poses. The self-localization pose corresponding to the first frame image can be the reference coordinate system S. L The origin of the self-localization coordinate system (i.e., the self-localization pose corresponding to the first frame image is denoted as...). Self-positioning pose With reference coordinate system S L The origins of the coordinate systems coincide. For each self-localization pose in the self-localization trajectory, it is in the reference coordinate system S. L The self-positioning pose is determined below.
[0119] The gray solid circle represents the global localization pose. The trajectory composed of multiple global localization poses is called the global localization trajectory. That is, the global localization trajectory includes multiple global localization poses, which can be the coordinate system S of the 3D visual map. G The poses below, that is, each global localization pose in the global localization trajectory, are in the 3D vision map coordinate system S. G The global positioning pose is the global positioning pose under the 3D visual map.
[0120] The white dashed circle represents the fused localization pose. The trajectory composed of multiple fused localization poses is called the fused localization trajectory. That is, the fused localization trajectory includes multiple fused localization poses, which can be in the 3D visual map coordinate system S. G The poses below, that is, each fused localization pose in the fused localization trajectory, are in the 3D visual map coordinate system S. G Fusion positioning pose under the 3D visual map, that is, fusion positioning pose under the 3D visual map.
[0121] See Figure 5 As shown, since the target image comprises multiple frames, each frame corresponds to a self-localization pose, and a subset of images is selected from these frames as the test image, each test image corresponds to a global localization pose. Therefore, the number of self-localization poses is greater than the number of global localization poses. When obtaining the fused localization trajectory based on the self-localization trajectory and the global localization trajectory, each self-localization pose corresponds to one fused localization pose (i.e., a one-to-one correspondence between self-localization poses and fused localization poses), meaning the number of self-localization poses is the same as the number of fused localization poses. Therefore, the number of fused localization poses is also greater than the number of global localization poses.
[0122] In one possible implementation, the server can implement trajectory fusion and pose transformation functions, see [link to relevant documentation]. Figure 6 As shown, the server can implement trajectory fusion and pose transformation functions using the following steps to obtain the fused positioning trajectory of the terminal device in the 3D visual map:
[0123] Step 601: Select N self-localization poses corresponding to the target time period from all self-localization poses included in the self-localization trajectory, and select P global localization poses corresponding to the target time period from all global localization poses included in the global localization trajectory. For example, N can be greater than P.
[0124] For example, when fusing the self-localization trajectory and global localization trajectory for a target time period, N self-localization poses corresponding to the target time period (i.e., self-localization poses determined based on images acquired during the target time period) can be determined, along with P global localization poses corresponding to the target time period (i.e., global localization poses determined based on images acquired during the target time period). See [link to relevant documentation]. Figure 5 As shown, it can be and The self-localized poses between the points can be used as N self-localized poses corresponding to the target time period, which can be used to... and The global positioning poses between the target time periods are used as the P global positioning poses corresponding to the target time period.
[0125] Step 602: Based on N self-localization poses and P global localization poses, determine N fused localization poses corresponding to the N self-localization poses, with a one-to-one correspondence between the N self-localization poses and the N fused localization poses.
[0126] For example, see Figure 5 As shown, the self-localization pose can be determined based on N self-localization poses and P global localization poses. Corresponding fusion positioning pose Determine self-positioning pose Corresponding fusion positioning pose Determine self-positioning pose Corresponding fusion positioning pose And so on.
[0127] In one possible implementation, assume there are N self-localized poses, P global localized poses, and N fused localized poses. The N self-localized poses are all known values, the P global localized poses are all known values, and the N fused localized poses are all unknown values, representing the pose values that need to be solved. For example... Figure 5 As shown, self-localization pose With fusion positioning pose Correspondingly, self-localization pose With fusion positioning pose Correspondingly, self-localization pose With fusion positioning pose Correspondingly, and so on. Global positioning pose. With fusion positioning pose Correspondingly, global positioning pose With fusion positioning pose Correspondingly, and so on.
[0128] A first constraint value can be determined based on N self-localized poses and N fused localized poses. This first constraint value represents the residual value between the fused localized pose and the self-localized pose. For example, it can be based on... and The difference and The difference, ... and The difference is used to calculate the first constraint value. The formula for calculating the first constraint value is not limited in this embodiment; it can be related to any of the aforementioned differences.
[0129] A second constraint value can be determined based on P global localization poses and P fused localization poses (i.e., selecting P fused localization poses corresponding to the P global localization poses from N fused localization poses). The second constraint value represents the residual value (i.e., the absolute difference) between the fused localization pose and the global localization pose. For example, it can be based on... and The difference, ... and The difference is used to calculate the second constraint value. The formula for calculating the second constraint value is not limited in this embodiment; it can be related to the differences mentioned above.
[0130] The target constraint value can be calculated based on the first and second constraint values, such as the sum of the first and second constraint values. Since the N self-localization poses and P global localization poses are all known values, while the N fused localization poses are all unknown values, the target constraint value is minimized by adjusting the values of the N fused localization poses. When the target constraint value is minimized, the values of the N fused localization poses are the final solved pose values. Thus, the values of the N fused localization poses are obtained.
[0131] In one possible implementation, the target constraint value can be calculated using formula (1):
[0132]
[0133] In formula (1), F(T) represents the target constraint value. The part before the plus sign (hereinafter referred to as the first part) is the first constraint value, and the part after the plus sign (hereinafter referred to as the second part) is the second constraint value. Of course, the above are just examples of the target constraint value, the first constraint value and the second constraint value, and there are no restrictions on them.
[0134] Ω i,i+1 It is the residual information matrix for self-localization pose, which can be configured empirically without restrictions. Ω kIt is a residual information matrix for global positioning pose, which can be configured based on experience and is not subject to any restrictions.
[0135] The first part represents the relative transformation constraints between the self-localized pose and the fused localized pose, which can be reflected by the first constraint value. N represents all self-localized poses in the self-localization trajectory, i.e., N self-localized poses. The second part represents the global localization constraints between the global localization pose and the fused localized pose, which can be reflected by the second constraint value. P represents all global localization poses in the global localization trajectory, i.e., P global localization poses.
[0136] For the first and second parts, they can also be expressed by formulas (2) and (3):
[0137]
[0138]
[0139] In formulas (2) and (3), and To fuse localization poses (where there is no corresponding global localization pose), and For self-positioning pose. For the relative pose change constraint between two self-localized poses, e i,i+1 for and Relative pose change and Constraint residuals.
[0140] To fuse localization poses (with corresponding global localization poses) ), for The corresponding global localization pose, e k Indicates fused localization pose Compared to global positioning pose The residual.
[0141] Since the self-localization pose and global localization pose are known, while the fused localization pose is unknown, the optimization objective can be to minimize the value of F(T) so that the fused localization pose can be obtained. That is, the fused localization trajectory in the 3D visual map coordinate system can be referred to formula (4): arg min F(T). By minimizing the value of F(T), the fused localization trajectory can be obtained, and the fused localization trajectory can include multiple fused localization poses.
[0142] For example, in order to minimize the value of F(T), algorithms such as Gauss-Newton, gradient descent, and LM (Levenberg-Marquardt) can be used to solve the problem and obtain the fused localization pose, which will not be elaborated here.
[0143] Step 603: Generate the fusion positioning trajectory of the terminal device in the 3D visual map based on N fusion positioning poses. The fusion positioning trajectory includes N fusion positioning poses in the 3D visual map.
[0144] At this point, the server obtains the fused positioning trajectory in the 3D visual map, that is, the fused positioning trajectory in the 3D visual map coordinate system. The number of fused positioning poses in this fused positioning trajectory is greater than the number of global positioning poses in the global positioning trajectory. In other words, a high frame rate fused positioning trajectory can be obtained.
[0145] Step 604: Select an initial fused positioning pose from the fused positioning trajectory, and select an initial self-positioning pose corresponding to the initial fused positioning pose from the self-positioning trajectory.
[0146] Step 605: Select the target self-localization pose from the self-localization trajectory, and determine the target fused localization pose based on the initial fused localization pose, the initial self-localization pose, and the target self-localization pose.
[0147] For example, after generating the fused positioning trajectory, it can be updated. During the trajectory update process, an initial fused positioning pose can be selected from the fused positioning trajectory, an initial self-positioning pose can be selected from the self-positioning trajectory, and a target self-positioning pose can be selected from the self-positioning trajectory. Based on this, a target fused positioning pose can be determined based on the initial fused positioning pose, the initial self-positioning pose, and the target self-positioning pose. Then, a new fused positioning trajectory can be generated based on the target fused positioning pose and the fused positioning trajectory to replace the original fused positioning trajectory.
[0148] For example, in steps 601-603, see... Figure 5 As shown, the self-localization trajectory includes and The self-localization pose and global localization trajectory between them include and The global positioning pose and fused positioning trajectory include and The fusion of the localized poses between the two, and then, if a new self-localized pose is obtained. However, since there is no corresponding global localization pose, it is impossible to base the pose on both the global localization pose and the self-localization pose. Determine self-positioning pose Corresponding fusion positioning pose Based on this, in this embodiment, the fused positioning pose can also be determined using the following formula (4).
[0149]
[0150] In formula (4), Indicates self-localized pose The corresponding fused localization pose, i.e., the target fused localization pose. This represents the fused positioning pose, which is the initial fused positioning pose selected from the fused positioning trajectory. The expression represents the self-localized pose, that is, the pose selected from the self-localized trajectory. The corresponding initial self-localization pose, This represents the self-localized pose, i.e., the target self-localized pose selected from the self-localization trajectory. In summary, it can be seen that the initial fused localization pose can be used as a basis for... The initial self-localization pose and the target's self-localized pose Determine target fusion positioning pose
[0151] After obtaining the target fusion positioning pose Afterwards, a new fused localization trajectory can be generated, which can include the target fused localization pose. This allows for the updating of the fused positioning trajectory.
[0152] In the above process, steps 601-603 are the trajectory fusion process, and steps 604-605 are the pose transformation process. Trajectory fusion is the process of registering and fusing the self-localized trajectory with the global localization trajectory, realizing the transformation of the self-localized trajectory from the self-localization coordinate system to the 3D visual map coordinate system. The trajectory is corrected using the global localization result. When a new frame can obtain the global localization trajectory, trajectory fusion is performed again. Since not all frames can successfully obtain the global localization trajectory, the pose of these frames is output as the fused localization pose in the 3D visual map coordinate system through pose transformation, i.e., the pose transformation process.
[0153] VI. 3D Visualization Map of the Target Scene. A 3D visualization map of the target scene needs to be pre-constructed and stored on a server. The server can then display the trajectory based on this 3D visualization map. The 3D visualization map is a 3D visualization of the target scene, primarily used for trajectory display. It can be obtained through laser scanning and manual modeling. It is a viewable visualization map, and there are no restrictions on the method of constructing this 3D visualization map; it can be obtained using mapping algorithms.
[0154] Based on the 3D visual map and the 3D visualization map of the target scene, registration is required to ensure spatial alignment. For example, the 3D visualization map is sampled, transforming it from a triangular patch format into a dense point cloud. This point cloud is then registered with the 3D point cloud of the 3D visual map using the ICP (Iterative Closest Point) algorithm, yielding a transformation matrix T from the 3D visualization map to the 3D visual map. Finally, the transformation matrix T is used to transform the 3D visualization map to the 3D visual map coordinate system, resulting in a 3D visualization map aligned with the 3D visual map.
[0155] For example, the transformation matrix T (denoted as the target transformation matrix) can be determined in the following way:
[0156] Method 1: When constructing a 3D visual map and a 3D visualization map, multiple calibration points can be deployed in the target scene (different calibration points can be distinguished by different shapes, thus enabling identification of calibration points from the image). Both the 3D visual map and the 3D visualization map can include multiple calibration points. For each calibration point, a corresponding coordinate pair can be determined. This coordinate pair includes the position coordinates of the calibration point in the 3D visual map and the position coordinates of the calibration point in the 3D visualization map. Based on the coordinate pairs corresponding to multiple calibration points, the target transformation matrix can be determined. For example, the target transformation matrix T can be an m*n dimensional transformation matrix. The transformation relationship between the 3D visual map and the 3D visualization map can be: W = Q*T, where W represents the position coordinates in the 3D visualization map and Q represents the position coordinates in the 3D visual map. Substituting the multiple coordinate pairs corresponding to multiple calibration points into the above formula (i.e., the position coordinates of the calibration points in the 3D visual map as Q, and the position coordinates of the calibration points in the 3D visualization map as W), the target transformation matrix T can be obtained. This process will not be elaborated further.
[0157] Method 2: Obtain the initial transformation matrix. Based on this initial transformation matrix, map the position coordinates in the 3D visual map to the mapped coordinates in the 3D visualization map. Determine whether the initial transformation matrix has converged based on the relationship between the mapped coordinates and the actual coordinates in the 3D visualization map. If yes, the initial transformation matrix can be determined as the target transformation matrix, i.e., the target transformation matrix is obtained. If not, the initial transformation matrix can be adjusted, and the adjusted transformation matrix can be used as the initial transformation matrix. Then, return to execute the operation of mapping the position coordinates in the 3D visual map to the mapped coordinates in the 3D visualization map based on the initial transformation matrix, and so on, until the target transformation matrix is obtained.
[0158] For example, we can first obtain an initial transformation matrix. There are no restrictions on how we obtain this initial transformation matrix. It can be a randomly set initial transformation matrix or an initial transformation matrix obtained by a certain algorithm. This initial transformation matrix is a matrix that needs to be iteratively optimized. That is, we continuously iterate and optimize the initial transformation matrix, and use the iteratively optimized initial transformation matrix as the target transformation matrix.
[0159] After obtaining the initial transformation matrix, the position coordinates in the 3D visual map can be mapped to the mapped coordinates in the 3D visualization map. For example, the transformation relationship between the 3D visual map and the 3D visualization map can be: W = Q * T. That is, by using the position coordinates in the 3D visual map as Q and the initial transformation matrix as T, the position coordinates in the 3D visualization map can be obtained (for convenience, we will refer to them as mapped coordinates). Then, the convergence of the initial transformation matrix is determined based on the relationship between the mapped coordinates and the actual coordinates in the 3D visualization map. For example, the mapped coordinates in the 3D visualization map are the coordinates transformed by the initial transformation matrix, and the actual coordinates in the 3D visualization map are the true coordinates in the 3D visualization map. The smaller the difference between the mapped coordinates and the actual coordinates, the higher the accuracy of the initial transformation matrix; the larger the difference, the lower the accuracy of the initial transformation matrix. Based on the above principle, the convergence of the initial transformation matrix can be determined based on the difference between the mapped coordinates and the actual coordinates.
[0160] For example, if the difference between the mapped coordinates and the actual coordinates (which can be the sum of multiple sets of differences, each set of differences corresponding to a difference between the mapped coordinates and the actual coordinates) is less than a threshold, then the initial transformation matrix is determined to have converged; if the difference between the mapped coordinates and the actual coordinates is not less than the threshold, then the initial transformation matrix is determined to have not converged.
[0161] If the initial transformation matrix does not converge, it can be adjusted. There are no restrictions on this adjustment process; for example, the ICP algorithm can be used to adjust the initial transformation matrix, and the adjusted transformation matrix can be used as the initial transformation matrix. The process then returns to perform the operation of mapping the position coordinates in the 3D visual map to the mapped coordinates in the 3D visualization map based on the initial transformation matrix. This process continues until the target transformation matrix is obtained. If the initial transformation matrix has converged, it is then determined as the target transformation matrix.
[0162] Method 3: Sample the 3D visualization map to obtain a first point cloud corresponding to the 3D visualization map; sample the 3D visual map to obtain a second point cloud corresponding to the 3D visual map. Use the ICP algorithm to register the first and second point clouds to obtain the target transformation matrix between the 3D visual map and the 3D visualization map. Clearly, since both the first and second point clouds are available, and both contain a large number of 3D points, the ICP algorithm can be used for registration based on these numerous 3D points. No restrictions are placed on this registration process.
[0163] VII. Trajectory Display. After obtaining the fused positioning trajectory, the server, for each fused positioning pose in the trajectory, can convert the fused positioning pose into a target positioning pose in the 3D visualization map based on the target transformation matrix between the 3D visual map and the 3D visualized map, and then display the target positioning pose through the 3D visualization map. Based on this, administrators can open a web browser and access the server via the network to view the target positioning poses displayed in the 3D visualization map; these target positioning poses form a trajectory. By reading and rendering the 3D visualization map, the server can display the target positioning poses of the terminal device in the 3D visualization map, allowing administrators to view the target positioning poses displayed in the 3D visualization map. Administrators can change the viewing perspective by dragging the mouse to achieve 3D viewing of the trajectory. For example, the server includes client software that reads and renders the 3D visualization map and displays the target positioning poses on the 3D visualization map. Based on this, users (such as administrators) can access the client software through a web browser to view the target positioning poses displayed in the 3D visualization map through the client software. For example, when viewing the target positioning poses displayed in the 3D visualization map through the client software, the viewing perspective of the 3D visualization map can be changed by dragging the mouse.
[0164] As can be seen from the above technical solutions, this application proposes a cloud-edge combined positioning and display method. The terminal device calculates a high-frame-rate self-positioning trajectory and only sends the trajectory and a small amount of image data to be measured, reducing the amount of data transmitted over the network. Global positioning is performed on the server, thereby reducing the computational and storage resource consumption of the terminal device. The cloud-edge fusion system architecture can distribute computational pressure, reduce the hardware cost of the terminal device, and reduce the amount of data transmitted over the network. The final positioning result can be displayed on a 3D visualization map, and administrators can interact with the display by accessing the server via a web interface.
[0165] Based on the same application concept as the above method, this application proposes a cloud-edge management system, which includes a terminal device and a server. The server includes a 3D visual map of a target scene, wherein: the terminal device is used to acquire a target image of the target scene and motion data of the terminal device during the movement of the target scene, and determine the self-localization trajectory of the terminal device based on the target image and the motion data; if the target image includes multiple frames, a portion of the images are selected as test images, and the test images and the self-localization trajectory are sent to the server; the server is used to generate a fused positioning trajectory of the terminal device in the 3D visual map based on the test images and the self-localization trajectory, the fused positioning trajectory including multiple fused positioning poses; for each fused positioning pose in the fused positioning trajectory, a target positioning pose corresponding to the fused positioning pose is determined, and the target positioning pose is displayed.
[0166] For example, the terminal device includes a visual sensor and a motion sensor; wherein the visual sensor is used to acquire a target image of the target scene, and the motion sensor is used to acquire motion data of the terminal device; wherein the terminal device is a wearable device, and the visual sensor and the motion sensor are deployed on the wearable device; or, the terminal device is a recorder, and the visual sensor and the motion sensor are deployed on the recorder; or, the terminal device is a camera, and the visual sensor and the motion sensor are deployed on the camera.
[0167] For example, when the server generates the fused positioning trajectory of the terminal device in the 3D visual map based on the image to be tested and the self-localization trajectory, it is specifically used for:
[0168] Target map points corresponding to the image to be tested are determined from the three-dimensional visual map, and the global positioning trajectory of the terminal device in the three-dimensional visual map is determined based on the target map points;
[0169] The terminal device generates a fused positioning trajectory in the 3D visual map based on the self-localization trajectory and the global positioning trajectory; wherein the frame rate of the fused positioning pose included in the fused positioning trajectory is greater than the frame rate of the global positioning pose included in the global positioning trajectory; and the frame rate of the fused positioning pose included in the fused positioning trajectory is equal to the frame rate of the self-localization pose included in the self-localization trajectory.
[0170] For example, when the server determines the target positioning pose corresponding to the fused positioning pose and displays the target positioning pose, it specifically performs the following: based on the target transformation matrix between the 3D visual map and the 3D visualization map, converts the fused positioning pose into the target positioning pose in the 3D visualization map, and displays the target positioning pose through the 3D visualization map;
[0171] The server includes client software, which reads and renders the 3D visualization map and displays the target positioning pose on the 3D visualization map.
[0172] Users access the client software via a web browser to view the target's location pose displayed on the 3D visualization map.
[0173] Specifically, when viewing the target's location pose displayed on the 3D visualization map through the client software, the viewing perspective of the 3D visualization map can be changed by dragging the mouse.
[0174] Based on the same concept as the above method, this application proposes a pose display device for use in a cloud-edge management system server. The server includes a 3D visual map of the target scene. See [link to relevant documentation]. Figure 7 The diagram shown is a structural diagram of the pose display device, which includes:
[0175] The acquisition module 71 is used to acquire the image to be tested and the self-localization trajectory; wherein, the self-localization trajectory is determined by the terminal device based on the target image of the target scene and the motion data of the terminal device, and the image to be tested is a portion of the multi-frame images included in the target image; the generation module 72 is used to generate the fused localization trajectory of the terminal device in the three-dimensional visual map based on the image to be tested and the self-localization trajectory, the fused localization trajectory including multiple fused localization poses; the display module 73 is used to determine the target localization pose corresponding to each fused localization pose in the fused localization trajectory, and display the target localization pose.
[0176] For example, when generating the fused positioning trajectory of the terminal device in the 3D visual map based on the image to be tested and the self-localization trajectory, the generation module 72 is specifically used to: determine the target map point corresponding to the image to be tested from the 3D visual map; determine the global positioning trajectory of the terminal device in the 3D visual map based on the target map point; generate the fused positioning trajectory of the terminal device in the 3D visual map based on the self-localization trajectory and the global positioning trajectory; the frame rate of the fused positioning pose included in the fused positioning trajectory is greater than the frame rate of the global positioning pose included in the global positioning trajectory; the frame rate of the fused positioning pose included in the fused positioning trajectory is equal to the frame rate of the self-localization pose included in the self-localization trajectory.
[0177] For example, the three-dimensional visual map includes at least one of the following: a pose matrix corresponding to a sample image, a global descriptor corresponding to a sample image, a local descriptor corresponding to a feature point in the sample image, and map point information; the generation module 72 determines the target map point corresponding to the image to be tested from the three-dimensional visual map, and when determining the global positioning trajectory of the terminal device in the three-dimensional visual map based on the target map point, it specifically performs the following: for each frame of the image to be tested, select candidate sample images from the multiple frame sample images based on the similarity between the image to be tested and the multiple frame sample images corresponding to the three-dimensional visual map; obtain multiple feature points from the image to be tested; for each feature point, determine the target map point corresponding to the feature point from the multiple map points corresponding to the candidate sample images; determine the global positioning pose in the three-dimensional visual map corresponding to the image to be tested based on the multiple feature points and the target map points corresponding to the multiple feature points; and generate the global positioning trajectory of the terminal device in the three-dimensional visual map based on the global positioning poses corresponding to all the images to be tested.
[0178] For example, when generating the fused positioning trajectory of the terminal device in the 3D visual map based on the self-localization trajectory and the global positioning trajectory, the generation module 72 is specifically used to: select N self-localization poses corresponding to the target time period from all self-localization poses included in the self-localization trajectory, and select P global positioning poses corresponding to the target time period from all global positioning poses included in the global positioning trajectory; N is greater than P; determine N fused positioning poses corresponding to the N self-localization poses based on the N self-localization poses and the P global positioning poses, with a one-to-one correspondence between the N self-localization poses and the N fused positioning poses; and generate the fused positioning trajectory of the terminal device in the 3D visual map based on the N fused positioning poses.
[0179] For example, when the display module 73 determines and displays the target positioning pose corresponding to the fused positioning pose, it is specifically used to: convert the fused positioning pose into a target positioning pose in the 3D visualization map based on the target transformation matrix between the 3D visual map and the 3D visualization map, and display the target positioning pose through the 3D visualization map; wherein, the display module, the display module 73, is further used to determine the target transformation matrix between the 3D visual map and the 3D visualization map in the following manner: for each of the multiple calibration points, determine the coordinate pair corresponding to the calibration point, the coordinate pair including the position coordinates of the calibration point in the 3D visual map and the position coordinates of the calibration point in the 3D visualization map; determine the target transformation matrix based on the coordinate pairs corresponding to the multiple calibration points; or, obtain an initial transformation matrix, and transform the fused positioning pose into a target positioning pose in the 3D visualization map based on the initial transformation matrix. The position coordinates in the 3D visual map are mapped to the mapped coordinates in the 3D visualization map. Based on the relationship between the mapped coordinates and the actual coordinates in the 3D visualization map, it is determined whether the initial transformation matrix has converged. If yes, the initial transformation matrix is determined as the target transformation matrix. If not, the initial transformation matrix is adjusted, and the adjusted transformation matrix is used as the initial transformation matrix. The operation of mapping the position coordinates in the 3D visual map to the mapped coordinates in the 3D visualization map based on the initial transformation matrix is then performed. Alternatively, the 3D visualization map is sampled to obtain a first point cloud corresponding to the 3D visualization map. The 3D visual map is also sampled to obtain a second point cloud corresponding to the 3D visual map. The ICP algorithm is used to register the first point cloud and the second point cloud to obtain the target transformation matrix between the 3D visual map and the 3D visualization map.
[0180] Based on the same application concept as the above method, this application proposes a server, which may include: a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the pose display method disclosed in the above examples of this application.
[0181] Based on the same concept as the above method, this application also provides a machine-readable storage medium storing a plurality of computer instructions, which, when executed by a processor, can implement the pose display method disclosed in the above examples of this application.
[0182] The aforementioned machine-readable storage medium can be any electronic, magnetic, optical, or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, machine-readable storage media can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or combinations thereof.
[0183] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, which can take the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email sending and receiving device, game console, tablet computer, wearable device, or any combination of these devices.
[0184] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.
[0185] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, embodiments of this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0186] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0187] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0188] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A pose display method, characterized in that, This is applied to a cloud-edge management system, which includes terminal devices and a server. The server includes a 3D visual map of the target scene, comprising: During the movement of the target scene, the terminal device acquires the target image of the target scene and the motion data of the terminal device, and determines the self-localization trajectory of the terminal device based on the target image and the motion data; if the target image includes multiple frames, a portion of the images are selected from the multiple frames as the test image, and the test image and the self-localization trajectory are sent to the server. The server generates a fused positioning trajectory of the terminal device in the 3D visual map based on the image to be tested and the self-localization trajectory, the fused positioning trajectory including multiple fused positioning poses; wherein, the server determines a target map point corresponding to the image to be tested from the 3D visual map, and determines a global positioning trajectory of the terminal device in the 3D visual map based on the target map point; the server generates a fused positioning trajectory of the terminal device in the 3D visual map based on the self-localization trajectory and the global positioning trajectory; wherein, the server selects N self-localization poses corresponding to a target time period from all self-localization poses included in the self-localization trajectory, and selects N self-localization poses corresponding to a target time period from the global positioning poses included in the self-localization trajectory. From all global positioning poses included in the positioning trajectory, P global positioning poses corresponding to the target time period are selected; wherein, N is greater than P; based on the N self-positioning poses and the P global positioning poses, N fused positioning poses corresponding to the N self-positioning poses are determined, and there is a one-to-one correspondence between the N self-positioning poses and the N fused positioning poses; based on the N fused positioning poses, a fused positioning trajectory of the terminal device in a 3D visual map is generated; wherein, the frame rate of the fused positioning poses included in the fused positioning trajectory is greater than the frame rate of the global positioning poses included in the global positioning trajectory; the frame rate of the fused positioning poses included in the fused positioning trajectory is equal to the frame rate of the self-positioning poses included in the self-positioning trajectory. For each fused positioning pose in the fused positioning trajectory, the server determines the target positioning pose corresponding to the fused positioning pose and displays the target positioning pose.
2. The method according to claim 1, characterized in that, The terminal device determines its self-localization trajectory based on the target image and the motion data, including: The terminal device traverses the multi-frame images to obtain the current frame image; determines the self-localization pose corresponding to the current frame image based on the self-localization pose corresponding to the K preceding frames, the map position of the terminal device in the self-localization coordinate system, and the motion data; and generates the self-localization trajectory of the terminal device in the self-localization coordinate system based on the self-localization pose corresponding to the multi-frame images. If the current frame image is a key image, a map position in a self-positioning coordinate system is generated based on the current position of the terminal device; if the number of matching feature points between the current frame image and the previous frame image does not reach a preset threshold, the current frame image is determined to be a key image.
3. The method according to claim 1, characterized in that, The three-dimensional visual map includes at least one of the following: a pose matrix corresponding to a sample image, a global descriptor corresponding to a sample image, a local descriptor corresponding to a feature point in the sample image, and map point information; the server determines a target map point corresponding to the image to be tested from the three-dimensional visual map, and determines the global positioning trajectory of the terminal device in the three-dimensional visual map based on the target map point, including: For each frame of the image to be tested, the server selects candidate sample images from the multi-frame sample images based on the similarity between the image to be tested and the multi-frame sample images corresponding to the 3D visual map. The server obtains multiple feature points from the image to be tested; for each feature point, it determines the target map point corresponding to that feature point from multiple map points corresponding to the candidate sample image; The global positioning pose of the image under test in the 3D visual map is determined based on the plurality of feature points and the target map points corresponding to the plurality of feature points; the global positioning trajectory of the terminal device in the 3D visual map is generated based on the global positioning poses corresponding to all images under test.
4. The method according to claim 1, characterized in that, The server determines and displays the target positioning pose corresponding to the fused positioning pose, including: converting the fused positioning pose into a target positioning pose in the 3D visualization map based on the target transformation matrix between the 3D visual map and the 3D visualization map, and displaying the target positioning pose through the 3D visualization map; wherein, the method for determining the target transformation matrix between the 3D visual map and the 3D visualization map includes: For each of the multiple calibration points, determine the coordinate pair corresponding to that calibration point. The coordinate pair includes the position coordinates of the calibration point in the 3D visual map and the position coordinates of the calibration point in the 3D visualization map. Determine the target transformation matrix based on the coordinate pairs corresponding to the multiple calibration points. Alternatively, an initial transformation matrix can be obtained. Based on the initial transformation matrix, the position coordinates in the 3D visual map are mapped to the mapped coordinates in the 3D visualization map. Based on the relationship between the mapped coordinates and the actual coordinates in the 3D visualization map, it is determined whether the initial transformation matrix has converged. If so, the initial transformation matrix is determined as the target transformation matrix. If not, the initial transformation matrix is adjusted, and the adjusted transformation matrix is used as the initial transformation matrix. Then, the operation of mapping the position coordinates in the 3D visual map to the mapped coordinates in the 3D visualization map based on the initial transformation matrix is returned. Alternatively, the three-dimensional visualization map is sampled to obtain a first point cloud corresponding to the three-dimensional visualization map; and the three-dimensional visual map is sampled to obtain a second point cloud corresponding to the three-dimensional visual map; the first point cloud and the second point cloud are registered using the ICP algorithm to obtain the target transformation matrix between the three-dimensional visual map and the three-dimensional visualization map.
5. A cloud-edge management system, characterized in that, The cloud-edge management system includes terminal devices and servers. The server includes a 3D visual map of the target scene, wherein: The terminal device is used to acquire a target image of the target scene and motion data of the terminal device during the movement of the target scene, and determine the self-localization trajectory of the terminal device based on the target image and the motion data; if the target image includes multiple frames, a portion of the images are selected from the multiple frames as the image to be tested, and the image to be tested and the self-localization trajectory are sent to the server. The server is configured to generate a fused positioning trajectory of the terminal device in the 3D visual map based on the image to be tested and the self-localization trajectory, the fused positioning trajectory including multiple fused positioning poses; for each fused positioning pose in the fused positioning trajectory, determine a target positioning pose corresponding to the fused positioning pose, and display the target positioning pose; wherein, when the server generates the fused positioning trajectory of the terminal device in the 3D visual map based on the image to be tested and the self-localization trajectory, it is specifically configured to: determine a target map point corresponding to the image to be tested from the 3D visual map; determine a global positioning trajectory of the terminal device in the 3D visual map based on the target map point; and generate a fused positioning trajectory of the terminal device in the 3D visual map based on the self-localization trajectory and the global positioning trajectory; wherein The server selects N self-localization poses corresponding to the target time period from all self-localization poses included in the self-localization trajectory, and selects P global positioning poses corresponding to the target time period from all global positioning poses included in the global positioning trajectory; wherein, N is greater than P; based on the N self-localization poses and the P global positioning poses, N fused positioning poses corresponding to the N self-localization poses are determined, and the N self-localization poses and N fused positioning poses correspond one-to-one; based on the N fused positioning poses, a fused positioning trajectory of the terminal device in a 3D visual map is generated; wherein, the frame rate of the fused positioning poses included in the fused positioning trajectory is greater than the frame rate of the global positioning poses included in the global positioning trajectory; the frame rate of the fused positioning poses included in the fused positioning trajectory is equal to the frame rate of the self-localization poses included in the self-localization trajectory.
6. The system according to claim 5, characterized in that, The terminal device includes a visual sensor and a motion sensor; wherein, the visual sensor is used to acquire a target image of the target scene, and the motion sensor is used to acquire motion data of the terminal device; Wherein, the terminal device is a wearable device, and the visual sensor and the motion sensor are deployed on the wearable device; or, the terminal device is a recorder, and the visual sensor and the motion sensor are deployed on the recorder; or, the terminal device is a camera, and the visual sensor and the motion sensor are deployed on the camera.
7. The system according to claim 5, characterized in that, When the server determines and displays the target positioning pose corresponding to the fused positioning pose, it specifically performs the following: based on the target transformation matrix between the 3D visual map and the 3D visualization map, converts the fused positioning pose into the target positioning pose in the 3D visualization map, and displays the target positioning pose through the 3D visualization map. The server includes client software, which reads and renders the 3D visualization map and displays the target positioning pose on the 3D visualization map. Users access the client software via a web browser to view the target's location pose displayed on the 3D visualization map. Specifically, when viewing the target's location pose displayed on the 3D visualization map through the client software, the viewing perspective of the 3D visualization map can be changed by dragging the mouse.
8. A pose display device, characterized in that, A server used in a cloud-edge management system, the server including a 3D visual map of the target scene, the device comprising: An acquisition module is used to acquire the image to be tested and the self-localization trajectory; wherein, the self-localization trajectory is determined by the terminal device based on the target image of the target scene and the motion data of the terminal device, and the image to be tested is a portion of the multiple frames included in the target image; The generation module is used to generate a fused positioning trajectory of the terminal device in the 3D visual map based on the image to be tested and the self-localization trajectory, wherein the fused positioning trajectory includes multiple fused positioning poses; specifically, when the generation module generates the fused positioning trajectory of the terminal device in the 3D visual map based on the image to be tested and the self-localization trajectory, it is used to: determine a target map point corresponding to the image to be tested from the 3D visual map; determine a global positioning trajectory of the terminal device in the 3D visual map based on the target map point; and generate the fused positioning trajectory of the terminal device in the 3D visual map based on the self-localization trajectory and the global positioning trajectory; wherein, when the generation module generates the fused positioning trajectory of the terminal device in the 3D visual map based on the self-localization trajectory and the global positioning trajectory, Specifically, it is used for: selecting N self-localization poses corresponding to a target time period from all self-localization poses included in the self-localization trajectory, and selecting P global positioning poses corresponding to the target time period from all global positioning poses included in the global positioning trajectory; N is greater than P; determining N fused positioning poses corresponding to the N self-localization poses based on the N self-localization poses and the P global positioning poses, with a one-to-one correspondence between the N self-localization poses and the N fused positioning poses; generating a fused positioning trajectory of the terminal device in a 3D visual map based on the N fused positioning poses; wherein, the frame rate of the fused positioning poses included in the fused positioning trajectory is greater than the frame rate of the global positioning poses included in the global positioning trajectory; the frame rate of the fused positioning poses included in the fused positioning trajectory is equal to the frame rate of the self-localization poses included in the self-localization trajectory. The display module is used to determine the target positioning pose corresponding to each fused positioning pose in the fused positioning trajectory, and to display the target positioning pose.
9. The apparatus according to claim 8, Its features are, The three-dimensional visual map includes at least one of the following: a pose matrix corresponding to a sample image, a global descriptor corresponding to a sample image, a local descriptor corresponding to a feature point in the sample image, and map point information. The generation module determines target map points corresponding to the image to be tested from the three-dimensional visual map. When determining the global positioning trajectory of the terminal device in the three-dimensional visual map based on the target map points, it specifically performs the following: selecting candidate sample images from multiple frames of sample images based on the similarity between the image to be tested and multiple frames of sample images corresponding to the three-dimensional visual map; acquiring multiple feature points from the image to be tested; for each feature point, determining a target map point corresponding to that feature point from multiple map points corresponding to the candidate sample images; determining the global positioning pose in the three-dimensional visual map corresponding to the image to be tested based on the multiple feature points and the target map points corresponding to the multiple feature points; and generating the global positioning trajectory of the terminal device in the three-dimensional visual map based on the global positioning poses corresponding to all images to be tested. Specifically, when the display module determines and displays the target positioning pose corresponding to the fused positioning pose, it is used to: convert the fused positioning pose into a target positioning pose in the 3D visualization map based on the target transformation matrix between the 3D visual map and the 3D visualization map, and display the target positioning pose through the 3D visualization map; wherein, the display module is further used to determine the target transformation matrix between the 3D visual map and the 3D visualization map in the following manner: for each of the multiple calibration points, determine the coordinate pair corresponding to the calibration point, the coordinate pair including the position coordinates of the calibration point in the 3D visual map and the position coordinates of the calibration point in the 3D visualization map; determine the target transformation matrix based on the coordinate pairs corresponding to the multiple calibration points; or, obtain an initial transformation matrix, and transform the 3D visual map based on the initial transformation matrix. The position coordinates in the map are mapped to the mapped coordinates in the 3D visualization map. Based on the relationship between the mapped coordinates and the actual coordinates in the 3D visualization map, it is determined whether the initial transformation matrix has converged. If yes, the initial transformation matrix is determined as the target transformation matrix. If not, the initial transformation matrix is adjusted, and the adjusted transformation matrix is used as the initial transformation matrix. The operation of mapping the position coordinates in the 3D visual map to the mapped coordinates in the 3D visualization map based on the initial transformation matrix is then performed. Alternatively, the 3D visualization map is sampled to obtain a first point cloud corresponding to the 3D visualization map. The 3D visual map is also sampled to obtain a second point cloud corresponding to the 3D visual map. The ICP algorithm is used to register the first point cloud and the second point cloud to obtain the target transformation matrix between the 3D visual map and the 3D visualization map.
Citation Information
Patent Citations
Wide area localization from SLAM maps
CN105143821A
Global drifting-free autonomous robot simultaneous localization and mapping method
CN112945233A