Method, device and apparatus for determining posture

By acquiring target images and motion data on the terminal device and integrating self-positioning and global positioning with three-dimensional visual maps, the problem of poor satellite navigation signals in indoor environments is solved, and high-frame rate and high-precision indoor positioning is achieved, which is suitable for the positioning needs of industries such as coal, electricity, and petrochemicals.

CN114120301BActive Publication Date: 2025-09-26HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111350622.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-15
Publication Date
2025-09-26
Estimated Expiration
2041-11-15

AI Technical Summary

Technical Problem

In indoor environments, the poor signal of GPS or Beidou satellite navigation systems makes it impossible for terminal devices to accurately locate themselves, especially in energy industries such as coal, electricity, and petrochemicals, where indoor positioning needs cannot be met.

Method used

By acquiring the target image and motion data of the target scene, the self-localization trajectory and the global positioning trajectory are fused using the three-dimensional visual map to generate a high-frame-rate fused positioning trajectory, eliminating the accumulated error of self-localization and achieving high-precision indoor positioning.

Benefits of technology

It achieves high frame rate and high-precision indoor positioning, which can quickly determine the location of personnel in energy industries such as coal, electricity, and petrochemicals, ensuring safety and achieving efficient management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114120301B_ABST
    Figure CN114120301B_ABST
Patent Text Reader

Abstract

The present application provides a posture determination method, apparatus, and device, which includes: obtaining a target image of a target scene and motion data of a terminal device; determining a self-positioning trajectory of the terminal device based on the target image and motion data; determining a target map point corresponding to the target image from a three-dimensional visual map, and determining a global positioning trajectory of the terminal device in the three-dimensional visual map based on the target map point; generating a fused positioning trajectory of the terminal device in the three-dimensional visual map based on the self-positioning trajectory and the global positioning trajectory, and outputting the fused positioning trajectory; wherein the frame rate of the fused positioning posture included in the fused positioning trajectory is greater than the frame rate of the global positioning posture included in the global positioning trajectory. Through the technical solution of the present application, a high-frame-rate and high-precision positioning function is achieved, and a globally consistent high-frame-rate positioning function is achieved indoors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision, and in particular to a method, apparatus, and device for determining a posture. Background Art

[0002] GPS (Global Positioning System) is a high-precision radio navigation and positioning system based on artificial Earth satellites. It provides accurate geographic location, vehicle speed, and precise time information anywhere in the world, including near-Earth space. The Beidou satellite navigation system, consisting of three segments: the space segment, the ground segment, and the user segment, provides users with high-precision, highly reliable positioning, navigation, and timing services around the world, 24 / 7, and possesses regional navigation, positioning, and timing capabilities.

[0003] Because terminal devices are equipped with GPS or BeiDou satellite navigation systems, they can be used to locate the terminal device when positioning is required. In outdoor environments, due to the relatively strong GPS or BeiDou signals, the GPS or BeiDou satellite navigation systems can be used to accurately locate the terminal device. However, indoors, due to the relatively poor GPS or BeiDou signals, the GPS or BeiDou satellite navigation systems cannot accurately locate the terminal device. For example, in energy industries such as coal, electricity, and petrochemicals, there is an increasing demand for positioning. These positioning needs are generally in indoor environments, where accurate positioning of the terminal device is difficult due to signal obstruction and other issues. Summary of the Invention

[0004] The present application provides a posture determination method, which is applied to a terminal device, wherein the terminal device includes a three-dimensional visual map of a target scene. When the terminal device moves in the target scene, the method includes:

[0005] Acquiring a target image of the target scene and motion data of the terminal device;

[0006] determining a self-positioning trajectory of the terminal device based on the target image and the motion data;

[0007] Determining a target map point corresponding to the target image from the three-dimensional visual map, and determining a global positioning trajectory of the terminal device in the three-dimensional visual map based on the target map point;

[0008] A fused positioning trajectory of the terminal device in the three-dimensional visual map is generated based on the self-positioning trajectory and the global positioning trajectory, and the fused positioning trajectory is output; wherein the frame rate of the fused positioning posture included in the fused positioning trajectory is greater than the frame rate of the global positioning posture included in the global positioning trajectory.

[0009] The present application provides a posture determination device, which is applied to a terminal device, wherein the terminal device includes a three-dimensional visual map of a target scene. When the terminal device moves in the target scene, the device includes:

[0010] An acquisition module, configured to acquire a target image of the target scene and motion data of the terminal device;

[0011] a determination module, configured to determine a self-positioning trajectory of the terminal device based on the target image and the motion data; determine a target map point corresponding to the target image from the three-dimensional visual map, and determine a global positioning trajectory of the terminal device in the three-dimensional visual map based on the target map point;

[0012] A generation module is used to generate a fused positioning trajectory of the terminal device in the three-dimensional visual map based on the self-positioning trajectory and the global positioning trajectory, and output the fused positioning trajectory; the frame rate of the fused positioning posture included in the fused positioning trajectory is greater than the frame rate of the global positioning posture included in the global positioning trajectory.

[0013] The present application provides a terminal device, comprising: a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the posture determination method disclosed in the above example of the present application.

[0014] The present application provides a terminal device, including:

[0015] a visual sensor, configured to acquire a target image of the target scene during movement of the terminal device in the target scene, and input the target image to a processor;

[0016] a motion sensor, configured to obtain motion data of the terminal device during movement of the terminal device in the target scene, and input the motion data to the processor;

[0017] A processor is configured to determine a self-positioning trajectory of the terminal device based on the target image and the motion data; determine a target map point corresponding to the target image from a three-dimensional visual map of the target scene, and determine a global positioning trajectory of the terminal device in the three-dimensional visual map based on the target map point; generate a fused positioning trajectory of the terminal device in the three-dimensional visual map based on the self-positioning trajectory and the global positioning trajectory, and output the fused positioning trajectory; the frame rate of the fused positioning pose included in the fused positioning trajectory is greater than the frame rate of the global positioning pose included in the global positioning trajectory.

[0018] It can be seen from the above technical solution that in the embodiment of the present application, during the movement of the terminal device in the target scene, the terminal device can determine the self-positioning trajectory of the terminal device based on the target image of the target scene and the motion data of the terminal device, and determine the global positioning trajectory of the terminal device in the three-dimensional visual map based on the target image of the target scene, and generate a fused positioning trajectory of the terminal device in the three-dimensional visual map based on the self-positioning trajectory and the global positioning trajectory, and output the fused positioning trajectory. In the above method, high-frame-rate self-positioning can be performed based on the target image and motion data to obtain a high-frame-rate self-positioning trajectory, and low-frame-rate global positioning can be performed based on the target image and the three-dimensional visual map to obtain a low-frame-rate global positioning trajectory. Then, the high-frame-rate self-positioning trajectory and the low-frame-rate global positioning trajectory are fused to eliminate the accumulated error of self-positioning and obtain a high-frame-rate fused positioning trajectory, that is, a high-frame-rate fused positioning trajectory in the three-dimensional visual map, to achieve high-frame-rate and high-precision positioning functions, and to achieve globally consistent high-frame-rate positioning functions indoors. In the above method, the target scene can be an indoor environment, and a high-precision, low-cost, and easy-to-deploy indoor positioning function can be achieved based on the target image and motion data. It is a vision-based indoor positioning method that can be applied to energy industries such as coal, electricity, and petrochemicals to achieve indoor positioning of personnel (such as workers, inspectors, etc.), quickly obtain personnel location information, ensure personnel safety, and achieve efficient personnel management. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 1 is a flow chart of a method for determining a posture in one embodiment of the present application;

[0020] Figure 2 is a schematic structural diagram of a terminal device in one embodiment of the present application;

[0021] Figure 3 This is a schematic diagram of a process for determining a self-positioning trajectory in one embodiment of the present application;

[0022] Figure 4 This is a schematic diagram of a process for determining a global positioning trajectory in one embodiment of the present application;

[0023] Figure 5 It is a schematic diagram of self-localization trajectory, global localization trajectory and fusion localization trajectory;

[0024] Figure 6 This is a schematic diagram of a process for determining a fused positioning trajectory in one embodiment of the present application;

[0025] Figure 7 It is a structural diagram of a posture determination device in one embodiment of the present application. DETAILED DESCRIPTION

[0026] The terms used in the embodiments of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The singular forms "a," "the," and "the" used in this application and claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to any or all possible combinations of one or more associated listed items.

[0027] It should be understood that although the terms first, second, third, etc. may be used to describe various information in the embodiments of the present application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" used may also be interpreted as "at the time of" or "when" or "in response to determining".

[0028] In an embodiment of the present application, a posture determination method is proposed. The method can be applied to a terminal device that includes a three-dimensional visual map of a target scene. For example, the terminal device downloads the three-dimensional visual map of the target scene from a server and stores the three-dimensional visual map of the target scene. As the terminal device moves within the target scene, the posture determination method is used to determine the posture of the terminal device and output the posture of the terminal device.

[0029] See also Figure 1 FIG. 1 is a flow chart of the posture determination method, which may include:

[0030] Step 101: Acquire a target image of a target scene and motion data of a terminal device.

[0031] Step 102: Determine a self-positioning trajectory of the terminal device based on the target image and the motion data.

[0032] Exemplarily, if the target image includes multiple frame images, the current frame image can be traversed from the multiple frame images; based on the self-positioning postures corresponding to the K frame images preceding the current frame image, the map position of the terminal device in the self-positioning coordinate system and the motion data, the self-positioning posture corresponding to the current frame image is determined; based on the self-positioning postures corresponding to the multiple frame images, the self-positioning trajectory of the terminal device in the self-positioning coordinate system is generated.

[0033] For example, if the current frame image is a key image, a map position in the self-positioning coordinate system can be generated based on the current position of the terminal device (i.e., the position corresponding to the current frame image). If the current frame image is a non-key image, there is no need to generate a map position in the self-positioning coordinate system based on the current position of the terminal device.

[0034] If the number of matching feature points between the current frame image and the previous frame image reaches an unpredictable threshold, the current frame image is determined to be a key image. If the number of matching feature points between the current frame image and the previous frame image reaches a preset threshold, the current frame image is determined to be a non-key image.

[0035] Step 103: Determine a target map point corresponding to the target image from the three-dimensional visual map of the target scene, and determine a global positioning trajectory of the terminal device in the three-dimensional visual map based on the target map point.

[0036] Exemplarily, if the target image includes multiple frames of images, M frames of images are selected from the multiple frames of images as the images to be tested, that is, part of the images in the multiple frames of images are selected as the images to be tested, and M can be a positive integer, such as 1, 2, 3, etc. For each frame of the image to be tested, a candidate sample image is selected from the multiple frames of sample images based on the similarity between the image to be tested and the multiple frames of sample images corresponding to the three-dimensional visual map. A plurality of feature points are obtained from the image to be tested; for each feature point, a target map point corresponding to the feature point is determined from the multiple map points corresponding to the candidate sample image. The global positioning pose in the three-dimensional visual map corresponding to the image to be tested is determined based on the multiple feature points and the target map points corresponding to the multiple feature points. The global positioning trajectory of the terminal device in the three-dimensional visual map is generated based on the global positioning pose corresponding to the M frames of the image to be tested.

[0037] Exemplarily, based on the similarity between the image to be tested and the multiple frames of sample images corresponding to the three-dimensional visual map, a candidate sample image is selected from the multiple frames of sample images, including but not limited to: determining a global descriptor to be tested corresponding to the image to be tested, and determining the distance between the global descriptor to be tested and a sample global descriptor corresponding to each frame of sample image corresponding to the three-dimensional visual map; wherein the three-dimensional visual map includes a sample global descriptor corresponding to each frame of sample image. Based on the distance between the global descriptor to be tested and each sample global descriptor, a candidate sample image is selected from the multiple frames of sample images; wherein the distance between the global descriptor to be tested and the sample global descriptor corresponding to the candidate sample image is the minimum distance; or, the distance between the global descriptor to be tested and the sample global descriptor corresponding to the candidate sample image is less than a distance threshold.

[0038] In one possible embodiment, determining the global descriptor to be tested corresponding to the image to be tested may include, but is not limited to: determining a bag-of-words vector corresponding to the image to be tested based on a trained dictionary model, and determining the bag-of-words vector as the global descriptor to be tested corresponding to the image to be tested; or, inputting the image to be tested into a trained deep learning model to obtain a target vector corresponding to the image to be tested, and determining the target vector as the global descriptor to be tested corresponding to the image to be tested. Of course, the above are only two examples of determining the global descriptor to be tested, and the method for determining the global descriptor to be tested is not limited.

[0039] Exemplarily, determining a target map point corresponding to a feature point from a plurality of map points corresponding to a candidate sample image may include, but is not limited to: determining a local descriptor to be tested corresponding to the feature point, the local descriptor to be tested being used to represent a feature vector of an image block where the feature point is located, and the image block may be located in the image to be tested. Determining the distance between the local descriptor to be tested and a sample local descriptor corresponding to each map point corresponding to the candidate sample image; wherein the three-dimensional visual map includes at least a sample local descriptor corresponding to each map point corresponding to the candidate sample image. Then, based on the distance between the local descriptor to be tested and each sample local descriptor, a target map point may be selected from the plurality of map points corresponding to the candidate sample image; wherein the distance between the local descriptor to be tested and the sample local descriptor corresponding to the target map point may be a minimum distance, and the minimum distance is less than a distance threshold.

[0040] Step 104 : Generate a fused positioning trajectory of the terminal device in the three-dimensional visual map based on the self-positioning trajectory and the global positioning trajectory, and output the fused positioning trajectory, such as displaying the fused positioning trajectory.

[0041] Exemplarily, the frame rate of the fused positioning poses included in the fused positioning trajectory can be greater than the frame rate of the global positioning poses included in the global positioning trajectory. That is, the frame rate of the fused positioning trajectory can be higher than the frame rate of the global positioning trajectory. The fused positioning trajectory can be a high-frame-rate pose in the 3D visual map, and the global positioning trajectory can be a low-frame-rate pose in the 3D visual map. The higher frame rate of the fused positioning trajectory than the global positioning trajectory indicates that the number of fused positioning poses is greater than the number of global positioning poses.

[0042] For example, the frame rate of the fused positioning poses included in the fused positioning trajectory can be equal to the frame rate of the self-positioning poses included in the self-positioning trajectory. In other words, the frame rate of the fused positioning trajectory can be equal to the frame rate of the self-positioning trajectory, that is, the self-positioning trajectory can be a high-frame-rate pose. The frame rate of the fused positioning trajectory being equal to the frame rate of the self-positioning trajectory means that the number of fused positioning poses is equal to the number of self-positioning poses.

[0043] For example, N self-positioning poses corresponding to the target time period can be selected from all self-positioning poses included in the self-positioning trajectory, and P global positioning poses corresponding to the target time period can be selected from all global positioning poses included in the global positioning trajectory; N is greater than P. Based on the N self-positioning poses and the P global positioning poses, N fused positioning poses corresponding to the N self-positioning poses are determined, and the N self-positioning poses correspond one-to-one to the N fused positioning poses. Based on the N fused positioning poses, a fused positioning trajectory of the terminal device in the three-dimensional visual map is generated, and the fused positioning trajectory is a high-frame-rate pose in the three-dimensional visual map.

[0044] Exemplarily, after generating a fused positioning trajectory of the terminal device in the three-dimensional visual map based on the self-positioning trajectory and the global positioning trajectory, an initial fused positioning pose can be selected from the fused positioning trajectory, and an initial self-positioning pose corresponding to the initial fused positioning pose can be selected from the self-positioning trajectory. A target self-positioning pose is selected from the self-positioning trajectory, and a target fused positioning pose is determined based on the initial fused positioning pose, the initial self-positioning pose, and the target self-positioning pose. Then, a new fused positioning trajectory is generated based on the target fused positioning pose and the fused positioning trajectory to replace the original fused positioning trajectory.

[0045] It can be seen from the above technical solutions that in the embodiment of the present application, high-frame-rate self-positioning can be performed based on the target image and motion data to obtain a high-frame-rate self-positioning trajectory, and low-frame-rate global positioning can be performed based on the target image and the three-dimensional visual map to obtain a low-frame-rate global positioning trajectory. Then, the high-frame-rate self-positioning trajectory and the low-frame-rate global positioning trajectory are fused to eliminate the accumulated error of self-positioning, and a high-frame-rate fused positioning trajectory is obtained, that is, a high-frame-rate fused positioning trajectory in the three-dimensional visual map, to achieve a high-frame-rate and high-precision positioning function, and to achieve a globally consistent high-frame-rate positioning function indoors. In the above method, the target scene can be an indoor environment, and a high-precision, low-cost, and easy-to-deploy indoor positioning function can be achieved based on the target image and motion data. It is a vision-based indoor positioning method that can be applied in energy industries such as coal, electricity, and petrochemicals to achieve indoor positioning of personnel (such as workers, patrol personnel, etc.), quickly obtain the location information of personnel, ensure personnel safety, and achieve efficient management of personnel.

[0046] The following describes the posture determination method of the embodiment of the present application in conjunction with specific embodiments.

[0047] In an embodiment of the present application, a method for determining a position and posture is proposed. As a terminal device moves through a target scene, a fused positioning trajectory of the terminal device within a three-dimensional visual map is determined and output. The target scene can be an indoor environment. That is, when the terminal device moves within the indoor environment, the fused positioning trajectory of the terminal device within the three-dimensional visual map is determined, thereby proposing a vision-based indoor positioning method. Of course, the target scene is not limited to indoor environments, and outdoor environments, for example, are not restricted to this.

[0048] See also Figure 2 The figure shows a schematic diagram of the structure of the terminal device, which may include a self-positioning module, a global positioning module and a fusion positioning module. The terminal device may also include a visual sensor and a motion sensor. The visual sensor may be a camera, etc. The visual sensor is used to collect images of the target scene during the movement of the terminal device. For the convenience of distinction, the image is recorded as a target image, and the target image may include multiple frames of images (i.e., multiple frames of real-time images during the movement of the terminal device). The motion sensor may be an IMU (Inertial Measurement Unit), etc. IMU generally refers to a measuring device including a gyroscope and an accelerometer. The motion sensor is used to collect motion data of the terminal device during the movement of the terminal device, for example, motion data such as acceleration and angular velocity.

[0049] For example, the terminal device can be a wearable device (such as a video helmet, smart watch, smart glasses, etc.), and the visual sensor and motion sensor are deployed on the wearable device; or the terminal device can be a recorder (such as a device carried by a worker while performing work, which integrates real-time audio and video acquisition, photography, recording, intercom, positioning, etc.), and the visual sensor and motion sensor are deployed on the recorder; or the terminal device can be a camera (such as a split camera), and the visual sensor and motion sensor are deployed on the camera. Of course, the above are just examples, and there is no limitation on the type of terminal device. For example, it can also be a smartphone, as long as it is equipped with a visual sensor and a motion sensor.

[0050] See also Figure 2 As shown, the self-positioning module can obtain a target image and motion data, perform high-frame-rate self-positioning based on the target image and motion data, obtain a high-frame-rate self-positioning trajectory (such as a 6DOF (six degrees of freedom) self-positioning trajectory), and send the high-frame-rate self-positioning trajectory to the fusion positioning module. Exemplarily, the self-positioning trajectory may include multiple self-positioning poses. Since the self-positioning trajectory is a high-frame-rate self-positioning trajectory, the number of self-positioning poses in the self-positioning trajectory is relatively large.

[0051] The global positioning module can acquire a target image and perform low-frame-rate global positioning based on the target image and the 3D visual map of the target scene, obtaining a low-frame-rate global positioning trajectory (i.e., the global positioning trajectory of the target image in the 3D visual map), and sending the low-frame-rate global positioning trajectory to the fusion positioning module. Exemplarily, the global positioning trajectory may include multiple global positioning poses. Since the global positioning trajectory is a low-frame-rate global positioning trajectory, the number of global positioning poses in the global positioning trajectory is relatively small.

[0052] The fused localization module can obtain a high-frame-rate self-localization trajectory and a low-frame-rate global localization trajectory. It then fuses these high-frame-rate self-localization trajectory and the low-frame-rate global localization trajectory to produce a high-frame-rate fused localization trajectory, i.e., a high-frame-rate fused localization trajectory within the 3D visual map, thereby achieving a high-frame-rate global localization result. The fused localization trajectory can include multiple fused localization poses. Because the fused localization trajectory is a high-frame-rate fused localization trajectory, it contains a relatively large number of fused localization poses.

[0053] In the above embodiments, the posture (such as self-positioning posture, global positioning posture, fusion positioning posture, etc.) can be position and posture, which are generally represented by rotation matrix and translation vector, and there is no limitation on this.

[0054] To sum up, in this embodiment, based on the target image and motion data, a globally unified high-frame-rate visual positioning function can be achieved, and a high-frame-rate fused positioning trajectory (such as 6DOF posture) in the three-dimensional visual map can be obtained. It is a high-frame-rate globally consistent positioning method, which realizes the high-frame-rate, high-precision, low-cost, and easy-to-deploy indoor positioning function of the terminal device, and realizes a globally consistent high-frame-rate positioning function indoors.

[0055] The functions of the self-positioning module, global positioning module and fusion positioning module are explained below.

[0056] 1. Self-positioning module: The self-positioning module is used to obtain a target image of a target scene and motion data of a terminal device, and determine a self-positioning trajectory of the terminal device based on the target image and the motion data.

[0057] The target image may include multiple frames of images. For each frame of image, the self-positioning module determines the self-positioning posture corresponding to the image, that is, multiple frames of image correspond to multiple self-positioning postures. The self-positioning trajectory of the terminal device may include multiple self-positioning postures. It can be understood that the self-positioning trajectory is a collection of multiple self-positioning postures.

[0058] For the first frame of the multi-frame image, the self-positioning module determines the self-positioning pose corresponding to the first frame of the image, and for the second frame of the multi-frame image, the self-positioning module determines the self-positioning pose corresponding to the second frame of the image, and so on. The self-positioning pose corresponding to the first frame of the image can be the coordinate origin of the reference coordinate system (i.e., the self-positioning pose corresponding to the second frame of the image is a pose point in the reference coordinate system, i.e., a pose point relative to the coordinate origin (i.e., the self-positioning pose corresponding to the first frame of the image), the self-positioning pose corresponding to the third frame of the image is a pose point in the reference coordinate system, i.e., a pose point relative to the coordinate origin, and so on. The self-positioning pose corresponding to each frame of the image is a pose point in the reference coordinate system.

[0059] In summary, after obtaining the self-positioning postures corresponding to each frame of image, these self-positioning postures can be combined into a self-positioning trajectory in the reference coordinate system, and the self-positioning trajectory includes these self-positioning postures.

[0060] In one possible implementation, see Figure 3 As shown, the following steps are used to determine the self-positioning trajectory:

[0061] Step 301: Acquire a target image of a target scene and motion data of a terminal device.

[0062] Step 302: If the target image includes multiple frame images, traverse the multiple frame images to find the current frame image.

[0063] When the first frame image is traversed from multiple frames as the current frame image, the self-positioning posture corresponding to the first frame image can be the coordinate origin of the reference coordinate system (i.e., the self-positioning coordinate system), that is, the self-positioning posture coincides with the coordinate origin. When the second frame image is traversed from multiple frames as the current frame image, the self-positioning posture corresponding to the second frame image can be determined using subsequent steps. When the third frame image is traversed from multiple frames as the current frame image, the self-positioning posture corresponding to the third frame image can be determined using subsequent steps, and so on, each frame image can be traversed as the current frame image.

[0064] Step 303: Calculate the feature point correlation between the current frame image and the previous frame image using an optical flow algorithm. The optical flow algorithm uses the temporal changes in pixels in the current frame image and the correlation between pixels in the previous frame image to find the corresponding relationship between the current frame image and the previous frame image, thereby calculating the motion information of the object between the current frame image and the previous frame image.

[0065] Step 304: Determine whether the current frame image is a key image based on the number of matching feature points between the current frame image and the previous frame image. For example, if the number of matching feature points between the current frame image and the previous frame image does not reach a preset threshold, it indicates that the change between the current frame image and the previous frame image is large, resulting in a relatively small number of matching feature points between the two frames. In this case, the current frame image is determined to be a key image, and step 305 is executed. If the number of matching feature points between the current frame image and the previous frame image reaches a preset threshold, it indicates that the change between the current frame image and the previous frame image is small, resulting in a relatively large number of matching feature points between the two frames. In this case, the current frame image is determined to be a non-key image, and step 306 is executed.

[0066] For example, a matching ratio between the current frame image and the previous frame image may be calculated based on the number of matching feature points between the current frame image and the previous frame image, for example, the ratio of the number of matching feature points to the total number of feature points. If the matching ratio does not reach a preset ratio, the current frame image is determined to be a key image; if the matching ratio reaches the preset ratio, the current frame image is determined to be a non-key image.

[0067] Step 305: If the current frame image is a key image, a map position in a self-positioning coordinate system (i.e., a reference coordinate system) is generated based on the current position of the terminal device (i.e., the position of the terminal device when the current frame image was captured), i.e., a new 3D map position is generated. If the current frame image is a non-key image, there is no need to generate a map position in a self-positioning coordinate system based on the current position of the terminal device.

[0068] Step 306: Determine the self-positioning posture corresponding to the current frame image based on the self-positioning posture corresponding to the K frame images preceding the current frame image, the map position of the terminal device in the self-positioning coordinate system, and the motion data of the terminal device. K can be a positive integer or a value configured based on experience, and there is no restriction on this.

[0069] For example, all motion data between the previous frame and the current frame can be pre-integrated to obtain the inertial measurement constraints between the two frames. Based on the self-localization pose and motion data (such as velocity, acceleration, angular velocity, etc.) corresponding to the K frames (such as a sliding window) preceding the current frame, the map position in the self-localization coordinate system, and the inertial measurement constraints (such as velocity, acceleration, angular velocity, etc. between the previous and current frames), a bundled optimization can be used to perform a joint optimization update to obtain the self-localization pose corresponding to the current frame. There are no restrictions on this bundled optimization process.

[0070] For example, in order to maintain the scale of the variable to be optimized, a certain frame and part of the map position within the sliding window can be marginalized, and these constraint information can be retained in a priori form.

[0071] Exemplarily, the self-positioning module can use a VIO (Visual Inertial Odometry) algorithm to determine the self-positioning pose. That is, the input data of the VIO algorithm is the target image and motion data, and the output data of the VIO algorithm is the self-positioning pose. For example, based on the target image and motion data, the VIO algorithm can obtain the self-positioning pose. For example, the VIO algorithm is used to execute steps 301-306 to obtain the self-positioning pose. The VIO algorithm may include but is not limited to VINS (Visual Inertial Navigation Systems), SVO (Semi-direct Visual Odometry), MSCKF (Multi State Constraint Kalman Filter), etc., and is not limited here, as long as the self-positioning pose can be obtained.

[0072] Step 307: Generate a self-positioning trajectory of the terminal device in the self-positioning coordinate system based on the self-positioning postures corresponding to the multiple frames of images, where the self-positioning trajectory includes multiple self-positioning postures in the self-positioning coordinate system.

[0073] At this point, the self-positioning module can obtain the self-positioning trajectory in the self-positioning coordinate system, which can include the self-positioning postures corresponding to multiple frames of images. Since the visual sensor can collect a large number of images, the self-positioning module can obtain the self-positioning postures corresponding to these images, that is, the self-positioning trajectory can include a large number of self-positioning postures, that is, the self-positioning module can obtain a self-positioning trajectory with a high frame rate.

[0074] 2. Global Positioning Module. Based on the acquired 3D visual map of the target scene, the global positioning module determines the target map points corresponding to the target image from the 3D visual map of the target scene after obtaining the target image. Based on the target map points, the global positioning trajectory of the terminal device in the 3D visual map is determined.

[0075] The target image can include multiple frames, and a portion of the images can be selected from the multiple frames as the images to be tested. In the following, M frames are selected as the images to be tested, where M is a positive integer. For each frame of the image to be tested, the global positioning module can determine the global positioning pose corresponding to the image to be tested. That is, M frames of the image to be tested correspond to M global positioning poses. The global positioning trajectory of the terminal device in the three-dimensional visual map can include M global positioning poses. It can be understood that the global positioning trajectory is a collection of M global positioning poses.

[0076] For the first of the M frames to be tested, the global positioning module determines the global positioning pose corresponding to the first frame to be tested. For the second frame to be tested, the global positioning module determines the global positioning pose corresponding to the second frame to be tested, and so on. For each global positioning pose, the global positioning pose is a pose point in the 3D visual map, that is, a pose point in the 3D visual map coordinate system.

[0077] In summary, after obtaining the global positioning poses corresponding to the M frames of images to be tested, these global positioning poses are combined into a global positioning trajectory in the three-dimensional visual map, and the global positioning trajectory includes the global positioning poses.

[0078] In one possible implementation, it is necessary to pre-build a three-dimensional visual map of the target scene and store the three-dimensional visual map on a server. When the terminal device needs to move in the target scene, the terminal device can download the three-dimensional visual map of the target scene from the server and store the three-dimensional visual map. In this way, during the movement of the terminal device, the global positioning trajectory of the terminal device in the three-dimensional visual map can be determined based on the three-dimensional visual map. The three-dimensional visual map is a storage method for the image information of the target scene. Multiple frames of sample images of the target scene can be collected, and a three-dimensional visual map can be constructed based on these sample images. For example, based on multiple frames of sample images of the target scene, a visual mapping algorithm such as SFM (Structure From Motion) or SLAM (Simultaneous Localization And Mapping) is used to construct a three-dimensional visual map of the target scene. There is no restriction on this construction method.

[0079] After obtaining the 3D visual map of the target scene, the 3D visual map may include the following information:

[0080] Sample image pose: The sample image is a representative image when constructing a three-dimensional visual map, that is, a three-dimensional visual map can be constructed based on the sample image, and the pose matrix of the sample image (referred to as the sample image pose) can be stored in the three-dimensional visual map, that is, the three-dimensional visual map includes the sample image pose.

[0081] Sample global descriptor: For each frame of sample image, the sample image can correspond to an image global descriptor, which is recorded as the sample global descriptor. The sample global descriptor uses a high-dimensional vector to represent the sample image. The sample global descriptor is used to distinguish the image features of different sample images.

[0082] For each sample image, a bag-of-words vector corresponding to the sample image can be determined based on the trained dictionary model, and the bag-of-words vector can be determined as the sample global descriptor corresponding to the sample image. For example, the bag-of-words method is a method for determining a global descriptor. In the bag-of-words method, a bag-of-words vector can be constructed. The bag-of-words vector is a vector representation method used for image similarity detection, and the bag-of-words vector can be used as the sample global descriptor corresponding to the sample image.

[0083] In the bag-of-visual-words method, a "dictionary" needs to be trained in advance, also called a dictionary model. Generally, feature point descriptors in a large number of images are clustered to obtain a classification tree. Each classification tree can represent a visual "word", and these visual "words" constitute the dictionary model.

[0084] For a sample image, all feature point descriptors in the sample image can be classified into "words" and the frequency of occurrence of all words can be counted. In this way, the frequency of each word in the dictionary can form a vector, which is the bag-of-words vector corresponding to the sample image. The bag-of-words vector can be used to measure the similarity between two images and the bag-of-words vector is used as the sample global descriptor corresponding to the sample image.

[0085] For each sample image frame, the sample image can be input into a trained deep learning model to obtain a target vector corresponding to the sample image, and the target vector is determined as the sample global descriptor corresponding to the sample image. For example, a deep learning method is a method for determining a global descriptor. In a deep learning method, a sample image can be subjected to multi-layer convolution by a deep learning model to ultimately obtain a high-dimensional target vector, which is used as the sample global descriptor corresponding to the sample image.

[0086] Deep learning methods require pre-training of deep learning models, such as CNN (Convolutional Neural Networks) models. These models are typically trained using a large number of images, and there are no restrictions on the training method. For a sample image, the sample image can be input into the deep learning model, which processes the sample image to obtain a high-dimensional target vector, which is used as the sample global descriptor corresponding to the sample image.

[0087] Sample local descriptors corresponding to feature points of the sample image: For each frame of the sample image, the sample image may include multiple feature points. The feature point may be a specific pixel position in the sample image. The feature point may correspond to an image local descriptor, which is recorded as a sample local descriptor. The sample local descriptor uses a vector to describe the features of the image block in the vicinity of the feature point (i.e., pixel position). The vector may also be called the descriptor of the feature point. In summary, the sample local descriptor is a feature vector used to represent the image block where the feature point is located, and the image block may be located in the sample image. It should be noted that the feature points of the sample image (i.e., two-dimensional feature points) correspond to the map points in the three-dimensional visual map (i.e., three-dimensional map points). Therefore, the sample local descriptor corresponding to the feature point of the sample image is the sample local descriptor corresponding to the map point corresponding to the feature point.

[0088] Among them, algorithms such as ORB (Oriented FAST and Rotated BRIEF), SIFT (Scale-Invariant Feature Transform), and SURF (Speeded Up Robust Features) can be used to extract feature points from the sample image and determine the sample local descriptors corresponding to the feature points. Deep learning algorithms (such as SuperPoint, DELF, D2-Net, etc.) can also be used to extract feature points from the sample image and determine the sample local descriptors corresponding to the feature points. There is no restriction on this, as long as feature points can be obtained and the sample local descriptors can be determined.

[0089] Map point information (i.e., feature point information): Map point information may include, but is not limited to: the 3D spatial position of the map point, all observed sample images, and the corresponding 2D feature point number.

[0090] Based on the three-dimensional visual map of the target scene, in one possible implementation, see Figure 4 As shown, the global positioning module uses the following steps to determine the global positioning trajectory of the terminal device in the three-dimensional visual map:

[0091] Step 401: Acquire a target image of a target scene. If the target image includes multiple frames of images, select M frames of images from the multiple frames of images as images to be tested, that is, the images to be tested are M frames, and M can be a positive integer.

[0092] For example, referring to step 304, the multiple frames of images include key images and non-key images. On this basis, the key images in the multiple frames of images are used as images to be tested, while the non-key images are not used as images to be tested.

[0093] For another example, images to be tested can be selected from multiple frames of images at a fixed interval. Assuming that the fixed interval is 5 (of course, the fixed interval can be arbitrarily configured based on experience and there is no restriction on this), the first frame of image can be used as the image to be tested, the sixth (1+5) frame of image can be used as the image to be tested, the eleventh (6+5) frame of image can be used as the image to be tested, and so on. A frame of image to be tested is selected every five frames of image.

[0094] Of course, the above-mentioned methods of selecting images to be tested are only two examples. As long as some images can be selected from multiple frames of images as images to be tested, there is no limitation on the method of selecting images to be tested.

[0095] Step 402: for each frame of the image to be tested, determine the global descriptor to be tested corresponding to the image to be tested.

[0096] Exemplarily, for each frame of the image to be tested, the image to be tested may correspond to an image global descriptor, and the image global descriptor may be recorded as the global descriptor to be tested. The global descriptor to be tested uses a high-dimensional vector to represent the image to be tested, and the global descriptor to be tested is used to distinguish the image features of different images to be tested.

[0097] For each image under test, a bag-of-words vector corresponding to the image under test is determined based on a trained dictionary model, and the bag-of-words vector is determined as the global descriptor under test corresponding to the image under test. Alternatively, for each image under test, the image under test is input into a trained deep learning model to obtain a target vector corresponding to the image under test, and the target vector is determined as the global descriptor under test corresponding to the image under test.

[0098] In summary, the global descriptor to be tested corresponding to the image to be tested can be determined based on the bag-of-visual-words method or the deep learning method. The determination method refers to the method for determining the sample global descriptor, which will not be repeated here.

[0099] Step 403: For each frame of the image to be tested, determine the similarity between the global descriptor to be tested corresponding to the image to be tested and the sample global descriptor corresponding to each frame of the sample image corresponding to the three-dimensional visual map.

[0100] Referring to the above embodiment, the three-dimensional visual map may include a sample global descriptor corresponding to each frame of the sample image. Therefore, the similarity between the global descriptor to be tested and each sample global descriptor can be determined. Taking the similarity as "distance similarity" as an example, the distance between the global descriptor to be tested and each sample global descriptor can be determined, such as the Euclidean distance, that is, the Euclidean distance between the two feature vectors is calculated.

[0101] Step 404: Based on the distance between the global descriptor to be tested and each sample global descriptor, a candidate sample image is selected from the multiple frames of sample images corresponding to the three-dimensional visual map; wherein the distance between the global descriptor to be tested and the sample global descriptor corresponding to the candidate sample image is a minimum distance; or, the distance between the global descriptor to be tested and the sample global descriptor corresponding to the candidate sample image is less than a distance threshold.

[0102] For example, assuming that the three-dimensional visual map corresponds to sample image 1, sample image 2 and sample image 3, the distance 1 between the global descriptor to be tested and the sample global descriptor corresponding to sample image 1 can be calculated, and the distance 2 between the global descriptor to be tested and the sample global descriptor corresponding to sample image 2 can be calculated, and the distance 3 between the global descriptor to be tested and the sample global descriptor corresponding to sample image 3 can be calculated.

[0103] In one possible implementation, if distance 1 is the minimum distance, sample image 1 is selected as the candidate sample image. Alternatively, if distance 1 is less than a distance threshold (which can be configured based on experience), and distance 2 is less than the distance threshold, but distance 3 is not less than the distance threshold, both sample image 1 and sample image 2 are selected as candidate sample images. Alternatively, if distance 1 is the minimum distance and less than the distance threshold, sample image 1 is selected as the candidate sample image. However, if distance 1 is the minimum distance and not less than the distance threshold, no candidate sample image can be selected, i.e., relocalization fails.

[0104] In summary, for each frame of the image to be tested, a candidate sample image corresponding to the image to be tested is selected from multiple frames of sample images corresponding to the three-dimensional visual map, and the number of candidate sample images may be at least one.

[0105] Step 405: For each frame of the image to be tested, multiple feature points are obtained from the image to be tested. For each feature point, a local descriptor to be tested corresponding to the feature point is determined. The local descriptor to be tested is used to represent the feature vector of the image block where the feature point is located, and the image block can be located in the image to be tested.

[0106] For example, the image to be tested may include multiple feature points. A feature point may be a specific pixel location in the image to be tested. The feature point may correspond to a local image descriptor, which is recorded as the local descriptor to be tested. The local descriptor to be tested uses a vector to describe the characteristics of the image block within the vicinity of the feature point (i.e., the pixel location). This vector can also be called the descriptor of the feature point. In summary, the local descriptor to be tested is a feature vector used to represent the image block where the feature point is located.

[0107] Among them, ORB, SIFT, SURF and other algorithms can be used to extract feature points from the image to be tested and determine the local descriptors to be tested corresponding to the feature points. Deep learning algorithms (such as SuperPoint, DELF, D2-Net, etc.) can also be used to extract feature points from the image to be tested and determine the local descriptors to be tested corresponding to the feature points. There is no limitation on this, as long as the feature points can be obtained and the local descriptors to be tested can be determined.

[0108] Step 406: For each feature point corresponding to the image to be tested, determine the distance between the local descriptor to be tested corresponding to the feature point and the sample local descriptor corresponding to each map point corresponding to the candidate sample image corresponding to the image to be tested, such as the Euclidean distance, that is, calculate the Euclidean distance between the two feature vectors.

[0109] Referring to the above embodiment, for each sample image frame, the three-dimensional visual map includes a sample local descriptor corresponding to each map point corresponding to the sample image (i.e., a sample local descriptor corresponding to each map point corresponding to each feature point in the sample image). Therefore, after obtaining a candidate sample image corresponding to the image to be tested, the sample local descriptor corresponding to each map point corresponding to the candidate sample image can be obtained from the three-dimensional visual map. After obtaining each feature point corresponding to the image to be tested, the distance between the local descriptor corresponding to the feature point to be tested and the sample local descriptor corresponding to each map point corresponding to the candidate sample image is determined.

[0110] Step 407: For each feature point, based on the distance between the local descriptor to be tested corresponding to the feature point and the sample local descriptor corresponding to each map point corresponding to the candidate sample image, select a target map point from the multiple map points corresponding to the candidate sample image; wherein the distance between the local descriptor to be tested and the sample local descriptor corresponding to the target map point is the minimum distance, and the minimum distance is less than the distance threshold.

[0111] For example, assuming that the candidate sample image corresponds to map point 1, map point 2, and map point 3, the distance 1 between the local descriptor to be tested corresponding to the feature point and the sample local descriptor corresponding to map point 1 can be calculated, and the distance 2 between the local descriptor to be tested and the sample local descriptor corresponding to map point 2 can be calculated, and the distance 3 between the local descriptor to be tested and the sample local descriptor corresponding to map point 3 can be calculated.

[0112] In one possible implementation, if Distance 1 is the minimum distance, then Map Point 1 can be selected as the target map point. Alternatively, if Distance 1 is less than a distance threshold (which can be configured based on experience), and Distance 2 is less than the distance threshold, but Distance 3 is not less than the distance threshold, then both Map Point 1 and Map Point 2 can be selected as the target map point. Alternatively, if Distance 1 is the minimum distance and Distance 1 is less than the distance threshold, then Map Point 1 can be selected as the target map point. However, if Distance 1 is the minimum distance and Distance 1 is not less than the distance threshold, then no target map point can be selected, i.e., relocalization fails.

[0113] In summary, for each feature point of the image to be tested, a target map point corresponding to the feature point is selected from the candidate sample images corresponding to the image to be tested, and a matching relationship between the feature point and the target map point is obtained.

[0114] Step 408: Determine a global positioning pose in the three-dimensional visual map corresponding to the image to be tested based on the multiple feature points corresponding to the image to be tested and the target map points corresponding to the multiple feature points.

[0115] For a frame of the image to be tested, the image to be tested can correspond to multiple feature points, and each feature point corresponds to a target map point. For example, the target map point corresponding to feature point 1 is map point 1, the target map point corresponding to feature point 2 is map point 2, and so on, thereby obtaining multiple matching relationship pairs, each matching relationship pair includes a feature point (i.e., a two-dimensional feature point) and a map point (i.e., a three-dimensional map point in a three-dimensional visual map). The feature point represents the two-dimensional position in the image to be tested, and the map point represents the three-dimensional position in the three-dimensional visual map, that is, the matching relationship pair includes a mapping relationship from a two-dimensional position to a three-dimensional position, that is, a mapping relationship from a two-dimensional position in the image to be tested to a three-dimensional position in the three-dimensional visual map.

[0116] If the total number of multiple matching relationship pairs does not meet the quantity requirement, it means that the global positioning pose in the three-dimensional visual map corresponding to the image to be tested cannot be determined based on the multiple matching relationship pairs. If the total number of multiple matching relationship pairs meets the quantity requirement (that is, the total number reaches the preset quantity value), it means that the global positioning pose in the three-dimensional visual map corresponding to the image to be tested can be determined based on the multiple matching relationship pairs, that is, the global positioning pose in the three-dimensional visual map corresponding to the image to be tested can be determined based on the multiple matching relationship pairs. For example, the PnP (Perspective N Point, n-point perspective) algorithm is used to calculate the global positioning pose of the image to be tested in the three-dimensional visual map, and there is no restriction on this calculation method. For example, the input data of the PnP algorithm is multiple matching relationship pairs. For each matching relationship pair, the matching relationship pair includes a two-dimensional position in the image to be tested and a three-dimensional position in the three-dimensional visual map. Based on multiple matching relationship pairs, the PnP algorithm can be used to calculate the pose of the image to be tested in the three-dimensional visual map, that is, the global positioning pose.

[0117] In summary, for each frame of the image to be tested, the global positioning pose in the three-dimensional visual map corresponding to the image to be tested can be obtained, that is, the global positioning pose of the image to be tested in the three-dimensional visual map coordinate system can be obtained.

[0118] In one possible implementation, after obtaining multiple matching pairs, valid matching pairs can be found from the multiple matching pairs. Based on these valid matching pairs, a PnP algorithm can be used to calculate the global positioning pose of the image to be tested in the three-dimensional visual map. For example, a RANSAC (RANdom SAmple Consensus) detection algorithm can be used to find valid matching pairs from all matching pairs, and there is no restriction on this process.

[0119] Step 409: Generate a global positioning trajectory of the terminal device in the three-dimensional visual map based on the global positioning poses corresponding to the M frames of images to be tested, and the global positioning trajectory includes multiple global positioning poses in the three-dimensional visual map. At this point, the global positioning module can obtain the global positioning trajectory in the three-dimensional visual map, that is, the global positioning trajectory in the three-dimensional visual map coordinate system. The global positioning trajectory can include the global positioning poses corresponding to the M frames of images to be tested, that is, the global positioning trajectory can include M global positioning poses. Since the M frames of images to be tested are part of the images selected from all the images, the global positioning trajectory can include a small number of global positioning poses, that is, the global positioning module obtains a global positioning trajectory with a low frame rate.

[0120] 3. Fusion positioning module. A high-frame-rate self-positioning trajectory can be obtained from the self-positioning module, and a low-frame-rate global positioning trajectory can be obtained from the global positioning module. The high-frame-rate self-positioning trajectory and the low-frame-rate global positioning trajectory are fused to obtain a high-frame-rate fused positioning trajectory in the three-dimensional visual map coordinate system, that is, the fused positioning trajectory of the terminal device in the three-dimensional visual map, and the fused positioning trajectory is output. Among them, the fused positioning trajectory is a high-frame-rate pose in the three-dimensional visual map, and the global positioning trajectory is a low-frame-rate pose in the three-dimensional visual map, that is, the frame rate of the fused positioning trajectory is higher than the frame rate of the global positioning trajectory, and the number of fused positioning poses included in the fused positioning trajectory is greater than the number of global positioning poses included in the global positioning trajectory.

[0121] See also Figure 5 As shown in the figure, the white solid circle represents the self-positioning posture, and the trajectory composed of multiple self-positioning postures is called a self-positioning trajectory, that is, the self-positioning trajectory includes multiple self-positioning postures. The self-positioning posture corresponding to the first frame image can be the reference coordinate system S L The coordinate origin of the self-positioning coordinate system is recorded as Self-positioning pose With reference coordinate system S L For each self-positioning posture in the self-positioning trajectory, in the reference coordinate system S L The self-localization pose under .

[0122] The gray solid circle represents the global positioning pose. The trajectory composed of multiple global positioning poses is called the global positioning trajectory, that is, the global positioning trajectory includes multiple global positioning poses. The global positioning pose can be a three-dimensional visual map coordinate system S G The pose under the global positioning trajectory is the three-dimensional visual map coordinate system S G The global positioning pose under , that is, the global positioning pose under the 3D visual map.

[0123] The white dotted circle represents the fused positioning pose. The trajectory composed of multiple fused positioning poses is called the fused positioning trajectory, that is, the fused positioning trajectory includes multiple fused positioning poses. The fused positioning pose can be the 3D visual map coordinate system S G The pose under the fusion positioning trajectory is the three-dimensional visual map coordinate system S G The fused positioning pose under , that is, the fused positioning pose under the 3D visual map.

[0124] See also Figure 5As shown in the figure, since the target image includes multiple frames, each frame corresponds to a self-localization pose, and some images are selected from the multiple frames as test images, each test frame corresponds to a global localization pose. Therefore, the number of self-localization poses is greater than the number of global localization poses. When a fused localization trajectory is obtained based on the self-localization trajectory and the global localization trajectory, each self-localization pose corresponds to a fused localization pose (i.e., the self-localization pose and the fused localization pose have a one-to-one correspondence). In other words, the number of self-localization poses is the same as the number of fused localization poses. Therefore, the number of fused localization poses is also greater than the number of global localization poses.

[0125] In a possible implementation, the fusion positioning module can realize the trajectory fusion function and the posture transformation function, see Figure 6 As shown, the fusion positioning module can implement the trajectory fusion function and posture transformation function by using the following steps to obtain the fusion positioning trajectory of the terminal device in the three-dimensional visual map:

[0126] Step 601: Select N self-positioning poses corresponding to the target time period from all self-positioning poses included in the self-positioning trajectory, and select P global positioning poses corresponding to the target time period from all global positioning poses included in the global positioning trajectory. Exemplarily, N can be greater than P.

[0127] For example, when fusing the self-localization trajectory and the global localization trajectory of the target time period, N self-localization poses corresponding to the target time period (i.e., the self-localization poses determined based on the images collected during the target time period) and P global localization poses corresponding to the target time period (i.e., the global localization poses determined based on the images collected during the target time period) can be determined. Figure 5 As shown, you can and The self-positioning poses between are taken as the N self-positioning poses corresponding to the target time period, and the and The global positioning poses between are taken as the P global positioning poses corresponding to the target time period.

[0128] Step 602: Determine N fused positioning poses corresponding to the N self-positioning poses based on the N self-positioning poses and the P global positioning poses, where the N self-positioning poses correspond to the N fused positioning poses in a one-to-one manner.

[0129] For example, see Figure 5 As shown, the self-positioning pose can be determined based on N self-positioning poses and P global positioning poses. Corresponding fusion positioning pose Determine self-positioning pose Corresponding fusion positioning pose Determine self-positioning pose Corresponding fusion positioning pose And so on.

[0130] In a possible implementation, assume that there are N self-positioning poses, P global positioning poses, and N fusion positioning poses. The N self-positioning poses are all known values, the P global positioning poses are all known values, and the N fusion positioning poses are all unknown values, which are the pose values ​​that need to be solved. Figure 5 As shown, the self-positioning posture and fusion positioning pose Corresponding, self-positioning pose and fusion positioning pose Corresponding, self-positioning pose and fusion positioning pose Corresponding, and so on. Global positioning pose and fusion positioning pose Time response, global positioning pose and fusion positioning pose Corresponding, and so on.

[0131] The first constraint value can be determined based on N self-positioning poses and N fusion positioning poses. The first constraint value is used to represent the residual value between the fusion positioning pose and the self-positioning pose. For example, and The difference, and The difference, ..., and The calculation formula of the first constraint value is not limited in this embodiment and can be related to the above-mentioned differences.

[0132] The second constraint value can be determined based on P global positioning poses and P fused positioning poses (i.e., P fused positioning poses corresponding to P global positioning poses are selected from N fused positioning poses). The second constraint value is used to represent the residual value (i.e., absolute difference) between the fused positioning pose and the global positioning pose. For example, it can be based on and The difference, ..., and The second constraint value is calculated based on the difference between the two values. The calculation formula of the second constraint value is not limited in this embodiment and can be related to the above-mentioned differences.

[0133] The target constraint value can be calculated based on the first and second constraint values. For example, the target constraint value can be the sum of the first and second constraint values. Since the N self-localization poses and the P global localization poses are known values, while the N fused localization poses are unknown values, the target constraint value is minimized by adjusting the values ​​of the N fused localization poses. When the target constraint value is minimized, the values ​​of the N fused localization poses become the final solved pose value, thus obtaining the values ​​of the N fused localization poses.

[0134] In one possible implementation, the target constraint value may be calculated using formula (1):

[0135]

[0136] In formula (1), F(T) represents the target constraint value, the part before the plus sign (hereinafter referred to as the first part) is the first constraint value, and the part after the plus sign (hereinafter referred to as the second part) is the second constraint value. Of course, the above are only examples of the target constraint value, the first constraint value, and the second constraint value, and there is no limitation on this.

[0137] Ω i,i+1 It is the residual information matrix for the self-positioning posture, which can be configured according to experience and is not restricted. k It is the residual information matrix for the global positioning pose, which can be configured based on experience and there is no restriction on it.

[0138] The first part represents the relative transformation constraint between the self-localization pose and the fused localization pose, which can be reflected by the first constraint value. N is all the self-localization poses in the self-localization trajectory, that is, N self-localization poses. The second part represents the global positioning constraint between the global positioning pose and the fused localization pose, which can be reflected by the second constraint value. P is all the global positioning poses in the global positioning trajectory, that is, P global positioning poses.

[0139] For the first and second parts, they can also be expressed by formula (2) and formula (3):

[0140]

[0141]

[0142] In formula (2) and formula (3), and For fused positioning poses (without corresponding global positioning poses), and is the self-positioning pose, is the relative pose change constraint between two self-localization poses, e i,i+1 for and Relative posture changes and Constrained residuals.

[0143] For fusion positioning pose (with corresponding global positioning pose ), for The corresponding global positioning pose, e k Represents the fusion positioning pose Relative to the global positioning pose The residual.

[0144] Since the self-localization pose and the global localization pose are known, and the fused localization pose is unknown, the optimization goal can be to minimize the value of F(T), so as to obtain the fused localization pose. That is, the fused localization trajectory in the three-dimensional visual map coordinate system can be referred to in formula (4): arg min F(T). By minimizing the value of F(T), the fused localization trajectory can be obtained, and the fused localization trajectory can include multiple fused localization poses.

[0145] For example, in order to minimize the value of F(T), algorithms such as Gauss-Newton, gradient descent, and LM (Levenberg-Marquardt) can be used to obtain the fused positioning pose, which will not be described in detail here.

[0146] Step 603: Generate a fused positioning trajectory of the terminal device in the 3D visual map based on the N fused positioning poses. The fused positioning trajectory includes the N fused positioning poses in the 3D visual map. At this point, the fused positioning module can obtain a fused positioning trajectory in the 3D visual map, i.e., a fused positioning trajectory in the 3D visual map coordinate system. The number of fused positioning poses in the fused positioning trajectory is greater than the number of global positioning poses in the global positioning trajectory, which means that a fused positioning trajectory with a high frame rate can be obtained.

[0147] Step 604: Select an initial fused positioning pose from the fused positioning trajectory, and select an initial self-positioning pose corresponding to the initial fused positioning pose from the self-positioning trajectory.

[0148] Step 605: Select a target self-positioning pose from the self-positioning trajectory, and determine a target fused positioning pose based on the initial fused positioning pose, the initial self-positioning pose, and the target self-positioning pose.

[0149] Exemplarily, after generating the fused positioning trajectory, the fused positioning trajectory can also be updated. During the trajectory update process, an initial fused positioning pose can be selected from the fused positioning trajectory, an initial self-positioning pose can be selected from the self-positioning trajectory, and a target self-positioning pose can be selected from the self-positioning trajectory. On this basis, a target fused positioning pose can be determined based on the initial fused positioning pose, the initial self-positioning pose, and the target self-positioning pose. Then, a new fused positioning trajectory can be generated based on the target fused positioning pose and the fused positioning trajectory to replace the original fused positioning trajectory.

[0150] For example, in steps 601 to 603, see Figure 5 As shown, the self-positioning trajectory includes and The self-localization pose between , the global positioning trajectory includes and The global positioning pose between them, the fusion positioning trajectory includes and The fusion positioning pose between them, after which, if a new self-positioning pose is obtained However, since there is no corresponding global positioning pose, it is impossible to use the global positioning pose and the self-positioning pose to Determine the self-positioning pose Corresponding fusion positioning pose On this basis, in this embodiment, the following formula (4) can also be used to determine the fusion positioning posture

[0151]

[0152] In formula (4), Represents the self-positioning pose The corresponding fusion positioning pose, that is, the target fusion positioning pose, represents the fusion positioning pose, that is, the initial fusion positioning pose selected from the fusion positioning trajectory, Represents the self-positioning pose, that is, the pose selected from the self-positioning trajectory The corresponding initial self-positioning pose, Represents the self-positioning pose, that is, the target self-positioning pose selected from the self-positioning trajectory. In summary, it can be seen that the initial fusion positioning pose can be The initial self-positioning pose and the target's self-localization pose Determine the target fusion positioning pose

[0153] After obtaining the target fusion positioning pose Afterwards, a new fusion positioning trajectory can be generated, that is, the new fusion positioning trajectory can include the target fusion positioning pose Thereby updating the fused positioning trajectory.

[0154] In the above process, steps 601-603 are the trajectory fusion process, and steps 604-605 are the pose transformation process. Trajectory fusion is the process of aligning and fusing the self-localization trajectory with the global positioning trajectory, realizing the conversion of the self-localization trajectory from the self-localization coordinate system to the 3D visual map coordinate system. The trajectory is corrected using the global positioning result. When a new frame can obtain the global positioning trajectory, trajectory fusion is performed again. Since not all frames can successfully obtain the global positioning trajectory, the pose of these frames is output as the fused positioning pose in the 3D visual map coordinate system through pose transformation, i.e., the pose transformation process.

[0155] It can be seen from the above technical solutions that in the embodiment of the present application, high-frame-rate self-positioning can be performed based on the target image and motion data to obtain a high-frame-rate self-positioning trajectory, and low-frame-rate global positioning can be performed based on the target image and the three-dimensional visual map to obtain a low-frame-rate global positioning trajectory. Then, the high-frame-rate self-positioning trajectory and the low-frame-rate global positioning trajectory are fused to eliminate the accumulated error of self-positioning, and a high-frame-rate fused positioning trajectory is obtained, that is, a high-frame-rate fused positioning trajectory in the three-dimensional visual map, to achieve a high-frame-rate and high-precision positioning function, and to achieve a globally consistent high-frame-rate positioning function indoors. In the above method, the target scene can be an indoor environment, and a high-precision, low-cost, and easy-to-deploy indoor positioning function can be achieved based on the target image and motion data. It is a vision-based indoor positioning method that can be applied in energy industries such as coal, electricity, and petrochemicals to achieve indoor positioning of personnel (such as workers, patrol personnel, etc.), quickly obtain the location information of personnel, ensure personnel safety, and achieve efficient management of personnel.

[0156] Based on the same application concept as the above method, an embodiment of the present application proposes a posture determination device, which is applied to a terminal device, wherein the terminal device includes a three-dimensional visual map of the target scene. When the terminal device moves in the target scene, see Figure 7 FIG. 1 is a structural diagram of the device, which includes:

[0157] An acquisition module 71 is used to acquire a target image of the target scene and motion data of the terminal device; a determination module 72 is used to determine a self-positioning trajectory of the terminal device based on the target image and the motion data; a target map point corresponding to the target image is determined from the three-dimensional visual map, and a global positioning trajectory of the terminal device in the three-dimensional visual map is determined based on the target map point; a generation module 73 is used to generate a fused positioning trajectory of the terminal device in the three-dimensional visual map based on the self-positioning trajectory and the global positioning trajectory, and output the fused positioning trajectory; the frame rate of the fused positioning posture included in the fused positioning trajectory is greater than the frame rate of the global positioning posture included in the global positioning trajectory.

[0158] Exemplarily, when determining the self-positioning trajectory of the terminal device based on the target image and the motion data, the determination module 72 is specifically used to: if the target image includes multiple frame images, traverse the current frame image from the multiple frame images; determine the self-positioning posture corresponding to the current frame image based on the self-positioning posture corresponding to the K frame images preceding the current frame image, the map position of the terminal device in the self-positioning coordinate system and the motion data; generate the self-positioning trajectory of the terminal device in the self-positioning coordinate system based on the self-positioning posture corresponding to the multiple frame images; if the current frame image is a key image, generate the map position in the self-positioning coordinate system based on the current position of the terminal device; if the number of matching feature points between the current frame image and the frame image before the current frame image does not reach a preset threshold, determine that the current frame image is a key image.

[0159] Exemplarily, the determination module 72 determines the target map point corresponding to the target image from the three-dimensional visual map, and determines the global positioning trajectory of the terminal device in the three-dimensional visual map based on the target map point, which is specifically used for: if the target image includes multiple frames of images, selecting M frames of images from the multiple frames of images as the images to be tested; for each frame of the image to be tested, based on the similarity between the image to be tested and the multiple frames of sample images corresponding to the three-dimensional visual map, selecting a candidate sample image from the multiple frames of sample images; obtaining multiple feature points from the image to be tested; for each feature point, determining the target map point corresponding to the feature point from the multiple map points corresponding to the candidate sample images; determining the global positioning pose in the three-dimensional visual map corresponding to the image to be tested based on the multiple feature points and the target map points corresponding to the multiple feature points; and generating the global positioning trajectory of the terminal device in the three-dimensional visual map based on the global positioning poses corresponding to the M frames of the images to be tested.

[0160] Exemplarily, the determination module 72 is specifically used to select candidate sample images from multiple frames of sample images based on the similarity between the image to be tested and the multiple frames of sample images corresponding to the three-dimensional visual map: determine the global descriptor to be tested corresponding to the image to be tested, and determine the distance between the global descriptor to be tested and the sample global descriptor corresponding to each frame of sample image corresponding to the three-dimensional visual map; wherein the three-dimensional visual map includes a sample global descriptor corresponding to each frame of sample image; based on the distance between the global descriptor to be tested and each sample global descriptor, select the candidate sample image from the multiple frames of sample images; the distance between the global descriptor to be tested and the sample global descriptor corresponding to the candidate sample image is the minimum distance; or the distance between the global descriptor to be tested and the sample global descriptor corresponding to the candidate sample image is less than a distance threshold.

[0161] Exemplarily, when determining the global descriptor to be tested corresponding to the image to be tested, the determination module 72 is specifically used to: determine the bag-of-words vector corresponding to the image to be tested based on a trained dictionary model, and determine the bag-of-words vector as the global descriptor to be tested corresponding to the image to be tested; or, input the image to be tested into a trained deep learning model to obtain a target vector corresponding to the image to be tested, and determine the target vector as the global descriptor to be tested corresponding to the image to be tested.

[0162] Exemplarily, when determining a target map point corresponding to the feature point from multiple map points corresponding to the candidate sample image, the determination module 72 is specifically used to: determine a local descriptor to be tested corresponding to the feature point, the local descriptor to be tested is used to represent the feature vector of the image block where the feature point is located, and the image block is located in the image to be tested; determine the distance between the local descriptor to be tested and the sample local descriptor corresponding to each map point corresponding to the candidate sample image; wherein the three-dimensional visual map includes at least a sample local descriptor corresponding to each map point corresponding to the candidate sample image; based on the distance between the local descriptor to be tested and each sample local descriptor, select a target map point from the multiple map points; the distance between the local descriptor to be tested and the sample local descriptor corresponding to the target map point is the minimum distance, and the minimum distance is less than a distance threshold.

[0163] Exemplarily, when the generation module 73 generates the fused positioning trajectory of the terminal device in the three-dimensional visual map based on the self-positioning trajectory and the global positioning trajectory, it is specifically used to: select N self-positioning postures corresponding to the target time period from all the self-positioning postures included in the self-positioning trajectory, and select P global positioning postures corresponding to the target time period from all the global positioning postures included in the global positioning trajectory; the N is greater than the P; determine the N fused positioning postures corresponding to the N self-positioning postures based on the N self-positioning postures and the P global positioning postures, and the N self-positioning postures correspond one-to-one to the N fused positioning postures; generate the fused positioning trajectory of the terminal device in the three-dimensional visual map based on the N fused positioning postures; wherein, the frame rate of the fused positioning postures included in the fused positioning trajectory is equal to the frame rate of the self-positioning postures included in the self-positioning trajectory.

[0164] Exemplarily, after the generation module 73 generates a fused positioning trajectory of the terminal device in the three-dimensional visual map based on the self-positioning trajectory and the global positioning trajectory, it is also used to: select an initial fused positioning pose from the fused positioning trajectory; select an initial self-positioning pose corresponding to the initial fused positioning pose from the self-positioning trajectory; select a target self-positioning pose from the self-positioning trajectory, and determine the target fused positioning pose based on the initial fused positioning pose, the initial self-positioning pose and the target self-positioning pose; and generate a new fused positioning trajectory based on the target fused positioning pose and the fused positioning trajectory.

[0165] Based on the same application concept as the above method, a terminal device is proposed in an embodiment of the present application, which may include: a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the posture determination method disclosed in the above example of this application.

[0166] Based on the same application concept as the above method, a terminal device is proposed in an embodiment of the present application, including: a visual sensor, used to obtain a target image of the target scene during the movement of the terminal device in the target scene, and input the target image to a processor; a motion sensor, used to obtain motion data of the terminal device during the movement of the terminal device in the target scene, and input the motion data to a processor; a processor, used to determine the self-positioning trajectory of the terminal device based on the target image and the motion data; determine the target map point corresponding to the target image from the three-dimensional visual map of the target scene, and determine the global positioning trajectory of the terminal device in the three-dimensional visual map based on the target map point; generate a fused positioning trajectory of the terminal device in the three-dimensional visual map based on the self-positioning trajectory and the global positioning trajectory, and output the fused positioning trajectory; the frame rate of the fused positioning posture included in the fused positioning trajectory is greater than the frame rate of the global positioning posture included in the global positioning trajectory. Exemplarily, the terminal device is a wearable device, and the visual sensor and the motion sensor are deployed on the wearable device; or, the terminal device is a recorder, and the visual sensor and the motion sensor are deployed on the recorder; or, the terminal device is a camera, and the visual sensor and the motion sensor are deployed on the camera.

[0167] Based on the same application concept as the above method, an embodiment of the present application also provides a machine-readable storage medium, on which a number of computer instructions are stored. When the computer instructions are executed by the processor, the posture determination method disclosed in the above example of the present application can be implemented.

[0168] The machine-readable storage medium may be any electronic, magnetic, optical, or other physical storage device that may contain or store information, such as executable instructions, data, and the like. For example, the machine-readable storage medium may be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, a storage drive (such as a hard disk drive), a solid-state drive, any type of storage disk (such as a CD, DVD, etc.), or similar storage media, or a combination thereof.

[0169] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer, which may be in the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email transceiver, game console, tablet computer, wearable device, or any combination of these devices.

[0170] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0171] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the embodiments of the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0172] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0173] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0174] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A method for determining a posture, characterized in that: Applied to a terminal device, the terminal device includes a three-dimensional visual map of a target scene, and the terminal device, during movement in the target scene, includes: Acquiring a target image of the target scene and motion data of the terminal device; determining a self-positioning trajectory of the terminal device based on the target image and the motion data; Determining a target map point corresponding to the target image from the three-dimensional visual map, and determining a global positioning trajectory of the terminal device in the three-dimensional visual map based on the target map point; Generating a fused positioning trajectory of the terminal device in the three-dimensional visual map based on the self-positioning trajectory and the global positioning trajectory, and outputting the fused positioning trajectory; wherein a frame rate of a fused positioning pose included in the fused positioning trajectory is greater than a frame rate of a global positioning pose included in the global positioning trajectory; wherein the self-positioning pose included in the self-positioning trajectory corresponds one-to-one to the fused positioning pose; Wherein, generating a fused positioning trajectory of the terminal device in the three-dimensional visual map based on the self-positioning trajectory and the global positioning trajectory includes: selecting N self-positioning postures corresponding to the target time period from all the self-positioning postures included in the self-positioning trajectory, and selecting P global positioning postures corresponding to the target time period from all the global positioning postures included in the global positioning trajectory; wherein, N is greater than P; determining N fused positioning postures corresponding to the N self-positioning postures based on the N self-positioning postures and the P global positioning postures, and the N self-positioning postures correspond one-to-one to the N fused positioning postures; generating a fused positioning trajectory of the terminal device in the three-dimensional visual map based on the N fused positioning postures; wherein, the frame rate of the fused positioning postures included in the fused positioning trajectory is equal to the frame rate of the self-positioning postures included in the self-positioning trajectory.

2. The method according to claim 1, characterized in that The determining the self-positioning trajectory of the terminal device based on the target image and the motion data includes: If the target image includes multiple frames, traverse the current frame image from the multiple frames; determine the self-positioning posture corresponding to the current frame image based on the self-positioning postures corresponding to the K frames of images preceding the current frame image, the map position of the terminal device in the self-positioning coordinate system, and the motion data; generate a self-positioning trajectory of the terminal device in the self-positioning coordinate system based on the self-positioning postures corresponding to the multiple frames of images; Among them, if the current frame image is a key image, a map position in the self-positioning coordinate system is generated based on the current position of the terminal device; if the number of matching feature points between the current frame image and the previous frame image of the current frame image does not reach a preset threshold, the current frame image is determined to be a key image.

3. The method according to claim 1, characterized in that Determining a target map point corresponding to the target image from the three-dimensional visual map, and determining a global positioning trajectory of the terminal device in the three-dimensional visual map based on the target map point, includes: If the target image includes multiple frames of images, M frames of images are selected from the multiple frames of images as the images to be tested; For each frame of the image to be tested, based on the similarity between the image to be tested and the multiple frames of sample images corresponding to the three-dimensional visual map, selecting a candidate sample image from the multiple frames of sample images; Acquire multiple feature points from the image to be tested; for each feature point, determine a target map point corresponding to the feature point from multiple map points corresponding to the candidate sample image; Based on the multiple feature points and the target map points corresponding to the multiple feature points, the global positioning pose in the three-dimensional visual map corresponding to the image to be tested is determined; based on the global positioning pose corresponding to the M frames of image to be tested, the global positioning trajectory of the terminal device in the three-dimensional visual map is generated.

4. The method according to claim 3, characterized in that The selecting a candidate sample image from the multiple frames of sample images based on the similarity between the image to be tested and the multiple frames of sample images corresponding to the three-dimensional visual map includes: Determining a global descriptor to be tested corresponding to the image to be tested, and determining a distance between the global descriptor to be tested and a sample global descriptor corresponding to each frame of the sample image corresponding to the three-dimensional visual map; wherein the three-dimensional visual map includes at least a sample global descriptor corresponding to each frame of the sample image; Based on the distance between the global descriptor to be tested and each sample global descriptor, a candidate sample image is selected from the multiple frames of sample images; wherein the distance between the global descriptor to be tested and the sample global descriptor corresponding to the candidate sample image is the minimum distance; or, the distance between the global descriptor to be tested and the sample global descriptor corresponding to the candidate sample image is less than a distance threshold.

5. The method according to claim 4, characterized in that The determining of the global descriptor to be tested corresponding to the image to be tested includes: Determine a bag-of-words vector corresponding to the image to be tested based on a trained dictionary model, and determine the bag-of-words vector as a global descriptor to be tested corresponding to the image to be tested; or The image to be tested is input into a trained deep learning model to obtain a target vector corresponding to the image to be tested, and the target vector is determined as a global descriptor to be tested corresponding to the image to be tested.

6. The method according to claim 3, characterized in that The step of determining a target map point corresponding to the feature point from a plurality of map points corresponding to the candidate sample image comprises: Determine a local descriptor to be tested corresponding to the feature point, where the local descriptor to be tested is used to represent a feature vector of an image block where the feature point is located, and the image block is located in the image to be tested; Determining a distance between the local descriptor to be tested and a sample local descriptor corresponding to each map point corresponding to the candidate sample image; wherein the three-dimensional visual map includes at least a sample local descriptor corresponding to each map point corresponding to the candidate sample image; Based on the distance between the local descriptor to be tested and each sample local descriptor, a target map point is selected from the multiple map points; the distance between the local descriptor to be tested and the sample local descriptors corresponding to the target map point is a minimum distance, and the minimum distance is less than a distance threshold.

7. The method according to claim 1, characterized in that After generating a fused positioning trajectory of the terminal device in the three-dimensional visual map based on the self-positioning trajectory and the global positioning trajectory, the method further includes: Selecting an initial fused positioning pose from the fused positioning trajectory; Selecting an initial self-positioning pose corresponding to the initial fused positioning pose from the self-positioning trajectory; Selecting a target self-positioning pose from the self-positioning trajectory, and determining a target fused positioning pose based on the initial fused positioning pose, the initial self-positioning pose, and the target self-positioning pose; A new fused positioning trajectory is generated based on the target fused positioning pose and the fused positioning trajectory.

8. A posture determination device, characterized in that: Applied to a terminal device, the terminal device includes a three-dimensional visual map of a target scene, and during movement of the terminal device in the target scene, includes: An acquisition module, configured to acquire a target image of the target scene and motion data of the terminal device; a determination module, configured to determine a self-positioning trajectory of the terminal device based on the target image and the motion data; determine a target map point corresponding to the target image from the three-dimensional visual map, and determine a global positioning trajectory of the terminal device in the three-dimensional visual map based on the target map point; a generation module, configured to generate a fused positioning trajectory of the terminal device in a three-dimensional visual map based on the self-positioning trajectory and the global positioning trajectory, and output the fused positioning trajectory; the frame rate of the fused positioning pose included in the fused positioning trajectory is greater than the frame rate of the global positioning pose included in the global positioning trajectory; wherein the self-positioning pose included in the self-positioning trajectory corresponds one-to-one to the fused positioning pose; Wherein, when the generation module generates the fused positioning trajectory of the terminal device in the three-dimensional visual map based on the self-positioning trajectory and the global positioning trajectory, it is specifically used to: select N self-positioning postures corresponding to the target time period from all the self-positioning postures included in the self-positioning trajectory, and select P global positioning postures corresponding to the target time period from all the global positioning postures included in the global positioning trajectory; the N is greater than the P; determine the N fused positioning postures corresponding to the N self-positioning postures based on the N self-positioning postures and the P global positioning postures, and the N self-positioning postures correspond one-to-one to the N fused positioning postures; generate the fused positioning trajectory of the terminal device in the three-dimensional visual map based on the N fused positioning postures; wherein, the frame rate of the fused positioning postures included in the fused positioning trajectory is equal to the frame rate of the self-positioning postures included in the self-positioning trajectory.

9. The device according to claim 8, It is characterized by: in, When determining the self-positioning trajectory of the terminal device based on the target image and the motion data, the determination module is specifically used to: if the target image includes multiple frame images, traverse the current frame image from the multiple frame images; determine the self-positioning posture corresponding to the current frame image based on the self-positioning postures corresponding to the K frame images preceding the current frame image, the map position of the terminal device in the self-positioning coordinate system and the motion data; generate the self-positioning trajectory of the terminal device in the self-positioning coordinate system based on the self-positioning postures corresponding to the multiple frame images; wherein, if the current frame image is a key image, generate the map position in the self-positioning coordinate system based on the current position of the terminal device; if the number of matching feature points between the current frame image and the frame image before the current frame image does not reach a preset threshold, determine that the current frame image is a key image; Among them, the determination module determines the target map point corresponding to the target image from the three-dimensional visual map, and determines the global positioning trajectory of the terminal device in the three-dimensional visual map based on the target map point, specifically for: if the target image includes multiple frames of images, select M frames of images from the multiple frames of images as the images to be tested; for each frame of the image to be tested, based on the similarity between the image to be tested and the multiple frames of sample images corresponding to the three-dimensional visual map, select a candidate sample image from the multiple frames of sample images; obtain multiple feature points from the image to be tested; for each feature point, determine the target map point corresponding to the feature point from the multiple map points corresponding to the candidate sample images; determine the global positioning pose in the three-dimensional visual map corresponding to the image to be tested based on the multiple feature points and the target map points corresponding to the multiple feature points; generate the global positioning trajectory of the terminal device in the three-dimensional visual map based on the global positioning poses corresponding to the M frames of images to be tested; Wherein, the determination module is specifically used to select candidate sample images from multiple frames of sample images based on the similarity between the image to be tested and multiple frames of sample images corresponding to the three-dimensional visual map: determine the global descriptor to be tested corresponding to the image to be tested, and determine the distance between the global descriptor to be tested and the sample global descriptor corresponding to each frame of sample image corresponding to the three-dimensional visual map; wherein the three-dimensional visual map includes a sample global descriptor corresponding to each frame of sample image; based on the distance between the global descriptor to be tested and each sample global descriptor, select the candidate sample image from the multiple frames of sample images; the distance between the global descriptor to be tested and the sample global descriptor corresponding to the candidate sample image is the minimum distance; or the distance between the global descriptor to be tested and the sample global descriptor corresponding to the candidate sample image is less than a distance threshold; Wherein, when determining the global descriptor to be tested corresponding to the image to be tested, the determination module is specifically used to: determine the bag-of-words vector corresponding to the image to be tested based on a trained dictionary model, and determine the bag-of-words vector as the global descriptor to be tested corresponding to the image to be tested; or input the image to be tested into a trained deep learning model to obtain a target vector corresponding to the image to be tested, and determine the target vector as the global descriptor to be tested corresponding to the image to be tested; Wherein, when the determination module determines the target map point corresponding to the feature point from the multiple map points corresponding to the candidate sample image, it is specifically used to: determine the local descriptor to be tested corresponding to the feature point, the local descriptor to be tested is used to represent the feature vector of the image block where the feature point is located, and the image block is located in the image to be tested; determine the distance between the local descriptor to be tested and the sample local descriptor corresponding to each map point corresponding to the candidate sample image; wherein the three-dimensional visual map includes at least the sample local descriptor corresponding to each map point corresponding to the candidate sample image; based on the distance between the local descriptor to be tested and each sample local descriptor, select the target map point from the multiple map points; the distance between the local descriptor to be tested and the sample local descriptor corresponding to the target map point is the minimum distance, and the minimum distance is less than the distance threshold; Among them, after the generation module generates a fused positioning trajectory of the terminal device in the three-dimensional visual map based on the self-positioning trajectory and the global positioning trajectory, it is also used to: select an initial fused positioning pose from the fused positioning trajectory; select an initial self-positioning pose corresponding to the initial fused positioning pose from the self-positioning trajectory; select a target self-positioning pose from the self-positioning trajectory, and determine the target fused positioning pose based on the initial fused positioning pose, the initial self-positioning pose and the target self-positioning pose; generate a new fused positioning trajectory based on the target fused positioning pose and the fused positioning trajectory.

10. A terminal device, characterized in that: include: a processor and a machine-readable storage medium storing machine-executable instructions capable of being executed by the processor; The processor is configured to execute machine-executable instructions to implement the method steps described in any one of claims 1-7.

11. A terminal device, characterized in that: include: a visual sensor, configured to acquire a target image of the target scene during movement of the terminal device in the target scene, and input the target image to a processor; a motion sensor, configured to obtain motion data of the terminal device during movement of the terminal device in the target scene, and input the motion data to the processor; a processor, configured to determine a self-positioning trajectory of the terminal device based on the target image and the motion data; Determining a target map point corresponding to the target image from the three-dimensional visual map of the target scene, and determining a global positioning trajectory of the terminal device in the three-dimensional visual map based on the target map point; generating a fused positioning trajectory of the terminal device in the three-dimensional visual map based on the self-positioning trajectory and the global positioning trajectory, and outputting the fused positioning trajectory; wherein the frame rate of the fused positioning pose included in the fused positioning trajectory is greater than the frame rate of the global positioning pose included in the global positioning trajectory; wherein the self-positioning pose included in the self-positioning trajectory corresponds one-to-one to the fused positioning pose; Wherein, generating a fused positioning trajectory of the terminal device in the three-dimensional visual map based on the self-positioning trajectory and the global positioning trajectory includes: selecting N self-positioning postures corresponding to the target time period from all the self-positioning postures included in the self-positioning trajectory, and selecting P global positioning postures corresponding to the target time period from all the global positioning postures included in the global positioning trajectory; wherein, N is greater than P; determining N fused positioning postures corresponding to the N self-positioning postures based on the N self-positioning postures and the P global positioning postures, and the N self-positioning postures correspond one-to-one to the N fused positioning postures; generating a fused positioning trajectory of the terminal device in the three-dimensional visual map based on the N fused positioning postures; wherein, the frame rate of the fused positioning postures included in the fused positioning trajectory is equal to the frame rate of the self-positioning postures included in the self-positioning trajectory.

12. The terminal device according to claim 11, characterized in that The terminal device is a wearable device, and the visual sensor and the motion sensor are deployed on the wearable device; or, the terminal device is a recorder, and the visual sensor and the motion sensor are deployed on the recorder; or, the terminal device is a camera, and the visual sensor and the motion sensor are deployed on the camera.

Citation Information

Patent Citations

  • Front-end and back-end architecture map positioning method based on computer vision technology

    CN110514198A

  • Multi-sensor fusion coal mine underground space positioning and mapping method and device

    CN113140040A