A three-dimensional visual navigation method and system based on multi-sensor information fusion

By using a multi-sensor information fusion method, the problems of low fusion accuracy and easy tracking loss between virtual targets and real scenes were solved, realizing real-time and accurate virtual target navigation and improving the accuracy and robustness of interactive perception.

CN120544103BActive Publication Date: 2025-12-23ANHUI RUIXIANG DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510686499.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-12-23
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

In existing technologies, registration based on image feature points is easily affected by changes in lighting, occlusion, and interference from similar textures, resulting in low accuracy of virtual target fusion with the real scene and easy loss of tracking, making it impossible to achieve stable and accurate virtual target navigation.

Method used

A multi-sensor information fusion method is adopted to generate virtual target image information by acquiring video frame information and sensor measurement information from multiple scenes. Feature extraction and parsing are performed through image information feature processing function, and registration and fusion are performed by combining multi-sensor measurement information to achieve a refined description and spatial alignment between the virtual target and the real scene, enhance physical consistency, and perform dynamic navigation through visualization tracking calculation.

Benefits of technology

It significantly improves the accuracy and robustness of virtual target and environment interaction perception, realizes real-time, accurate and spatially consistent virtual target visualization information generation, and solves the problem of inaccurate virtual target navigation in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544103B_ABST
    Figure CN120544103B_ABST
Patent Text Reader

Abstract

The application provides a three-dimensional visual navigation method and system based on multi-sensor information fusion, which is suitable for the field of image processing technology, and the method comprises the following steps: obtaining virtual target image feature information and scene video frame feature information according to virtual target image information and scene video frame information, and obtaining virtual target image registration information and scene target image feature registration information through registration processing; generating virtual target image transformation information according to multi-sensor measurement information, virtual target image registration information and scene target image feature registration information; obtaining scene target tracking information according to multi-sensor measurement information, virtual target image transformation information and scene target image feature registration information, so that the virtual target image transformation information is navigated based on the scene target tracking information, and virtual target visual information is generated. The application can improve the real-time positioning accuracy of the virtual target in a dynamic scene, so as to realize accurate navigation of the virtual target visual information.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of image processing, and particularly relates to a three-dimensional visual navigation method and system based on multi-sensor information fusion. BACKGROUND

[0002] In today's digital era, through the application of technology that skillfully fuses virtual information with the real world, the application scenarios are expanding towards diversification with the improvement of hardware device performance, and the application scenarios are gradually expanding from consumer-level entertainment to professional fields such as industrial maintenance, medical navigation, and intelligent security.

[0003] In the prior art, the optical flow method is usually used to track the dynamically changing target in the scene to determine the approximate position of the virtual target in the real scene, and then registration processing is performed, and after the registration is completed, the virtual target is superimposed into the scene picture.

[0004] However, in the prior art, registration based on image feature points alone is easily disturbed by factors such as changes in scene lighting, occlusion, and similar textures, resulting in low registration accuracy, poor fusion effect of virtual targets and real scenes, and misalignment, drift, and other phenomena. In the tracking process, only image information is used for tracking, and when the target temporarily leaves the field of view or the scene changes greatly, the tracking is easily lost, making it difficult to achieve stable and continuous visual tracking. The navigation of the virtual target based on the scene target tracking information is not accurate enough, and the virtual target visual information cannot provide smooth and reliable navigation for the user. SUMMARY

[0005] Therefore, the embodiments of the present application provide a three-dimensional visual navigation method and system based on multi-sensor information fusion, aiming to solve the problems of low registration accuracy, easy disturbance, easy loss of tracking, and inaccurate navigation in the prior art when fusing virtual and real information.

[0006] The first aspect of the embodiments of the present application provides a three-dimensional visual navigation method based on multi-sensor information fusion, comprising:

[0007] Obtaining a plurality of scene video frame information and a plurality of sensor measurement information;

[0008] Generating at least one virtual target image information;

[0009] Based on a plurality of preset image information feature processing functions, performing feature extraction and analysis processing on the virtual target image information and the scene video frame information to obtain virtual target image feature information and scene video frame feature information;

[0010] According to the virtual target image feature information and the scene video frame feature information, registration processing is performed to obtain virtual target image registration information and scene target image feature registration information;

[0011] According to the multiple sensor measurement information, image fusion processing is performed on the virtual target image registration information and the scene target image feature registration information to generate virtual target image transformation information;

[0012] According to the multiple sensor measurement information and the virtual target image transformation information, visual tracking calculation is performed on the scene target image feature registration information to obtain scene target tracking information, so that the virtual target image transformation information is navigated based on the scene target tracking information to generate virtual target visual information.

[0013] A second aspect of the embodiment of the application provides a three-dimensional visual navigation system based on multi-sensor information fusion, comprising:

[0014] An information acquisition module is configured to acquire multiple scene video frame information and multiple sensor measurement information;

[0015] A virtual target image information generation module is configured to generate at least one virtual target image information;

[0016] An image feature information generation module is configured to perform feature extraction and analysis processing on the virtual target image information and the scene video frame information based on multiple preset image information feature processing functions to obtain virtual target image feature information and scene video frame feature information;

[0017] An image registration information generation module is configured to perform registration processing on the virtual target image feature information and the scene video frame feature information to obtain virtual target image registration information and scene target image feature registration information;

[0018] A virtual target image transformation information generation module is configured to perform image fusion processing on the virtual target image registration information and the scene target image feature registration information according to the multiple sensor measurement information to generate virtual target image transformation information;

[0019] A virtual target visual information generation module is configured to perform visual tracking calculation on the scene target image feature registration information according to the multiple sensor measurement information and the virtual target image transformation information to obtain scene target tracking information, so that the virtual target image transformation information is navigated based on the scene target tracking information to generate virtual target visual information.

[0020] The third aspect of the embodiment of the present application provides a terminal device, the terminal device comprises a memory and a processor, the memory stores a computer program capable of running on the processor, and the processor implements the steps of the three-dimensional visualization navigation method based on multi-sensor information fusion in the first aspect.

[0021] Compared with the prior art, the embodiment of the present application has the beneficial effects that: the present application enriches the environment perception data by acquiring multiple sensor measurement information, generates virtual target image information, and extracts and analyzes the features of the virtual target image information and the scene video frame information through an image information feature processing function, to realize fine description of the virtual target and the real scene picture. The spatial alignment of the virtual target image and the scene target image is realized based on the feature-based registration processing, the registered virtual and real images are fused in combination with the multi-sensor measurement information, to enhance the physical consistency of the virtual target transformation information and the real scene, and then the scene target feature registration information is visualized and tracked by the multi-sensor data and the virtual target image transformation information, to realize dynamic navigation of the virtual target based on the scene target tracking information, thereby generating real-time, accurate and spatially consistent virtual target visualization information, and significantly improving the interactive perception accuracy and robustness of the identified target and the environment. BRIEF DESCRIPTION OF DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0023] Figure 1 is an implementation flowchart of the three-dimensional visualization navigation method based on multi-sensor information fusion provided by the first embodiment of the present application;

[0024] Figure 2 is an implementation flowchart of the three-dimensional visualization navigation method based on multi-sensor information fusion provided by the second embodiment of the present application;

[0025] Figure 3 is an implementation flowchart of the three-dimensional visualization navigation method based on multi-sensor information fusion provided by the third embodiment of the present application;

[0026] Figure 4 is an implementation flowchart of the three-dimensional visualization navigation method based on multi-sensor information fusion provided by the fourth embodiment of the present application;

[0027] Figure 5is an implementation flowchart of the three-dimensional visual navigation method based on multi-sensor information fusion provided in Embodiment Five of the present application;

[0028] Figure 6 is an implementation flowchart of the three-dimensional visual navigation method based on multi-sensor information fusion provided in Embodiment Six of the present application;

[0029] Figure 7 is an implementation flowchart of the three-dimensional visual navigation method based on multi-sensor information fusion provided in Embodiment Seven of the present application;

[0030] Figure 8 is an implementation flowchart of the three-dimensional visual navigation method based on multi-sensor information fusion provided in Embodiment Eight of the present application;

[0031] Figure 9 is a structural schematic diagram of the three-dimensional visual navigation system based on multi-sensor information fusion provided in the present application;

[0032] Figure 10 is a schematic diagram of the terminal device provided in the present application. DETAILED DESCRIPTION

[0033] In the following description, specific details are set forth in order to provide a thorough understanding of the embodiments of the present application. However, persons skilled in the art will understand that the present application can be implemented in other embodiments without these specific details. In other cases, well-known systems, devices, circuits, and methods have not been described in detail in order not to obscure the description of the present application with unnecessary detail.

[0034] In order to illustrate the technical solutions described in the present application, the following will be described through specific embodiments.

[0035] Figure 1 An implementation flowchart of the three-dimensional visual navigation method based on multi-sensor information fusion provided in Embodiment One of the present application is shown, which is described in detail as follows:

[0036] In step S101, a plurality of scene video frame information and a plurality of sensor measurement information are acquired.

[0037] In the embodiment, the scene video frame information can be obtained by photographing the scene through a photographing device, can be obtained by photographing through an RGB camera, can be obtained by photographing through a depth camera, and can be obtained by collecting continuous video frames output by the photographing device at a fixed frame rate. The plurality of sensors can include an inertial measurement unit (IMU), a GPS, and a depth sensor (such as a Kinect sensor). The plurality of sensor measurement information can include acceleration information, angular velocity information, and the like collected by the inertial measurement unit at a high frequency, can include latitude and longitude information of the photographed location obtained by the GPS, can be a distance between an object and the device measured by the depth sensor, and further generate depth map information based on the distance. The video frame and the sensor data can be aligned by a timestamp, for example, using a hardware synchronization trigger or a software interpolation algorithm to compensate for the sampling frequency difference.

[0038] In step S102, at least one virtual target image information is generated.

[0039] In the embodiment, a three-dimensional model of the virtual object can be designed manually using modeling software (such as Blender or Maya), and output in.obj,.stl, or the like. The three-dimensional model can include vertex coordinates, texture maps, and triangular mesh surface information, and a two-dimensional image of the virtual target can be generated by loading the three-dimensional model through a rendering engine and manually setting parameters such as lighting and material.

[0040] In step S103, a plurality of preset image information feature processing functions are used to perform feature extraction and analysis processing on the virtual target image information and the scene video frame information, to obtain virtual target image feature information and scene video frame feature information.

[0041] In the embodiment, the plurality of preset image information feature processing functions can be designed manually, and can be designed based on a Gaussian function to perform feature extraction and analysis processing on the virtual target image information and the scene video frame information. The feature extraction and analysis processing can first convert the virtual target image information and the scene video frame information into a vector form, then use the vector as an independent variable of the plurality of preset image information feature processing functions, and then output a plurality of virtual target image feature information and a plurality of scene video frame feature information through calculation of the plurality of preset image information feature processing functions.

[0042] In step S104, registration processing is performed according to the virtual target image feature information and the scene video frame feature information, to obtain virtual target image registration information and scene target image feature registration information.

[0043] In the embodiment, the cosine similarity of the virtual target image feature information and the scene video frame feature information can be calculated, and the feature matching can be completed when the cosine similarity is greater than 0.8, that is, the registration of the virtual target image feature information and the scene video frame feature information is successful, and then the registered virtual target image feature information and the scene video frame feature information are taken as the virtual target image registration information and the scene target image feature registration information.

[0044] In step S105, the virtual target image registration information and the scene target image feature registration information are image fused according to the sensor measurement information, and virtual target image transformation information is generated.

[0045] In the embodiment, the acceleration and angular velocity measured by the IMU can be integrated to predict the position and motion attitude of the device, so as to compensate for the motion error in the registration process. The global three-dimensional map of the scene can be constructed by fusing the GPS data and the depth value measured by the depth sensor, and the absolute spatial reference for the virtual target can be provided. The virtual target can be rendered according to the predicted position and motion attitude of the device, and the rendered image information can be taken as the virtual target image transformation information.

[0046] In step S106, the scene target tracking information is obtained by visualizing tracking calculation of the scene target image feature registration information according to the sensor measurement information and the virtual target image transformation information, so that the virtual target image transformation information is navigated based on the scene target tracking information, and the virtual target visualization information is generated.

[0047] In the embodiment, the feature matching algorithm can be used to track the feature points in the scene target image feature registration information to obtain the coordinate information of the feature points in the new frame. The data measured by the IMU can be taken as the state input, and the two-dimensional motion and three-dimensional space information of the feature points can be fused by the particle filtering algorithm to calculate the real motion parameters of the target. Then, the two-dimensional pixel offset of the target is calculated based on the displacement of the feature points, and the position and motion attitude change data calculated by the sensor data prediction are combined to obtain the new position of the target in the image, so as to realize the visual tracking processing. The position of the target obtained by tracking is fed back to the pose control module of the virtual target to update the rendering parameters, so that the virtual target image transformation information can follow the scene target in real time, and finally the dynamic alignment virtual target visualization information is generated.

[0048] The three-dimensional visual navigation method based on multi-sensor information fusion provided in the embodiments of the present application enriches the environmental perception data by acquiring multi-sensor measurement information, generates virtual target image information, and performs feature extraction and analysis on the virtual target image information and scene video frame information through an image information feature processing function, so as to realize fine description of the virtual target and the real scene picture. The spatial alignment of the virtual target image and the scene target image is realized based on feature-based registration processing, the registered virtual-real image is fused in combination with the multi-sensor measurement information, so as to enhance the physical consistency of the virtual target transformation information and the real scene. Then, the scene target feature registration information is visualized and tracked through the multi-sensor data and the virtual target image transformation information, the dynamic navigation of the virtual target based on the scene target tracking information is realized, real-time, accurate and spatially consistent virtual target visual information is generated, and the interactive perception accuracy and robustness for identifying the target and the environmental condition are significantly improved.

[0049] Figure 2 The implementation flowchart of the three-dimensional visual navigation method based on multi-sensor information fusion provided in the second embodiment of the present application is shown, which is different from the first embodiment described above in that:

[0050] The plurality of preset image information feature processing functions include a preset first scale image information feature processing function, a preset second scale image information feature processing function, and a plurality of preset image information feature vector processing functions.

[0051] The step S103 specifically includes:

[0052] In step S201, the virtual target image information and the scene video frame information are format-converted to obtain a virtual target pixel matrix and a scene video frame pixel matrix.

[0053] In this embodiment, the virtual target image information and the scene video frame information can be converted into a standard pixel matrix format, or the virtual target image information and the scene video frame information can be read through OpenCV or PIL library and converted into NumPy arrays, so that the NumPy arrays are used as the virtual target pixel matrix and the scene video frame pixel matrix.

[0054] In step S202, the virtual target pixel matrix and the scene video frame pixel matrix are analyzed and calculated according to the preset first scale image information feature processing function and the preset second scale image information feature processing function to obtain first scale virtual target pixel matrix variable information, second scale virtual target pixel matrix variable information, first scale scene video frame pixel matrix variable information, and second scale scene video frame pixel matrix variable information.

[0055] In the embodiment, the preset first scale image information feature processing function and the preset second scale image information feature processing function can be artificially set, can be set based on a Gaussian kernel function, can be designed based on a radial basis function, and can be designed based on a structure of a weight part and a bias part. The first scale can be greater than the second scale, or the second scale can be greater than the first scale. It can be understood that the first scale and the second scale are only used to illustrate that image information feature processing functions with different processing scales are used to analyze and calculate the virtual target pixel matrix and the scene video frame pixel matrix. In actual application, there can be a third scale image information feature processing function, a fourth scale image information feature processing function, and the like. The virtual target pixel matrix and the scene video frame pixel matrix can be respectively taken as the independent variables of the preset first scale image information feature processing function and the preset second scale image information feature processing function. The calculation results are respectively taken as the first scale virtual target pixel matrix variable information, the second scale virtual target pixel matrix variable information, the first scale scene video frame pixel matrix variable information, and the second scale scene video frame pixel matrix variable information.

[0056] In step S203, the first scale virtual target pixel matrix variable information and the second scale virtual target pixel matrix variable information are subjected to difference calculation to obtain scale difference virtual target pixel matrix variable information.

[0057] In the embodiment, when the first scale virtual target pixel matrix variable information is greater than the second scale virtual target pixel matrix variable information, the first scale virtual target pixel matrix variable information is subtracted from the second scale virtual target pixel matrix variable information to perform difference calculation, so as to obtain the scale difference virtual target pixel matrix variable information. When the second scale virtual target pixel matrix variable information is greater than the first scale virtual target pixel matrix variable information, the second scale virtual target pixel matrix variable information is subtracted from the first scale virtual target pixel matrix variable information to perform difference calculation, so as to obtain the scale difference virtual target pixel matrix variable information.

[0058] In step S204, the first scale scene video frame pixel matrix variable information and the second scale scene video frame pixel matrix variable information are subjected to difference calculation to obtain scale difference scene video frame pixel matrix variable information.

[0059] In the embodiment, when the first scale scene video frame pixel matrix variable information is greater than the second scale scene video frame pixel matrix variable information, the difference calculation is performed by subtracting the second scale scene video frame pixel matrix variable information from the first scale scene video frame pixel matrix variable information to obtain the scale difference scene video frame pixel matrix variable information; when the second scale scene video frame pixel matrix variable information is greater than the first scale scene video frame pixel matrix variable information, the difference calculation is performed by subtracting the first scale scene video frame pixel matrix variable information from the second scale scene video frame pixel matrix variable information to obtain the scale difference scene video frame pixel matrix variable information.

[0060] In step S205, the extreme value of the scale difference virtual target pixel matrix variable information and the extreme value of the scale difference scene video frame pixel matrix variable information are extracted to obtain the virtual target image feature center information and the scene video frame feature center information.

[0061] In the embodiment, it can be understood that the scale difference virtual target pixel matrix variable information and the scale difference scene video frame pixel matrix variable information are both in the form of matrix, and the maximum value element can be determined from the matrix and taken as the virtual target image feature center information and the scene video frame feature center information.

[0062] In step S206, the virtual target image feature information and the scene video frame feature information are obtained according to the virtual target image feature center information, the scene video frame feature center information, the preset image information feature extraction radius information and the plurality of preset image information feature vector processing functions.

[0063] In the embodiment, the preset image information feature extraction radius information and the plurality of preset image information feature vector processing functions can be artificially set, wherein the preset image information feature extraction radius information can take the value 64, the plurality of preset image information feature vector processing functions can be designed based on the structure of the weight coefficient and the bias term, and the calculation process of the preset image information feature vector processing function can be that the independent variable is multiplied by the weight coefficient, then the multiplication result is added to the bias term, and the addition result is taken as the output of the image information feature vector processing function. The virtual target image feature center information and the scene video frame feature center information can be taken as the centers of two parts of a circle respectively, the preset image information feature extraction radius information can be taken as a feature extraction range, the feature points falling into the feature extraction range are extracted, and the feature points are format-converted to generate feature vectors, so that the feature vectors are calculated layer by layer through the plurality of preset image information feature vector processing functions, and the calculation result output by the last image information feature vector processing function is the virtual target image feature information and the scene video frame feature information.

[0064] The three-dimensional visual navigation method based on multi-sensor information fusion provided in the embodiments of the present application uses preset first scale and second scale image information feature processing functions to perform multi-scale analysis on the pixel matrix, respectively captures fine details and coarse-grained structure features, so that the feature capture of the virtual target image information and the scene video frame information is more comprehensive and detailed, highlights the key features of a specific size by performing difference calculation on pixel matrix variable information of different scales, enhances the adaptability of the features to scale changes, determines the feature center by extracting the extreme value of the scale difference matrix variable information, realizes accurate positioning of the feature points, and then generates feature information based on the virtual target image feature center information, the scene video frame feature center information, the preset image information feature extraction radius information, and a plurality of preset image information feature vector processing functions, ensures that the feature vector contains not only the spatial structure information of the local neighborhood but also the global expression capability, so as to realize stable extraction and description of the virtual target and the scene video frame features in a complex scene, improve the precision and robustness of feature matching, and provide a reliable feature basis for subsequent registration, fusion and tracking tasks.

[0065] Figure 3 An implementation flowchart of the three-dimensional visual navigation method based on multi-sensor information fusion provided in the third embodiment of the present application is shown, which is different from the second embodiment described above in that the step S206 specifically includes:

[0066] In step S301, based on the virtual target image feature center information and the scene video frame feature center information, a plurality of virtual target image feature extraction point information and a plurality of scene video frame feature extraction point information are obtained according to the preset image information feature extraction radius information.

[0067] In this embodiment, the preset image information feature extraction radius information can be artificially set and can be an interval distance with a value of 64 elements. The virtual target image feature center information and the scene video frame feature center information are taken as the center of a circle, and the preset image information feature extraction radius information is taken as the radius of the circle to generate a circle center region. All pixel point coordinates in the region are extracted by coordinate traversal as the plurality of virtual target image feature extraction point information and the plurality of scene video frame feature extraction point information.

[0068] In step S302, virtual target image feature extraction vector information and scene video frame feature extraction vector information are generated according to the plurality of virtual target image feature extraction point information and the plurality of scene video frame feature extraction point information.

[0069] In the embodiment, pixel values of each virtual target image feature extraction point are unfolded into a one-dimensional vector in channel order (such as RGB), and spliced into virtual target image feature extraction vector information; similarly, pixel values of scene video frame feature extraction points are converted into scene video frame feature extraction vector information, and the vector length is the number of extraction points x the number of channels.

[0070] In step S303, virtual target image feature information and scene video frame feature information are obtained according to the virtual target image feature extraction vector information, the scene video frame feature extraction vector information, and a plurality of preset image information feature vector processing functions.

[0071] In the embodiment, the virtual target image feature extraction vector information and the scene video frame feature extraction vector information are respectively input into a plurality of preset image information feature vector processing functions. The output result after layer-by-layer calculation by the plurality of preset image information feature vector processing functions can be weighted and summed, and the weighted sum result can be normalized by a ReLU function, so as to finally output virtual target image feature information and scene video frame feature information of a fixed dimension.

[0072] The three-dimensional visual navigation method based on multi-sensor information fusion provided by the embodiment of the application uses the feature center positioning and the radius-constrained feature extraction mechanism, takes the virtual target image feature center information and the scene video frame feature center information as spatial anchor points, combines the preset image information feature extraction radius information, ensures that the feature extraction points accurately cover the target local structure, so that the generated virtual target image feature extraction vector information and scene video frame feature extraction vector information contain not only the core response value of the feature points, but also the neighborhood spatial context relationship such as texture distribution and gradient direction, avoiding the one-sidedness of single pixel features. Through layer-by-layer calculation of a plurality of preset image information feature vector processing functions, the feature information is gradually abstracted from the original pixel value to a high-level semantic representation with discriminability, thereby improving the robustness of the feature information to scale changes and spatial position changes, and reducing local noise interference through the global expression ability of the feature vector, so as to provide a feature basis with detail accuracy and global consistency for subsequent registration, and comprehensively improve the registration accuracy and stability.

[0073] Figure 4 An implementation flowchart of the three-dimensional visual navigation method based on multi-sensor information fusion provided by the fourth embodiment of the application is shown, which is different from the third embodiment described above in that:

[0074] The plurality of preset image information feature vector processing functions include a preset first image information feature vector processing function, a preset second image information feature vector processing function, and a preset third image information feature vector processing function.

[0075] The step S303 specifically comprises:

[0076] The step S401 obtains first virtual target image feature vector transformation information and first scene video frame feature vector transformation information according to the virtual target image feature extraction vector information, the scene video frame feature extraction vector information and a preset first image information feature vector processing function.

[0077] In the embodiment, the preset first image information feature vector processing function can be artificially set, can be a linear transformation function, or can be a matrix multiplication of the virtual target image feature extraction vector information and the scene video frame feature extraction vector information by the first image information feature vector processing function, and the result of the matrix multiplication is taken as the first virtual target image feature vector transformation information and the first scene video frame feature vector transformation information.

[0078] The step S402 obtains second virtual target image feature vector transformation information and second scene video frame feature vector transformation information according to the first virtual target image feature vector transformation information, the first scene video frame feature vector transformation information and a preset second image information feature vector processing function.

[0079] In the embodiment, the preset second image information feature vector processing function can be artificially set, can be a nonlinear function, can be a hyperbolic tangent function or a Sigmoid function, and the difference from the preset first image information feature vector processing function can be different coefficients. The second image information feature vector processing function can be a matrix multiplication of the first virtual target image feature vector transformation information and the first scene video frame feature vector transformation information, and the result of the matrix multiplication is taken as the second virtual target image feature vector transformation information and the second scene video frame feature vector transformation information.

[0080] The step S403 obtains third virtual target image feature vector transformation information and third scene video frame feature vector transformation information according to the second virtual target image feature vector transformation information, the second scene video frame feature vector transformation information and a preset third image information feature vector processing function.

[0081] In the embodiment, the preset third image information feature vector processing function can be artificially set, and can be a mean aggregation function, that is, used for automatically batching all independent variables, automatically calculating the mean value by batch, and outputting a plurality of mean value calculation results. It can be that the second virtual target image feature vector transformation information and the second scene video frame feature vector transformation information are taken as independent variables of the preset third image information feature vector processing function, and the second virtual target image feature vector transformation information and the second scene video frame feature vector transformation information are processed by dimension reduction through calculation of the third image information feature vector processing function, and the third virtual target image feature vector transformation information and the third scene video frame feature vector transformation information of fixed dimensions are output.

[0082] In step S404, the average values of the third virtual target image feature vector transformation information and the third scene video frame feature vector transformation information are calculated to obtain virtual target image feature vector representation information and scene video frame feature vector representation information.

[0083] In the embodiment, the average values of the elements of the third virtual target image feature vector transformation information and the third scene video frame feature vector transformation information can be calculated, which is used to enhance the feature stability, and the obtained element-level average values are taken as the virtual target image feature vector representation information and the scene video frame feature vector representation information.

[0084] In step S405, the virtual target image feature information and the scene video frame feature information are obtained according to the virtual target image feature vector representation information and the scene video frame feature vector representation information.

[0085] In the embodiment, the virtual target image feature vector representation information and the scene video frame feature vector representation information can be format-converted to the virtual target image feature information and the scene video frame feature information in numerical form, which is in vector form, and used for subsequent registration processing.

[0086] The three-dimensional visual navigation method based on multi-sensor information fusion provided in the embodiments of the present application extracts the feature extraction vector information of the virtual target image feature and the feature extraction vector information of the scene video frame through a first image information feature vector processing function, compresses the redundant dimensions while enhancing the relevance between the features, introduces a nonlinear transformation through a second image information feature vector processing function to suppress the invalid feature information in the preliminary extracted features and amplify the key response feature information, so that the feature vector has nonlinear expression capability, the feature information after the invalid information suppression processing is compressed in spatial dimension through a third image information feature vector processing function, the local features are aggregated into a vector with global representation to reduce the computational complexity, and then the feature vector average value is calculated to effectively smooth the single sample noise, generate stable virtual target image feature vector representation information and scene video frame feature vector representation information, so as to improve the feature discrimination while ensuring the dimension consistency of the feature vector, and lay a foundation for the efficiency and accuracy of subsequent registration.

[0087] Figure 5 The implementation flowchart of the three-dimensional visual navigation method based on multi-sensor information fusion provided in the fifth embodiment of the present application is shown, which is different from the fourth embodiment described above in that the step S405 specifically includes:

[0088] In step S501, the virtual target image feature vector representation information and the scene video frame feature vector representation information are scaled according to the preset feature vector representation information scaling coefficient, to obtain virtual target image feature vector dimension transformation information and scene video frame feature vector dimension transformation information.

[0089] In this embodiment, the preset feature vector representation information scaling coefficient can be artificially set, which can be 0.5, or the virtual target image feature vector representation information and the scene video frame feature vector representation information are scaled in dimension by the feature vector representation information scaling coefficient, and the vector after the dimension scaling is taken as the virtual target image feature vector dimension transformation information and the scene video frame feature vector dimension transformation information, to reduce the computational complexity.

[0090] In step S502, the virtual target feature representation variable and the scene feature representation variable are obtained according to the virtual target image feature vector dimension transformation information, the scene video frame feature vector dimension transformation information, the preset feature vector coefficient information and the preset feature vector compensation information.

[0091] In the embodiment, the preset feature vector coefficient information and the preset feature vector compensation information can be artificially set. The virtual target image feature vector dimension transformation information and the scene video frame feature vector dimension transformation information can be weighted and summed through the preset feature vector coefficient information, and then the weighted and summed result is added to the preset feature vector compensation information, so as to obtain the virtual target feature representation variable and the scene feature representation variable.

[0092] In step S503, the virtual target feature representation variable and the scene feature representation variable are subjected to dimension reduction processing according to the preset feature vector representation information scaling coefficient, so as to obtain the virtual target image feature vector dimension reduction information and the scene video frame feature vector dimension reduction information.

[0093] In the embodiment, the virtual target feature representation variable and the scene feature representation variable can be subjected to dimension reduction processing according to the inverse of the preset feature vector representation information scaling coefficient, so as to restore the virtual target feature representation variable and the scene feature representation variable to the original feature dimension. After restoration, the virtual target image feature vector dimension reduction information and the scene video frame feature vector dimension reduction information are obtained.

[0094] In step S504, the virtual target image feature information and the scene video frame feature information are obtained according to the virtual target image feature vector dimension reduction information, the scene video frame feature vector dimension reduction information, the preset dimension reduction feature vector coefficient information and the preset dimension reduction feature vector compensation information.

[0095] In the embodiment, the preset dimension reduction feature vector coefficient information and the preset dimension reduction feature vector compensation information can be artificially set. The virtual target image feature vector dimension reduction information and the scene video frame feature vector dimension reduction information can be weighted and summed according to the dimension reduction feature vector coefficient information, and then the weighted and summed result is added to the dimension reduction feature vector compensation information, and then the added result is taken as the virtual target image feature information and the scene video frame feature information.

[0096] The three-dimensional visual navigation method based on multi-sensor information fusion provided in the embodiments of the present application realizes flexible dimension reduction of virtual target image feature vector representation information and scene video frame feature vector representation information by using a preset feature vector representation information scaling coefficient, reduces the amount of calculation while retaining key feature semantics, calibrates feature distribution in a low-dimensional space by introducing a preset feature vector coefficient information and compensation information, enhances the physical meaning of the features and the scene adaptability, restores the original expression space of the feature vector by dimension reduction, and further optimizes the feature distribution by using a preset dimension reduction feature vector coefficient information and compensation information, thereby solving the calculation redundancy or semantic loss problem caused by fixed dimension in the prior art, making the generated virtual target image feature information and scene video frame feature information not only adapt to the computing power of different hardware platforms, but also maintain semantic consistency across scenes, and providing effective data support for subsequent accurate registration processing.

[0097] Figure 6 An implementation flowchart of the three-dimensional visual navigation method based on multi-sensor information fusion provided in the sixth embodiment of the present application is shown, which is different from the first embodiment described above in that the step S104 specifically includes:

[0098] In step S601, based on a preset image feature aggregation window, the virtual target image feature information and the scene video frame feature information are respectively aggregated to obtain a plurality of virtual target image feature aggregation information and a plurality of scene video frame feature aggregation information.

[0099] In this embodiment, the preset image feature aggregation window can be set by a person, and can be a window with consistent length and width, and the length and width can both be the distance of three elements. The virtual target image feature information and the scene video frame feature information can be aggregated by a sliding window based on the preset image feature aggregation window, the mean or maximum value of the feature vector in the window is calculated, and a plurality of local aggregation features are obtained.

[0100] In step S602, the similarity of the plurality of virtual target image feature aggregation information and the plurality of scene video frame feature aggregation information is calculated to obtain a plurality of image feature aggregation information local similarity information.

[0101] In this embodiment, the cosine similarity of the virtual target image feature aggregation information and the scene video frame feature aggregation information can be calculated, and the cosine similarity is taken as the image feature aggregation information local similarity information.

[0102] In step S603, it is judged whether the local similarity information of the image feature aggregation information is greater than the preset image local feature similarity threshold information. If yes, the virtual target image feature aggregation information corresponding to the local similarity information of the image feature aggregation information and the scene video frame feature aggregation information are obtained as the virtual target image registration information and the scene target image feature registration information. If no, the virtual target image feature aggregation information corresponding to the local similarity information of the image feature aggregation information and the scene video frame feature aggregation information are skipped.

[0103] In the embodiment, the preset image local feature similarity threshold can be artificially set and can be 0.7. When the local similarity information of the image feature aggregation information is greater than the preset image local feature similarity threshold information, the virtual target image feature aggregation information corresponding to the local similarity information of the image feature aggregation information and the scene video frame feature aggregation information are taken as the feature pair of successful registration, and the feature pair is taken as the virtual target image registration information and the scene target image feature registration information. When the local similarity information of the image feature aggregation information is less than or equal to the preset image local feature similarity threshold information, the virtual target image feature aggregation information corresponding to the local similarity information of the image feature aggregation information and the scene video frame feature aggregation information do not need to be taken as the feature pair of successful registration, and thus the virtual target image feature aggregation information corresponding to the local similarity information of the image feature aggregation information and the scene video frame feature aggregation information can be ignored.

[0104] The three-dimensional visual navigation method based on multi-sensor information fusion provided in the embodiment divides the virtual target image feature information and the scene video frame feature information into a plurality of local regions, and then suppresses the interference of a single noise point through the aggregation processing based on the image feature aggregation window, enhances the spatial consistency of the features, filters the low-confidence matching pairs through the preset image local feature similarity threshold information, effectively eliminates the false matching caused by the local texture similarity, and thus decomposes the global registration into the independent matching of a plurality of local regions, reduces the calculation complexity, improves the reliability of the matching, and improves the accuracy, robustness and spatial consistency of the registration result.

[0105] Figure 7 An implementation flowchart of the three-dimensional visual navigation method based on multi-sensor information fusion provided in the seventh embodiment of the application is shown, which is different from the first embodiment in that:

[0106] The multi-sensor measurement information includes acceleration measurement information and angular velocity measurement information.

[0107] The step S105 specifically includes:

[0108] Step S701, extract the multi-dimensional coordinate information of the virtual target image registration information to obtain virtual target image registration coordinate range information.

[0109] In this embodiment, the three-dimensional coordinate range can be extracted from the virtual target image registration information as the virtual target image registration coordinate range information.

[0110] Step S702, extract the pixel coordinate information of the scene target image feature registration information to obtain a plurality of scene target image feature registration coordinate information and scene target image feature registration coordinate range information.

[0111] In this embodiment, the two-dimensional pixel coordinates and coordinate range of all feature points can be extracted from the scene target image feature registration information as the scene target image feature registration coordinate information and range information.

[0112] Step S703, according to the scene target image feature registration coordinate range information, the virtual target image registration coordinate range information is scaled to generate virtual target image registration scaling information; the virtual target image registration scaling information includes a plurality of virtual target image registration scaling coordinate information.

[0113] In this embodiment, according to the aspect ratio of the scene target image feature registration coordinate range, the virtual target image registration coordinate range can be scaled in proportion to generate the scaled virtual target image registration scaling coordinate information.

[0114] Step S704, according to the acceleration measurement information and the angular velocity measurement information, the image feature angle change information and the image feature translation change information are generated.

[0115] In this embodiment, the acceleration measurement information can be integrated to obtain the image feature translation change information, and the image feature angle change information can be calculated by quaternion integral combined with the angular velocity measurement information, so as to obtain the image feature translation change and the angle change.

[0116] Step S705, according to the image feature angle change information and the image feature translation change information, the coordinate transformation processing is performed on a plurality of virtual target image registration scaling coordinate information to generate virtual target image transformation information.

[0117] In this embodiment, the scaled virtual target coordinate can be translated and rotated to generate the virtual target image transformation information, so as to realize the position compensation in the dynamic scene.

[0118] The three-dimensional visual navigation method based on multi-sensor information fusion provided in the embodiments of the present application solves the scale adjustment problem of virtual targets and real objects by using the scale adjustment of virtual target image registration coordinate range information and scene target image feature registration coordinate range information, generates image feature angle change information and translation change information through integral operation based on the motion compensation mechanism of acceleration measurement information and angular velocity measurement information, to update the virtual target image registration scaling coordinate information in real time, effectively compensates the registration drift caused by device motion, thereby combining the motion sensing capability of the sensor and the geometric transformation of the virtual target, realizing the consistency maintenance of virtual and real spaces in a dynamic scene, and significantly improving the accuracy and stability of navigation.

[0119] Figure 8 An implementation flowchart of the three-dimensional visual navigation method based on multi-sensor information fusion provided in the eighth embodiment of the present application is shown, which is different from the first embodiment described above in that the step S106 specifically includes:

[0120] In step S801, the virtual target image transformation information and the scene target image feature registration information are converted in format and spliced to generate image registration joint feature vector information.

[0121] In this embodiment, the virtual target image transformation information and the scene target image feature registration information are converted into one-dimensional feature vectors and then spliced to generate joint feature vector information, to fuse virtual and real features.

[0122] In step S802, the multi-sensor measurement information is encoded and dimensionally transformed according to the dimension information of the image registration joint feature vector information, to obtain sensor measurement feature vector information.

[0123] In this embodiment, the sensor measurement information can be encoded into a vector with the same dimension as the joint feature vector, or the sensor measurement information can be used as the independent variable of the ReLU function, and then the dimension transformation and encoding conversion processing is performed through the calculation of the ReLU function, to obtain the sensor measurement feature vector information.

[0124] In step S803, the scene target tracking information is obtained according to the image registration joint feature vector information, the sensor measurement feature vector information, and a preset feature vector tracking calculation function, so that the virtual target image transformation information is navigated based on the scene target tracking information to generate virtual target visual information.

[0125] In the embodiment, the preset feature vector tracking calculation function can be artificially set, can be designed based on the combination of the weight coefficient and the bias term, or can be a Transformer encoder for capturing the time dependence in the sequence data of the image registration joint feature vector information and the sensor measurement feature vector information. The image registration joint feature vector information and the sensor measurement feature vector information can be taken as input information of the feature vector tracking calculation function, and the target position prediction value is output to obtain the scene target tracking information in combination with the initial coordinates, and then drive the virtual target to update the pose.

[0126] The three-dimensional visual navigation method based on multi-sensor information fusion provided by the embodiments of the present application can realize cross-domain feature expression of the target by transforming and splicing the virtual target image transformation information and the scene target image feature registration information to construct joint feature vector information containing virtual and real domain features. The encoding and dimension transformation of the sensor measurement information enable different sensor measurement information and visual features to interact in the same space, making up for the deficiency of pure visual tracking in occlusion or low-texture scenes. The preset feature vector tracking calculation function captures the long-term dependence between time features, realizes the recognition of complex motion trajectories, and deeply integrates the spatial positioning capability of visual features and the motion prediction capability of sensors. The generated scene target tracking information can reflect the real-time position of the target and predict its future motion trajectory, ensuring that the dynamic navigation of the virtual target image transformation information has accuracy, real-time performance and stability.

[0127] The method of the above embodiment, Figure 9 A structure block diagram of the three-dimensional visual navigation system based on multi-sensor information fusion provided by the embodiments of the present application is shown. For ease of illustration, only parts related to the embodiments of the present application are shown. Figure 9 The three-dimensional visual navigation system based on multi-sensor information fusion can be an execution subject of the three-dimensional visual navigation method based on multi-sensor information fusion provided by the first embodiment.

[0128] Referring to Figure 9 The three-dimensional visual navigation system based on multi-sensor information fusion includes:

[0129] The information acquisition module 910 is configured to acquire a plurality of scene video frame information and a plurality of sensor measurement information.

[0130] The virtual target image information generation module 920 is configured to generate at least one virtual target image information.

[0131] The image feature information generation module 930 is configured to perform feature extraction and analysis processing on the virtual target image information and the scene video frame information based on a plurality of preset image information feature processing functions, to obtain virtual target image feature information and scene video frame feature information.

[0132] The image registration information generation module 940 is configured to perform registration processing on the virtual target image feature information and the scene video frame feature information, to obtain virtual target image registration information and scene target image feature registration information.

[0133] The virtual target image transformation information generation module 950 is configured to perform image fusion processing on the virtual target image registration information and the scene target image feature registration information based on a plurality of sensor measurement information, to generate virtual target image transformation information.

[0134] The virtual target visualization information generation module 960 is configured to perform visualization tracking calculation on the scene target image feature registration information based on a plurality of sensor measurement information and virtual target image transformation information, to obtain scene target tracking information, so that the virtual target image transformation information is navigated based on the scene target tracking information, to generate virtual target visualization information.

[0135] The process in which each module of the three-dimensional visualization navigation system based on multi-sensor information fusion provided in the embodiments of the present application realizes its own function is specifically referable to the description of the aforementioned embodiment one, and will not be described herein again. Figure 1

[0136] It should be understood that the size of the serial number of each step in the above embodiments does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0137] It should be understood that when used in the present application and the appended claims, the term "comprising" indicates the presence of the described features, whole, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, whole, steps, operations, elements, components and / or sets thereof.

[0138] It should also be understood that the term "and / or" used in the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0139] ​As used in the specification and in the claims, the term “if’ can be interpreted as meaning “when” or “upon” or “in response to a determination” or “in response to a detection” depending on the context. Similarly, the phrase “if it is determined” or “if [the described condition or event] is detected” can be interpreted as meaning “upon a determination” or “in response to a determination” or “upon a detection of [the described condition or event]” or “in response to a detection of [the described condition or event]” depending on the context.

[0140] In addition, in the description and the accompanying claims of the application, the terms “first”, “second”, “third”, and the like are used merely to distinguish descriptions and are not intended to imply or suggest relative importance. It should also be understood that, although the terms “first”, “second”, and the like are used in the text to describe various elements in some embodiments of the application, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, the first table can be named as the second table, and similarly, the second table can be named as the first table, without departing from the scope of various described embodiments. The first table and the second table are both tables, but they are not the same table.

[0141] In the present application, the reference to “one embodiment” or “some embodiments” and the like means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. Thus, the appearance of the phrases “in one embodiment”, “in some embodiments”, “in other embodiments”, “in additional embodiments”, and the like, in various places throughout the specification is not necessarily all referring to the same embodiment, unless otherwise specifically noted. The terms “comprise”, “include”, “have” and their conjugates mean “including but not limited to”, unless otherwise specifically noted.

[0142] The three-dimensional visualization navigation method based on multi-sensor information fusion provided by the embodiments of the application can be applied to terminal devices such as tablet computers, augmented reality (AR) / virtual reality (VR) devices, notebook computers, ultra-mobile personal computers (UMPC), netbooks, personal digital assistants (PDA), etc. The embodiments of the application do not make any limitation on the specific type of terminal device.

[0143] For example, the terminal device may be a station (STAION, ST) in a WLAN, a cellular phone, a cordless phone, a Session Initiation Protocol (SIP) phone, a Wireless Local Loop (WLL) station, a Personal Digital Assistant (PDA) device, a handheld device with wireless communication capabilities, a computing device or other processing device connected to a wireless modem, an in-vehicle device, a vehicle networking terminal, a computer, a laptop computer, a handheld communication device, a handheld computing device, a satellite wireless device, a wireless modem card, a set-top box (STB), customer premises equipment (CPE), and / or other devices used for communication over a wireless system, as well as next-generation communication systems, such as mobile terminals in 5G networks or mobile terminals in future evolved Public Land Mobile Network (PLMN) networks.

[0144] As an example and not a limitation, when the terminal device is a wearable device, the term "wearable device" can also refer to any device that utilizes wearable technology to intelligently design and develop everyday wearables, such as glasses, gloves, watches, clothing, and shoes. Wearable devices are portable devices worn directly on the body or integrated into a user's clothing or accessories. Wearable devices are not merely hardware devices; they achieve powerful functions through software support, data interaction, and cloud interaction. Broadly defined, wearable smart devices include those with comprehensive functions, large sizes, and the ability to perform complete or partial functions without relying on a smartphone, such as smartwatches or smart glasses, as well as those focused on a specific application function that require interaction with other devices such as smartphones, such as various smart bracelets and smart jewelry for vital sign monitoring.

[0145] Figure 10 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. For example... Figure 10 As shown, the terminal device 100 of this embodiment includes: at least one processor 1000 ( Figure 10 (Only one is shown in the image) A memory 1001 stores a computer program 1002 that can run on the processor 1000. When the processor 1000 executes the computer program 1002, it implements the steps in the various embodiments of the three-dimensional visualization navigation method based on multi-sensor information fusion described above, for example... Figure 1The processor 1000 performs the steps S101-S106 illustrated. Alternatively, the processor 1000 implements functions of various modules / units in the above-described various system embodiments when the processor 1000 executes the computer program 1002, for example. Figure 9 The functions of the modules 910-960 are described above.

[0146] The terminal device 100 can be a desktop computer, a notebook computer, a palm computer, a cloud server, and the like. The terminal device can include, but is not limited to, the processor 1000, the memory 1001. Those skilled in the art can understand that the terminal device 100 can include more or less components, or combine some components, or include different components, for example, the terminal device can further include an input / output device, a network access device, a bus, and the like. Figure 10 The terminal device 100 is only an example and does not constitute a limitation on the terminal device 100, and can include more or less components, or combine some components, or include different components, for example, the terminal device can further include an input / output device, a network access device, a bus, and the like.

[0147] The processor 1000 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0148] The memory 1001 can be an internal storage unit of the terminal device 100 in some embodiments, for example, a hard disk or a memory of the terminal device 100. The memory 1001 can also be an external storage device of the terminal device 100, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 1001 can include both the internal storage unit and the external storage device of the terminal device 100. The memory 1001 is used to store an operating system, application programs, a boot loader, data, and other programs, for example, program codes of the computer program, etc. The memory 1001 can also be used to temporarily store data that has been or will be transmitted.

[0149] In addition, each of the function units in each of the embodiments of the present application can be integrated in one processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function unit.

[0150] The embodiments of the present application further provide a terminal device, which comprises at least one memory, at least one processor, and a computer program stored in the at least one memory and executable on the at least one processor, and the processor executes the computer program to enable the terminal device to implement the steps in any of the above method embodiments.

[0151] The integrated module / unit, if realized in the form of a software function unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on such understanding, all or part of the flow of the above-mentioned embodiment methods can also be implemented by a computer program instructing related hardware to complete, and the computer program can be stored in a computer-readable storage medium. The computer program can implement the steps of each method embodiment when executed by a processor. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms. The computer-readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0152] In the above embodiments, the description of each embodiment has its own focus, and the parts not described or recorded in detail in a certain embodiment can be referred to the relevant description of other embodiments.

[0153] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0154] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, may be located in one place, or may also be distributed to multiple network units. Part or all of the units can be selected to achieve the purpose of the embodiment scheme according to actual needs.

[0155] The above embodiments are only used to illustrate the technical solutions of the present application, but not limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can still be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A three-dimensional visual navigation method based on multi-sensor information fusion, characterized in that, The method comprises the following steps: acquiring a plurality of scene video frame information and a plurality of sensor measurement information; generating at least one virtual target image information; after format conversion of the virtual target image information and the scene video frame information, first scale virtual target pixel matrix variable information and second scale virtual target pixel matrix variable information, first scale scene video frame pixel matrix variable information and second scale scene video frame pixel matrix variable information are obtained according to a preset first scale image information feature processing function and a preset second scale image information feature processing function, difference calculation is performed respectively, scale difference virtual target pixel matrix variable information and scale difference scene video frame pixel matrix variable information are obtained, extreme values are extracted respectively, virtual target image feature center information and scene video frame feature center information are obtained, and virtual target image feature extraction vector information and scene video frame feature extraction vector information are generated in combination with preset image information feature extraction radius information; first virtual target image feature vector transformation information and first scene video frame feature vector transformation information are obtained according to the virtual target image feature extraction vector information, the scene video frame feature extraction vector information and a preset first image information feature vector processing function, second virtual target image feature vector transformation information and second scene video frame feature vector transformation information are obtained in combination with a preset second image information feature vector processing function, third virtual target image feature vector transformation information and third scene video frame feature vector transformation information are obtained in combination with a preset third image information feature vector processing function, virtual target image feature vector representation information and scene video frame feature vector representation information are calculated, the virtual target image feature vector representation information and the scene video frame feature vector representation information are scaled to obtain virtual target image feature vector dimension transformation information and scene video frame feature vector dimension transformation information according to a preset feature vector representation information scaling coefficient, virtual target feature representation variable and scene feature representation variable are obtained in combination with preset feature vector coefficient information and preset feature vector compensation information, the virtual target image feature vector dimension reduction information and the scene video frame feature vector dimension reduction information are obtained by performing dimension reduction processing in combination with a preset feature vector representation information scaling coefficient, and finally, virtual target image feature information and scene video frame feature information are obtained in combination with preset dimension reduction feature vector coefficient information and preset dimension reduction feature vector compensation information; registration processing is performed according to the virtual target image feature information and the scene video frame feature information, virtual target image registration information and scene target image feature registration information are obtained; image fusion processing is performed on the virtual target image registration information and the scene target image feature registration information according to a plurality of the sensor measurement information, and virtual target image transformation information is generated; According to the virtual target image feature information and the scene video frame feature information, registration processing is performed to obtain virtual target image registration information and scene target image feature registration information.

2. The multi-sensor information fusion based three-dimensional visual navigation method of claim 1, wherein, The step of performing registration processing according to the virtual target image feature information and the scene video frame feature information to obtain virtual target image registration information and scene target image feature registration information specifically includes: Based on a preset image feature aggregation window, the virtual target image feature information and the scene video frame feature information are respectively aggregated to obtain a plurality of virtual target image feature aggregation information and a plurality of scene video frame feature aggregation information; The similarity of the plurality of virtual target image feature aggregation information and the plurality of scene video frame feature aggregation information is calculated to obtain a plurality of image feature aggregation information local similarity information; When the image feature aggregation information local similarity information is greater than a preset image local feature similarity threshold information, then according to the virtual target image feature aggregation information and the scene video frame feature aggregation information corresponding to the image feature aggregation information local similarity information, virtual target image registration information and scene target image feature registration information are obtained.

3. The three-dimensional visual navigation method based on multi-sensor information fusion according to claim 1, wherein The plurality of sensor measurement information includes acceleration measurement information and angular velocity measurement information; The step of performing image fusion processing on the virtual target image registration information and the scene target image feature registration information according to the plurality of sensor measurement information to generate virtual target image transformation information specifically includes: Multi-dimensional coordinate information of the virtual target image registration information is extracted to obtain virtual target image registration coordinate range information; Pixel coordinate information of the scene target image feature registration information is extracted to obtain a plurality of scene target image feature registration coordinate information and scene target image feature registration coordinate range information; According to the scene target image feature registration coordinate range information, scaling processing is performed on the virtual target image registration coordinate range information to generate virtual target image registration scaling information; the virtual target image registration scaling information includes a plurality of virtual target image registration scaling coordinate information; According to the acceleration measurement information and the angular velocity measurement information, image feature angle change information and image feature translation change information are generated; According to the image feature angle change information and the image feature translation change information, coordinate transformation processing is performed on the plurality of virtual target image registration scaling coordinate information to generate virtual target image transformation information.

4. The multi-sensor information fusion based three-dimensional visual navigation method of claim 1, wherein, The step of performing visual tracking calculation on the scene target image feature registration information according to the plurality of sensor measurement information and the virtual target image transformation information to obtain scene target tracking information, so that the virtual target image transformation information is navigated based on the scene target tracking information to generate virtual target visual information specifically includes: The virtual target image transformation information and the scene target image feature registration information are format-converted and spliced to generate image registration joint feature vector information; According to the dimension information of the image registration joint feature vector information, the multiple sensor measurement information is encoded and dimensionally transformed to obtain sensor measurement feature vector information; According to the image registration joint feature vector information, the sensor measurement feature vector information, and a preset feature vector tracking calculation function, scene target tracking information is obtained, so that the virtual target image transformation information is navigated based on the scene target tracking information to generate virtual target visualization information.

5. A three-dimensional visual navigation system based on multi-sensor information fusion, characterized in that, It comprises: An information acquisition module is configured to acquire multiple scene video frame information and multiple sensor measurement information; A virtual target image information generation module is configured to generate at least one virtual target image information; An image feature information generation module is configured to, after format-converting the virtual target image information and the scene video frame information, obtain first-scale virtual target pixel matrix variable information and second-scale virtual target pixel matrix variable information, first-scale scene video frame pixel matrix variable information and second-scale scene video frame pixel matrix variable information according to a preset first-scale image information feature processing function and a preset second-scale image information feature processing function, respectively perform difference calculation to obtain scale-difference virtual target pixel matrix variable information and scale-difference scene video frame pixel matrix variable information, respectively extract extreme values to obtain virtual target image feature center information and scene video frame feature center information, and combine preset image information feature extraction radius information to generate virtual target image feature extraction vector information and scene video frame feature extraction vector information. The first virtual target image feature vector transformation information and the first scene video frame feature vector transformation information are obtained by combining a preset second image information feature vector processing function, the third virtual target image feature vector transformation information and the third scene video frame feature vector transformation information are obtained by combining a preset third image information feature vector processing function, and the virtual target image feature vector representation information and the scene video frame feature vector representation information are calculated; the virtual target image feature vector representation information and the scene video frame feature vector representation information are scaled to obtain the virtual target image feature vector dimension transformation information and the scene video frame feature vector dimension transformation information according to a preset feature vector representation information scaling coefficient, the virtual target feature representation variable and the scene feature representation variable are obtained by combining a preset feature vector coefficient information and a preset feature vector compensation information, the virtual target image feature vector dimension reduction information and the scene video frame feature vector dimension reduction information are obtained by combining the preset feature vector representation information scaling coefficient for dimension reduction processing, and finally the virtual target image feature information and the scene video frame feature information are obtained by combining a preset dimension reduction feature vector coefficient information and a preset dimension reduction feature vector compensation information. An image registration information generation module is configured to perform registration processing on the virtual target image feature information and the scene video frame feature information to obtain virtual target image registration information and scene target image feature registration information. A virtual target image transformation information generation module is configured to perform image fusion processing on the virtual target image registration information and the scene target image feature registration information according to the sensor measurement information to generate virtual target image transformation information. A virtual target visualization information generation module is configured to perform visual tracking calculation on the scene target image feature registration information according to the sensor measurement information and the virtual target image transformation information to obtain scene target tracking information, so that the virtual target image transformation information is navigated based on the scene target tracking information to generate virtual target visualization information.

6. A terminal device, characterized by comprising: The terminal device comprises a memory and a processor, the memory stores a computer program capable of running on the processor, and the processor implements the steps of the method according to any one of claims 1 to 4 when executing the computer program.

Citation Information

Patent Citations

  • Three-dimensional visual navigation method based on multi-sensor information fusion

    CN101833104A

  • Large-scene cross-border head target tracking method and system based on three-dimensional geographic information

    CN110930507A