Three-dimensional visual navigation method and system based on multi-sensor information fusion

Through the multi-sensor information fusion method, the problem of low fusion accuracy between virtual targets and real scenes and easy to lose tracking is solved, real-time and accurate virtual target visual navigation is achieved, and the physical consistency and interactive perception ability between virtual targets and real scenes is enhanced.

CN120544103AActive Publication Date: 2025-08-26ANHUI RUIXIANG DIGITAL TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510686499.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-08-26
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

In the prior art, the registration based on image feature points is easily disturbed by light changes, occlusions and similar textures, resulting in low fusion accuracy between virtual targets and real scenes, easy to lose tracking of virtual targets, and inaccurate navigation.

Method used

The multi-sensor information fusion method is adopted to obtain multiple scene video frame information and sensor measurement information, generate virtual target image information, and use image information feature processing functions for feature extraction and analysis, and combine multi-sensor measurement information for registration and fusion to achieve refined description and spatial alignment of virtual targets and real scenes, enhance physical consistency, and perform visual tracking and calculations to achieve dynamic navigation.

Benefits of technology

The accuracy and robustness of interaction perception between virtual targets and environments are improved, real-time, accurate and spatially consistent virtual target visualization information is generated, and the accuracy and robustness of interaction perception between the targets and environments are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544103A_ABST
    Figure CN120544103A_ABST
Patent Text Reader

Abstract

The invention provides a three-dimensional visual navigation method and system based on multi-sensor information fusion, which are suitable for the technical field of image processing, and the method comprises the following steps: obtaining virtual target image feature information and scene video frame feature information according to virtual target image information and scene video frame information, performing registration processing to obtain virtual target image registration information and scene target image feature registration information; generating virtual target image transformation information according to the multi-sensor measurement information, the virtual target image registration information and the scene target image feature registration information; and according to the multi-sensor measurement information, the virtual target image transformation information and the scene target image feature registration information, scene target tracking information is obtained, so that the virtual target image transformation information is navigated based on the scene target tracking information, and virtual target visualization information is generated. According to the invention, the real-time positioning precision of the virtual target in a dynamic scene can be improved, so that accurate navigation of visual information of the virtual target is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of image processing technology, and in particular relates to a three-dimensional visualization navigation method and system based on multi-sensor information fusion. Background Art

[0002] In today's digital age, technological applications that cleverly integrate virtual information with the real world are expanding into diversified application scenarios as hardware equipment performance improves. Application scenarios are gradually expanding rapidly from consumer-grade entertainment to professional fields such as industrial maintenance, medical navigation, and intelligent security.

[0003] In the existing technology, the optical flow method is usually used to track dynamically changing targets in the scene to determine the approximate position of the virtual target in the real scene, and then perform alignment processing. After the alignment is completed, the virtual target is superimposed on the scene image.

[0004] However, in the existing technology, the registration based solely on image feature points is easily affected by factors such as lighting changes, occlusion, and similar textures in the scene, resulting in low registration accuracy, poor fusion of virtual targets and real scenes, and the occurrence of misalignment, drift, etc. Moreover, during the tracking process, tracking is only based on image information. When the target temporarily leaves the field of view or the scene changes significantly, the tracking is easily lost, and it is difficult to achieve stable and continuous visual tracking. As a result, the navigation of the virtual target based on the scene target tracking information is not accurate enough, and it is impossible to provide users with smooth and reliable virtual target visualization information. Summary of the Invention

[0005] In view of this, the embodiments of the present application provide a three-dimensional visualization navigation method and system based on multi-sensor information fusion, which aims to solve the problems existing in the prior art when virtual and reality are merged, such as low accuracy and susceptibility to interference, easy loss of tracking, and inaccurate navigation based solely on image feature points.

[0006] A first aspect of an embodiment of the present application provides a three-dimensional visualization navigation method based on multi-sensor information fusion, comprising:

[0007] Obtain multiple scene video frame information and multiple sensor measurement information;

[0008] generating at least one virtual target image information;

[0009] Based on a plurality of preset image information feature processing functions, feature extraction and analysis processing are performed on the virtual target image information and the scene video frame information to obtain virtual target image feature information and scene video frame feature information;

[0010] Performing registration processing based on the virtual target image feature information and the scene video frame feature information to obtain virtual target image registration information and scene target image feature registration information;

[0011] performing image fusion processing on the virtual target image registration information and the scene target image feature registration information according to the plurality of sensor measurement information to generate virtual target image transformation information;

[0012] According to the plurality of sensor measurement information and the virtual target image transformation information, the scene target image feature registration information is visually tracked and calculated to obtain scene target tracking information, so that the virtual target image transformation information is navigated based on the scene target tracking information to generate virtual target visualization information.

[0013] A second aspect of the embodiments of the present application provides a three-dimensional visual navigation system based on multi-sensor information fusion, including:

[0014] An information acquisition module, used to acquire multiple scene video frame information and multiple sensor measurement information;

[0015] A virtual target image information generating module, configured to generate at least one virtual target image information;

[0016] An image feature information generation module is used to perform feature extraction and analysis processing on the virtual target image information and the scene video frame information based on a plurality of preset image information feature processing functions to obtain virtual target image feature information and scene video frame feature information;

[0017] An image registration information generating module is used to perform registration processing based on the virtual target image feature information and the scene video frame feature information to obtain virtual target image registration information and scene target image feature registration information;

[0018] a virtual target image transformation information generating module, configured to perform image fusion processing on the virtual target image registration information and the scene target image feature registration information based on the plurality of sensor measurement information, and generate virtual target image transformation information;

[0019] A virtual target visualization information generation module is used to perform visual tracking calculation on the scene target image feature registration information based on the multiple sensor measurement information and virtual target image transformation information to obtain scene target tracking information, so that the virtual target image transformation information can be navigated based on the scene target tracking information to generate virtual target visualization information.

[0020] The third aspect of an embodiment of the present application provides a terminal device, which includes a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, it implements the steps of the three-dimensional visualization navigation method based on multi-sensor information fusion as described in the first aspect above.

[0021] Compared with the prior art, the beneficial effects of the embodiments of the present application are as follows: the present application obtains measurement information of multiple sensors to enrich environmental perception data, generates virtual target image information, and performs feature extraction and analysis on the virtual target image information and scene video frame information through image information feature processing functions to achieve a refined description of the virtual target and the real scene picture. The feature-based registration processing enables the virtual target image and the scene target image to be spatially aligned, and the registered virtual and real images are fused in combination with multi-sensor measurement information to enhance the physical consistency between the virtual target transformation information and the real scene. Then, the scene target feature registration information is visualized and tracked and calculated through multi-sensor data and virtual target image transformation information, thereby achieving dynamic navigation of the virtual target based on the scene target tracking information, thereby generating real-time, accurate and spatially consistent virtual target visualization information, thereby significantly improving the interactive perception accuracy and robustness of identifying targets and environmental conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0023] Figure 1 This is a schematic diagram of the implementation process of the three-dimensional visual navigation method based on multi-sensor information fusion provided in Example 1 of the present application;

[0024] Figure 2 This is a schematic diagram of the implementation process of the three-dimensional visual navigation method based on multi-sensor information fusion provided in Example 2 of the present application;

[0025] Figure 3 This is a schematic diagram of the implementation process of the three-dimensional visual navigation method based on multi-sensor information fusion provided in Example 3 of the present application;

[0026] Figure 4 This is a schematic diagram of the implementation process of the three-dimensional visual navigation method based on multi-sensor information fusion provided in Example 4 of the present application;

[0027] Figure 5This is a schematic diagram of the implementation process of the three-dimensional visual navigation method based on multi-sensor information fusion provided in Example 5 of the present application;

[0028] Figure 6 This is a schematic diagram of the implementation process of the three-dimensional visual navigation method based on multi-sensor information fusion provided in Example 6 of the present application;

[0029] Figure 7 This is a schematic diagram of the implementation process of the three-dimensional visual navigation method based on multi-sensor information fusion provided in Example 7 of the present application;

[0030] Figure 8 This is a schematic diagram of the implementation process of the three-dimensional visual navigation method based on multi-sensor information fusion provided in Example 8 of the present application;

[0031] Figure 9 Schematic diagram of the structure of a three-dimensional visual navigation system based on multi-sensor information fusion provided in an embodiment of the present application;

[0032] Figure 10 It is a schematic diagram of the terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0033] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0034] In order to illustrate the technical solution described in this application, specific embodiments are provided below.

[0035] Figure 1 The following is a flowchart of the implementation of the three-dimensional visualization navigation method based on multi-sensor information fusion provided in Example 1 of the present application, which is detailed as follows:

[0036] Step S101: Acquire multiple scene video frame information and multiple sensor measurement information.

[0037] In this embodiment, the scene video frame information may be obtained by photographing the scene with a photographic device, may be obtained by photographing with an RGB camera, may be obtained by photographing with a depth camera, or may be obtained by collecting continuous video frames output by the photographic device at a fixed frame rate. The multiple sensors may include an inertial measurement unit (IMU), a GPS, a depth sensor (such as a Kinect sensor). The multiple sensor measurement information may include acceleration information, angular velocity information, etc. collected at a high frequency by the inertial measurement unit, may include the longitude and latitude information of the shooting location obtained by the GPS, may be the distance between the object and the device measured by the depth sensor, and then the depth map information generated based on the distance. The video frame and the sensor data may be aligned by timestamps, for example, using hardware synchronization triggering or software interpolation algorithms to compensate for sampling frequency differences.

[0038] Step S102: Generate at least one virtual target image information.

[0039] In this embodiment, a three-dimensional model of a virtual object can be manually designed using modeling software (such as Blender, Maya), and the output can be in formats such as .obj and .stl, which can include vertex coordinates, texture maps, and triangular mesh surface information. The three-dimensional model can then be loaded through a rendering engine, and a two-dimensional image of the virtual target can be generated by manually setting parameters such as lighting and materials.

[0040] Step S103 : Based on a plurality of preset image information feature processing functions, feature extraction and analysis processing are performed on the virtual target image information and the scene video frame information to obtain virtual target image feature information and scene video frame feature information.

[0041] In this embodiment, the plurality of preset image information feature processing functions may be manually designed, and may be designed based on Gaussian functions, and are used to perform feature extraction and analysis processing on the virtual target image information and the scene video frame information. The feature extraction and analysis processing may include first converting the virtual target image information and the scene video frame information into a vector form, and then using the vector as an independent variable of the plurality of preset image information feature processing functions. The plurality of preset image information feature processing functions are then calculated to output a plurality of virtual target image feature information and a plurality of scene video frame feature information.

[0042] Step S104 : performing registration processing according to the virtual target image feature information and the scene video frame feature information to obtain virtual target image registration information and scene target image feature registration information.

[0043] In this embodiment, the cosine similarity of the virtual target image feature information and the scene video frame feature information can be calculated, and feature matching can be completed when the cosine similarity is greater than 0.8, that is, the registration of the virtual target image feature information and the scene video frame feature information is successfully achieved, and then the successfully registered virtual target image feature information and the scene video frame feature information are registered as virtual target image registration information and scene target image feature registration information.

[0044] Step S105 : performing image fusion processing on the virtual target image registration information and the scene target image feature registration information according to the plurality of sensor measurement information to generate virtual target image transformation information.

[0045] In this embodiment, the acceleration and angular velocity measured by the IMU can be integrated to predict the device's position and motion posture to compensate for motion errors during the alignment process. GPS data and depth values ​​measured by the depth sensor can be integrated to construct a global three-dimensional map of the scene, providing an absolute spatial reference for the virtual target. The virtual target can be rendered based on the predicted device position and motion posture, and the rendered image information can be used as the virtual target image transformation information.

[0046] Step S106, performing visual tracking calculation on the scene target image feature registration information based on the multiple sensor measurement information and the virtual target image transformation information to obtain scene target tracking information, so that the virtual target image transformation information is navigated based on the scene target tracking information to generate virtual target visualization information.

[0047] In this embodiment, feature points in the scene target image feature registration information may be tracked by a feature matching algorithm to obtain the coordinate information of the feature point in the new frame. Data obtained by IMU measurement may be used as state input, and the two-dimensional motion of the feature points and the three-dimensional spatial information may be fused by a particle filtering algorithm to calculate the real motion parameters of the target. The two-dimensional pixel offset of the target is then calculated based on the displacement of the feature points. The position and motion posture change data obtained by predicting and calculating the sensor data are combined to obtain the new position of the target in the image to realize visual tracking processing. The tracked target position is fed back to the posture control module of the virtual target to update its rendering parameters so that the virtual target image transformation information follows the movement of the scene target in real time, and finally dynamically aligned virtual target visualization information is generated.

[0048] The three-dimensional visualization navigation method based on multi-sensor information fusion provided in the embodiment of the present application obtains measurement information of multiple sensors to enrich environmental perception data, generates virtual target image information, and performs feature extraction and analysis on the virtual target image information and scene video frame information through image information feature processing function to achieve a refined description of the virtual target and the real scene picture. The feature-based registration processing enables the virtual target image and the scene target image to be spatially aligned, and the registered virtual and real images are fused in combination with multi-sensor measurement information to enhance the physical consistency between the virtual target transformation information and the real scene. Then, the scene target feature registration information is visualized and tracked and calculated through multi-sensor data and virtual target image transformation information, thereby achieving dynamic navigation of the virtual target based on the scene target tracking information, thereby generating real-time, accurate and spatially consistent virtual target visualization information, thereby significantly improving the interactive perception accuracy and robustness of the identified targets and environmental conditions.

[0049] Figure 2 The following is a flowchart of a three-dimensional visualization navigation method based on multi-sensor information fusion according to the second embodiment of the present application. The difference between the second embodiment and the first embodiment is that:

[0050] The plurality of preset image information feature processing functions include a preset first-scale image information feature processing function, a preset second-scale image information feature processing function, and a plurality of preset image information feature vector processing functions;

[0051] The step S103 specifically includes:

[0052] Step S201 : performing format conversion on the virtual target image information and the scene video frame information to obtain a virtual target pixel matrix and a scene video frame pixel matrix.

[0053] In this embodiment, the virtual target image information and the scene video frame information can be converted into a standard pixel matrix format, and the virtual target image information and the scene video frame information can be read through the OpenCV or PIL library and converted into NumPy arrays, so that the NumPy arrays are used as the virtual target pixel matrix and the scene video frame pixel matrix.

[0054] Step S202: Analyze and calculate the virtual target pixel matrix and the scene video frame pixel matrix according to a preset first-scale image information feature processing function and a preset second-scale image information feature processing function to obtain first-scale virtual target pixel matrix variable information, second-scale virtual target pixel matrix variable information, first-scale scene video frame pixel matrix variable information, and second-scale scene video frame pixel matrix variable information.

[0055] In this embodiment, the preset first-scale image information feature processing function and the preset second-scale image information feature processing function can both be manually set, and can be set based on a Gaussian kernel function, or can be designed based on a radial basis function, or can be designed based on the structure of a weight part and a bias part. The first scale can be larger than the second scale, or the second scale can be larger than the first scale. It is understandable that the description of the first scale and the second scale is merely used to illustrate the analytical calculation of the virtual target pixel matrix and the scene video frame pixel matrix by image information feature processing functions with different processing scales. In actual applications, there may also be a third-scale image information feature processing function, a fourth-scale image information feature processing function, and so on. The virtual target pixel matrix and the scene video frame pixel matrix can be used as the independent variables of the preset first-scale image information feature processing function and the preset second-scale image information feature processing function, respectively. After the calculations of the preset first-scale image information feature processing function and the preset second-scale image information feature processing function, the calculation results are used as the first-scale virtual target pixel matrix variable information, the second-scale virtual target pixel matrix variable information, the first-scale scene video frame pixel matrix variable information, and the second-scale scene video frame pixel matrix variable information, respectively.

[0056] Step S203 : performing difference calculation on the first-scale virtual target pixel matrix variable information and the second-scale virtual target pixel matrix variable information to obtain scale-difference virtual target pixel matrix variable information.

[0057] In this embodiment, when the first-scale virtual target pixel matrix variable information is greater than the second-scale virtual target pixel matrix variable information, a difference calculation is performed by subtracting the second-scale virtual target pixel matrix variable information from the first-scale virtual target pixel matrix variable information to obtain scale-differential virtual target pixel matrix variable information; when the second-scale virtual target pixel matrix variable information is greater than the first-scale virtual target pixel matrix variable information, a difference calculation is performed by subtracting the first-scale virtual target pixel matrix variable information from the second-scale virtual target pixel matrix variable information to obtain scale-differential virtual target pixel matrix variable information.

[0058] Step S204 : performing difference calculation on the first-scale scene video frame pixel matrix variable information and the second-scale scene video frame pixel matrix variable information to obtain scale-difference scene video frame pixel matrix variable information.

[0059] In this embodiment, when the first-scale scene video frame pixel matrix variable information is greater than the second-scale scene video frame pixel matrix variable information, a difference calculation is performed by subtracting the second-scale scene video frame pixel matrix variable information from the first-scale scene video frame pixel matrix variable information to obtain scale difference scene video frame pixel matrix variable information; when the second-scale scene video frame pixel matrix variable information is greater than the first-scale scene video frame pixel matrix variable information, a difference calculation is performed by subtracting the first-scale scene video frame pixel matrix variable information from the second-scale scene video frame pixel matrix variable information to obtain scale difference scene video frame pixel matrix variable information.

[0060] Step S205 , extracting the extreme values ​​of the scale-difference virtual target pixel matrix variable information and the extreme values ​​of the scale-difference scene video frame pixel matrix variable information to obtain virtual target image feature center information and scene video frame feature center information.

[0061] In this embodiment, it can be understood that the scale difference virtual target pixel matrix variable information and the scale difference scene video frame pixel matrix variable information are both in matrix form. The maximum element can be determined from these matrices and the maximum element can be used as the virtual target image feature center information and the scene video frame feature center information.

[0062] Step S206 , obtaining virtual target image feature information and scene video frame feature information according to the virtual target image feature center information, scene video frame feature center information, preset image information feature extraction radius information and multiple preset image information feature vector processing functions.

[0063] In this embodiment, the preset image information feature extraction radius information and multiple preset image information feature vector processing functions can be manually set, wherein the preset image information feature extraction radius information can take a value of 64, and the multiple preset image information feature vector processing functions can be designed based on the structure of weight coefficients and bias terms. The calculation process of the preset image information feature vector processing function can be to first multiply the independent variable by the weight coefficient, and then add the multiplication result to the bias term, so that the addition result is used as the output of the image information feature vector processing function. The virtual target image feature center information and the scene video frame feature center information can be used as the center of the two parts respectively, and the preset image information feature extraction radius information is used as the feature extraction range. The feature points falling within the feature extraction range are extracted, and these feature points are formatted to generate feature vectors, so that the feature vectors are calculated layer by layer by multiple preset image information feature vector processing functions, and the calculation result output by the last image information feature vector processing function is the virtual target image feature information and the scene video frame feature information.

[0064] The three-dimensional visualization navigation method based on multi-sensor information fusion provided in the embodiment of the present application uses preset first-scale and second-scale image information feature processing functions to perform multi-scale analysis on the pixel matrix, capturing fine details and coarse-grained structural features respectively, so that the feature capture of virtual target image information and scene video frame information is more comprehensive and detailed. By performing difference calculation on variable information of pixel matrices of different scales, key features of specific sizes are highlighted, and the adaptability of features to scale changes is enhanced. The feature center is determined by extracting the extreme value of the scale difference matrix variable information to achieve precise positioning of feature points. Feature information is then generated based on the virtual target image feature center information, the scene video frame feature center information, the preset image information feature extraction radius information, and multiple preset image information feature vector processing functions. It is ensured that the feature vector contains both spatial structure information of the local neighborhood and global expression capabilities, thereby achieving stable extraction and description of virtual target and scene video frame features in complex scenes, improving the accuracy and robustness of feature matching, and providing a reliable feature basis for subsequent alignment, fusion and tracking tasks.

[0065] Figure 3 The flowchart of the implementation of the three-dimensional visual navigation method based on multi-sensor information fusion provided in the third embodiment of the present application is shown. The difference between the third embodiment and the second embodiment is that the step S206 specifically includes:

[0066] Step S301 , based on the virtual target image feature center information and the scene video frame feature center information, and according to preset image information feature extraction radius information, obtain multiple virtual target image feature extraction point information and multiple scene video frame feature extraction point information.

[0067] In this embodiment, the preset image information feature extraction radius information can be manually set and can be an interval distance of 64 elements. With the virtual target image feature center information and the scene video frame feature center information as the center of the circle, and the preset image information feature extraction radius information as the radius, a center area is generated. The coordinates of all pixel points within the area are extracted by coordinate traversal, and are used as multiple virtual target image feature extraction point information and multiple scene video frame feature extraction point information.

[0068] Step S302 : generating virtual target image feature extraction vector information and scene video frame feature extraction vector information based on the plurality of virtual target image feature extraction point information and the plurality of scene video frame feature extraction point information.

[0069] In this embodiment, the pixel values ​​of each virtual target image feature extraction point are expanded into a one-dimensional vector in channel order (such as RGB) and spliced ​​into virtual target image feature extraction vector information; similarly, the pixel values ​​of the scene video frame feature extraction point are converted into scene video frame feature extraction vector information, and the vector length is the number of extraction points × the number of channels.

[0070] Step S303 , obtaining virtual target image feature information and scene video frame feature information according to the virtual target image feature extraction vector information, the scene video frame feature extraction vector information and a plurality of preset image information feature vector processing functions.

[0071] In this embodiment, the virtual target image feature extraction vector information and the scene video frame feature extraction vector information are respectively input into a plurality of preset image information feature vector processing functions. The results output after layer-by-layer calculations by a plurality of preset image information feature vector processing functions are weighted summed, and then the weighted summation results can be normalized by the ReLU function, and finally the virtual target image feature information and scene video frame feature information of fixed dimensions are output.

[0072] The three-dimensional visualization navigation method based on multi-sensor information fusion provided in the embodiment of the present application uses a feature extraction mechanism of feature center positioning and radius constraint, takes the feature center information of the virtual target image and the feature center information of the scene video frame as spatial anchor points, and combines the preset image information feature extraction radius information to ensure that the feature extraction points accurately cover the local structure of the target, so that the generated virtual target image feature extraction vector information and scene video frame feature extraction vector information contain both the core response value of the feature point and its neighborhood spatial contextual relationship, such as texture distribution, gradient direction, etc., avoiding the one-sidedness of single pixel features. Through the layer-by-layer calculation of multiple preset image information feature vector processing functions, the feature information is gradually abstracted from the original pixel value to a discriminative high-level semantic representation, thereby improving the robustness of the feature information to scale changes and spatial position changes, and reducing local noise interference through the global expression ability of the feature vector, providing a feature basis with both detail accuracy and global consistency for subsequent registration, which is used to comprehensively improve the registration accuracy and stability.

[0073] Figure 4 The following is a flowchart of a three-dimensional visualization navigation method based on multi-sensor information fusion according to the fourth embodiment of the present application. The difference between the fourth embodiment and the third embodiment is that:

[0074] The plurality of preset image information feature vector processing functions include a preset first image information feature vector processing function, a preset second image information feature vector processing function, and a preset third image information feature vector processing function;

[0075] The step S303 specifically includes:

[0076] Step S401 , obtaining first virtual target image feature vector transformation information and first scene video frame feature vector transformation information according to the virtual target image feature extraction vector information, the scene video frame feature extraction vector information and a preset first image information feature vector processing function.

[0077] In this embodiment, the preset first image information feature vector processing function can be artificially set, can be a linear transformation function, or can be a first image information feature vector processing function that performs matrix multiplication on the virtual target image feature extraction vector information and the scene video frame feature extraction vector information, and uses the result of the matrix multiplication as the first virtual target image feature vector transformation information and the first scene video frame feature vector transformation information.

[0078] Step S402: Obtain second virtual target image feature vector transformation information and second scene video frame feature vector transformation information according to the first virtual target image feature vector transformation information, the first scene video frame feature vector transformation information and a preset second image information feature vector processing function.

[0079] In this embodiment, the preset second image information feature vector processing function can be artificially set, and can be a nonlinear function, a hyperbolic tangent function, or a sigmoid function. The difference from the preset first image information feature vector processing function can be the difference in coefficients. The second image information feature vector processing function can be used to perform matrix multiplication on the first virtual target image feature vector transformation information and the first scene video frame feature vector transformation information, and the result of the matrix multiplication is used as the second virtual target image feature vector transformation information and the second scene video frame feature vector transformation information.

[0080] Step S403 , obtaining third virtual target image feature vector transformation information and third scene video frame feature vector transformation information according to the second virtual target image feature vector transformation information, the second scene video frame feature vector transformation information and a preset third image information feature vector processing function.

[0081] In this embodiment, the preset third image information feature vector processing function can be manually set and can be a mean aggregation function, that is, a function for automatically dividing all independent variables into batches, automatically calculating the mean of each batch, and outputting multiple mean calculation results. Alternatively, the function can be configured to use the second virtual target image feature vector transformation information and the second scene video frame feature vector transformation information as independent variables of the preset third image information feature vector processing function, perform dimensionality reduction processing on the second virtual target image feature vector transformation information and the second scene video frame feature vector transformation information through calculation by the third image information feature vector processing function, and output third virtual target image feature vector transformation information and third scene video frame feature vector transformation information of fixed dimensions.

[0082] Step S404 : calculating an average value of the third virtual target image feature vector transformation information and the third scene video frame feature vector transformation information to obtain virtual target image feature vector representation information and scene video frame feature vector representation information.

[0083] In this embodiment, the average value of each element of the third virtual target image feature vector transformation information and the third scene video frame feature vector transformation information can be calculated to enhance feature stability, and the obtained element-level average value can be used as the virtual target image feature vector representation information and the scene video frame feature vector representation information.

[0084] Step S405 : obtaining virtual target image feature information and scene video frame feature information according to the virtual target image feature vector representation information and the scene video frame feature vector representation information.

[0085] In this embodiment, the virtual target image feature vector representation information and the scene video frame feature vector representation information may be format converted, and the vector form may be converted into the numerical form of the virtual target image feature information and the scene video frame feature information for subsequent registration processing.

[0086] The three-dimensional visualization navigation method based on multi-sensor information fusion provided in the embodiment of the present application performs preliminary extraction of virtual target image feature extraction vector information and scene video frame feature extraction vector information through a first image information feature vector processing function, compresses redundant dimensions while enhancing the correlation between features, introduces nonlinear transformation through a second image information feature vector processing function to suppress invalid feature information in the preliminary extracted features and amplifies key response feature information, so that the feature vector has nonlinear expression capability, and performs spatial dimension compression on the feature information after invalid information suppression processing through a third image information feature vector processing function, aggregates local features into a globally representative vector to reduce computational complexity, and then effectively smoothes single sample noise through feature vector average value calculation to generate stable virtual target image feature vector representation information and scene video frame feature vector representation information, thereby ensuring the dimensional consistency of the feature vector while improving feature discrimination, laying the foundation for the efficiency and accuracy of subsequent alignment.

[0087] Figure 5 The flowchart of the implementation of the three-dimensional visual navigation method based on multi-sensor information fusion provided in the fifth embodiment of the present application is shown. The difference between the fifth embodiment and the fourth embodiment is that the step S405 specifically includes:

[0088] Step S501: scaling the virtual target image feature vector representation information and the scene video frame feature vector representation information according to a preset feature vector representation information scaling coefficient to obtain virtual target image feature vector dimensional transformation information and scene video frame feature vector dimensional transformation information.

[0089] In this embodiment, the preset feature vector representation information scaling coefficient can be set manually, which can be 0.5. The feature vector representation information scaling coefficient can be used to scale the virtual target image feature vector representation information and the scene video frame feature vector representation information. The vectors after dimension scaling are used as the virtual target image feature vector dimension transformation information and the scene video frame feature vector dimension transformation information to reduce the computational complexity.

[0090] Step S502 , obtaining a virtual target feature representation variable and a scene feature representation variable according to the virtual target image feature vector dimensional transformation information, the scene video frame feature vector dimensional transformation information, the preset feature vector coefficient information and the preset feature vector compensation information.

[0091] In this embodiment, the preset eigenvector coefficient information and the preset eigenvector compensation information can both be manually set. This can be accomplished by first performing a weighted summation of the virtual target image eigenvector dimensional transformation information and the scene video frame eigenvector dimensional transformation information using the preset eigenvector coefficient information, and then adding the weighted summation result to the preset eigenvector compensation information to obtain the virtual target feature representation variable and the scene feature representation variable.

[0092] Step S503 , performing dimensionality restoration processing on the virtual target feature representation variable and the scene feature representation variable according to a preset feature vector representation information scaling coefficient, to obtain virtual target image feature vector dimensionality restoration information and scene video frame feature vector dimensionality restoration information.

[0093] In this embodiment, the virtual target feature representation variable and the scene feature representation variable can be dimensionally restored according to the inverse of the preset feature vector representation information scaling coefficient to restore the virtual target feature representation variable and the scene feature representation variable to the original feature dimension. After restoration, the virtual target image feature vector dimension restoration information and the scene video frame feature vector dimension restoration information are obtained.

[0094] Step S504, obtaining virtual target image feature information and scene video frame feature information according to the virtual target image feature vector dimensional restoration information, the scene video frame feature vector dimensional restoration information, the preset dimensional restoration feature vector coefficient information and the preset dimensional restoration feature vector compensation information.

[0095] In this embodiment, the preset dimensionality restoration feature vector coefficient information and the preset dimensionality restoration feature vector compensation information can both be manually set. This can be accomplished by first performing a weighted summation process on the virtual target image feature vector dimensionality restoration information and the scene video frame feature vector dimensionality restoration information based on the dimensionality restoration feature vector coefficient information, then adding the weighted summation result to the dimensionality restoration feature vector compensation information, and then using the added result as the virtual target image feature information and the scene video frame feature information.

[0096] The three-dimensional visualization navigation method based on multi-sensor information fusion provided in the embodiment of the present application realizes flexible dimensionality reduction of virtual target image feature vector representation information and scene video frame feature vector representation information through a preset feature vector representation information scaling coefficient, reduces the amount of calculation while retaining the semantics of key features, calibrates the feature distribution in a low-dimensional space by introducing preset feature vector coefficient information and compensation information to enhance the physical meaning and scene adaptability of the features, restores the original expression space of the feature vector through dimensionality restoration, and further optimizes the feature distribution by using the preset dimensionality restoration feature vector coefficient information and compensation information, thereby solving the problem of computational redundancy or semantic loss caused by dimensionality fixation in the prior art, so that the generated virtual target image feature information and scene video frame feature information can adapt to the computing power of different hardware platforms and maintain semantic consistency across scenes, thereby providing effective data support for subsequent precise alignment processing.

[0097] Figure 6 The flowchart of the implementation of the three-dimensional visual navigation method based on multi-sensor information fusion provided in the sixth embodiment of the present application is shown. The difference between the sixth embodiment and the first embodiment is that the step S104 specifically includes:

[0098] Step S601 : Based on a preset image feature aggregation window, the virtual target image feature information and the scene video frame feature information are aggregated to obtain a plurality of virtual target image feature aggregation information and a plurality of scene video frame feature aggregation information.

[0099] In this embodiment, the preset image feature aggregation window can be set manually, and can be a window with the same length and width, and the length and width can be the distance of three elements. It can be based on the preset image feature aggregation window to perform sliding window aggregation on the virtual target image feature information and the scene video frame feature information, calculate the mean or maximum value of the feature vector in the window, and obtain multiple local aggregation features.

[0100] Step S602 : calculating similarities between the plurality of virtual target image feature aggregation information and the plurality of scene video frame feature aggregation information to obtain local similarity information of the plurality of image feature aggregation information.

[0101] In this embodiment, the cosine similarity of the virtual target image feature aggregation information and the scene video frame feature aggregation information may be calculated, and the cosine similarity is used as the local similarity information of the image feature aggregation information.

[0102] Step S603, determine whether the local similarity information of the image feature aggregation information is greater than the preset image local feature similarity threshold information; if so, obtain virtual target image registration information and scene target image feature registration information based on the virtual target image feature aggregation information corresponding to the local similarity information of the image feature aggregation information and the scene video frame feature aggregation information; if not, skip the virtual target image feature aggregation information corresponding to the local similarity information of the image feature aggregation information and the scene video frame feature aggregation information.

[0103] In this embodiment, the preset image local feature similarity threshold can be manually set, such as 0.7. When the local similarity information of the image feature aggregation information is greater than the preset image local feature similarity threshold information, the virtual target image feature aggregation information and the scene video frame feature aggregation information corresponding to the local similarity information of the image feature aggregation information are regarded as a successfully registered feature pair, and the feature pair is used as the virtual target image registration information and the scene target image feature registration information; when the local similarity information of the image feature aggregation information is less than or equal to the preset image local feature similarity threshold information, the virtual target image feature aggregation information and the scene video frame feature aggregation information corresponding to the local similarity information of the image feature aggregation information do not need to be regarded as a successfully registered feature pair, and thus the virtual target image feature aggregation information and the scene video frame feature aggregation information corresponding to the local similarity information of the image feature aggregation information can be ignored.

[0104] The three-dimensional visualization navigation method based on multi-sensor information fusion provided in the embodiment of the present application divides the virtual target image feature information and the scene video frame feature information into multiple local areas, and then suppresses the interference of single noise points through aggregation processing based on the image feature aggregation window, enhances the spatial consistency of features, and filters low-confidence matching pairs through a local feature matching mechanism that calculates cosine similarity through preset image local feature similarity threshold information, effectively eliminating false matches caused by local texture similarity, thereby decomposing the global alignment into independent matching of multiple local areas, reducing the computational complexity while improving the reliability of matching, and improving the accuracy, robustness and spatial consistency of the alignment results.

[0105] Figure 7 The following is a flowchart of a three-dimensional visualization navigation method based on multi-sensor information fusion according to the seventh embodiment of the present application. The difference between the seventh embodiment and the first embodiment is that:

[0106] The plurality of sensor measurement information includes acceleration measurement information and angular velocity measurement information;

[0107] The step S105 specifically includes:

[0108] Step S701 : extracting multi-dimensional coordinate information of the virtual target image registration information to obtain virtual target image registration coordinate range information.

[0109] In this embodiment, the three-dimensional coordinate range may be extracted from the virtual target image registration information as the virtual target image registration coordinate range information.

[0110] Step S702 : extracting pixel coordinate information of the scene object image feature registration information to obtain a plurality of scene object image feature registration coordinate information and scene object image feature registration coordinate range information.

[0111] In this embodiment, the two-dimensional pixel coordinates and coordinate ranges of all feature points may be extracted from the scene object image feature registration information as the scene object image feature registration coordinate information and range information.

[0112] Step S703 , scaling the virtual target image registration coordinate range information according to the scene target image feature registration coordinate range information to generate virtual target image registration scaling information; the virtual target image registration scaling information includes a plurality of virtual target image registration scaling coordinate information.

[0113] In this embodiment, the virtual target image registration coordinate range may be scaled proportionally according to the aspect ratio of the scene target image feature registration coordinate range to generate scaled virtual target image registration scaling coordinate information.

[0114] Step S704: Generate image feature angle change information and image feature translation change information according to the acceleration measurement information and the angular velocity measurement information.

[0115] In this embodiment, the acceleration measurement information may be integrated to obtain the image feature translation change information, and the image feature angle change information may be calculated by quaternion integration in combination with the angular velocity measurement information, thereby obtaining the image feature translation change amount and angle change amount.

[0116] Step S705 : performing coordinate transformation processing on the plurality of virtual target image registration and scaling coordinate information according to the image feature angle change information and the image feature translation change information to generate virtual target image transformation information.

[0117] In this embodiment, the scaled virtual target coordinates may be subjected to translation and rotation transformation to generate virtual target image transformation information, thereby achieving position compensation in a dynamic scene.

[0118] The three-dimensional visualization navigation method based on multi-sensor information fusion provided in the embodiment of the present application uses the scale scaling of the virtual target image registration coordinate range information and the scene target image feature registration coordinate range information to solve the problem of proportional imbalance between the virtual target and the real object. The motion compensation mechanism based on acceleration measurement information and angular velocity measurement information generates image feature angle change information and translation change information through integral operation to update the virtual target image registration scaling coordinate information in real time, effectively compensate for the registration drift caused by device movement, thereby combining the sensor's motion perception capability with the geometric transformation of the virtual target, realizing the maintenance of virtual and real space consistency in dynamic scenes, and significantly improving the accuracy and stability of navigation.

[0119] Figure 8 The flowchart of the implementation of the three-dimensional visual navigation method based on multi-sensor information fusion provided in the eighth embodiment of the present application is shown. The difference between the eighth embodiment and the first embodiment is that the step S106 specifically includes:

[0120] Step S801 : performing format conversion and splicing processing on the virtual target image transformation information and the scene target image feature registration information to generate image registration joint feature vector information.

[0121] In this embodiment, the virtual target image transformation information and the scene target image feature registration information are converted into one-dimensional feature vectors and then spliced ​​to generate joint feature vector information to fuse virtual and real features.

[0122] Step S802 : encoding and dimension transformation processing are performed on the plurality of sensor measurement information according to the dimension information of the image registration joint feature vector information to obtain sensor measurement feature vector information.

[0123] In this embodiment, the sensor measurement information may be encoded into a vector of the same dimension as the joint feature vector, or the sensor measurement information may be used as an independent variable of a ReLU function, and then dimension transformation and encoding conversion processing may be performed through calculation of the ReLU function to obtain the sensor measurement feature vector information.

[0124] Step S803, obtaining scene target tracking information based on the image registration joint feature vector information, the sensor measurement feature vector information and the preset feature vector tracking calculation function, so that the virtual target image transformation information is navigated based on the scene target tracking information to generate virtual target visualization information.

[0125] In this embodiment, the preset feature vector tracking calculation function can be manually set, can be designed based on a combination of weight coefficients and bias terms, or can be a Transformer encoder for capturing the temporal dependencies in the sequence data of the image registration joint feature vector information and the sensor measurement feature vector information. The image registration joint feature vector information and the sensor measurement feature vector information can be used as input information of the feature vector tracking calculation function, and the scene target tracking information is obtained by outputting the target position prediction value and combining it with the initial coordinates, thereby driving the virtual target to update its posture.

[0126] The three-dimensional visualization navigation method based on multi-sensor information fusion provided in the embodiment of the present application converts the format of virtual target image transformation information and scene target image feature registration information into different formats and splices them together to construct joint feature vector information containing virtual and real dual-domain features, so as to realize cross-domain feature expression of the target. Based on the encoding and dimensional transformation of sensor measurement information, different sensor measurement information and visual features interact in the same space, which makes up for the shortcomings of pure visual tracking in occluded or low-texture scenes. The long-term dependency between temporal features is captured through a preset feature vector tracking calculation function, and the recognition of complex motion trajectories is realized, thereby deeply integrating the spatial positioning capability of visual features with the motion prediction capability of sensors. The generated scene target tracking information can not only reflect the real-time position of the target, but also predict its future motion trajectory, ensuring that the dynamic navigation of the virtual target image transformation information is accurate, real-time and stable.

[0127] Corresponding to the method of the above embodiment, Figure 9 A structural block diagram of a three-dimensional visual navigation system based on multi-sensor information fusion provided in an embodiment of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown. Figure 9 The exemplary three-dimensional visualization navigation system based on multi-sensor information fusion may be an execution subject of the three-dimensional visualization navigation method based on multi-sensor information fusion provided in the aforementioned first embodiment.

[0128] Reference Figure 9 , the 3D visualization navigation system based on multi-sensor information fusion includes:

[0129] An information acquisition module 910 is configured to acquire multiple scene video frame information and multiple sensor measurement information;

[0130] A virtual target image information generating module 920 is configured to generate at least one virtual target image information;

[0131] An image feature information generating module 930 is configured to perform feature extraction and analysis processing on the virtual target image information and the scene video frame information based on a plurality of preset image information feature processing functions to obtain virtual target image feature information and scene video frame feature information;

[0132] An image registration information generating module 940 is configured to perform registration processing based on the virtual target image feature information and the scene video frame feature information to obtain virtual target image registration information and scene target image feature registration information;

[0133] A virtual target image transformation information generating module 950 is configured to perform image fusion processing on the virtual target image registration information and the scene target image feature registration information based on the plurality of sensor measurement information to generate virtual target image transformation information;

[0134] The virtual target visualization information generation module 960 is used to perform visual tracking calculations on the scene target image feature registration information based on the multiple sensor measurement information and the virtual target image transformation information to obtain scene target tracking information, so that the virtual target image transformation information can be navigated based on the scene target tracking information to generate virtual target visualization information.

[0135] The process of each module realizing its own function in the 3D visualization navigation system based on multi-sensor information fusion provided in the embodiment of the present application can be specifically referred to the aforementioned Figure 1 The description of the first embodiment will not be repeated here.

[0136] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0137] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.

[0138] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0139] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0140] In addition, in the description of the present specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish descriptions and should not be understood as indicating or implying relative importance. It should also be understood that although the terms "first", "second", etc. are used in the text to describe various elements in some embodiments of the present application, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first table can be named a second table, and similarly, a second table can be named a first table without departing from the scope of the various described embodiments. Both the first table and the second table are tables, but they are not the same table.

[0141] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0142] The three-dimensional visualization navigation method based on multi-sensor information fusion provided in the embodiments of the present application can be applied to terminal devices such as tablet computers, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs). The embodiments of the present application do not impose any restrictions on the specific types of terminal devices.

[0143] For example, the terminal device can be a station (STAION, ST) in a WLAN, a cellular phone, a cordless phone, a Session Initiation Protocol (SIP) phone, a Wireless Local Loop (WLL) station, a Personal Digital Assistant (PDA) device, a handheld device with wireless communication function, a computing device or other processing device connected to a wireless modem, a vehicle-mounted device, a vehicle networking terminal, a computer, a laptop computer, a handheld communication device, a handheld computing device, a satellite wireless device, a wireless modem card, a TV set-top box (STB), customer premise equipment (CPE) and / or other devices for communicating on a wireless system and a next-generation communication system, such as a mobile terminal in a 5G network or a mobile terminal in a future evolved Public Land Mobile Network (PLMN) network.

[0144] As an example and not a limitation, when the terminal device is a wearable device, the wearable device can also be a general term for wearable devices that are intelligently designed and developed using wearable technology for daily wear, such as glasses, gloves, watches, clothing and shoes. A wearable device is a portable device that is worn directly on the body or integrated into the user's clothes or accessories. Wearable devices are not only hardware devices, but also achieve powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable smart devices include those that are full-featured, large in size, and can achieve complete or partial functions without relying on smartphones, such as smart watches or smart glasses, as well as those that only focus on a certain type of application function and need to be used in conjunction with other devices such as smartphones, such as various smart bracelets and smart jewelry for vital sign monitoring.

[0145] Figure 10 This is a schematic diagram of the structure of a terminal device provided by an embodiment of the present application. Figure 10 As shown, the terminal device 100 of this embodiment includes: at least one processor 1000 ( Figure 10 Only one is shown), a memory 1001, wherein the memory 1001 stores a computer program 1002 that can be run on the processor 1000. When the processor 1000 executes the computer program 1002, the steps in the above-mentioned embodiments of the three-dimensional visualization navigation method based on multi-sensor information fusion are implemented, such as Figure 1Alternatively, when the processor 1000 executes the computer program 1002, the functions of the modules / units in the above-mentioned system embodiments are realized, for example, Figure 9 Functions of modules 910 to 960 are shown.

[0146] The terminal device 100 may be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The terminal device may include, but is not limited to, a processor 1000 and a memory 1001. Those skilled in the art will understand that Figure 10 It is only an example of the terminal device 100 and does not constitute a limitation of the terminal device 100. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the terminal device may also include an input and sending device, a network access device, a bus, etc.

[0147] The processor 1000 may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0148] In some embodiments, the memory 1001 may be an internal storage unit of the terminal device 100, such as a hard disk or memory of the terminal device 100. The memory 1001 may also be an external storage device of the terminal device 100, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the terminal device 100. Furthermore, the memory 1001 may include both an internal storage unit of the terminal device 100 and an external storage device. The memory 1001 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program. The memory 1001 may also be used to temporarily store data that has been sent or is to be sent.

[0149] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0150] An embodiment of the present application also provides a terminal device, which includes at least one memory, at least one processor, and a computer program stored in the at least one memory and executable on the at least one processor. When the processor executes the computer program, the terminal device implements the steps of any of the above-mentioned method embodiments.

[0151] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application can implement all or part of the processes in the above-mentioned embodiment method by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium.

[0152] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0153] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0154] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0155] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A three-dimensional visualization navigation method based on multi-sensor information fusion, characterized in that: include: Obtain multiple scene video frame information and multiple sensor measurement information; generating at least one virtual target image information; Based on a plurality of preset image information feature processing functions, feature extraction and analysis processing are performed on the virtual target image information and the scene video frame information to obtain virtual target image feature information and scene video frame feature information; Performing registration processing based on the virtual target image feature information and the scene video frame feature information to obtain virtual target image registration information and scene target image feature registration information; performing image fusion processing on the virtual target image registration information and the scene target image feature registration information according to the plurality of sensor measurement information to generate virtual target image transformation information; According to the plurality of sensor measurement information and the virtual target image transformation information, the scene target image feature registration information is visually tracked and calculated to obtain scene target tracking information, so that the virtual target image transformation information is navigated based on the scene target tracking information to generate virtual target visualization information.

2. The three-dimensional visualization navigation method based on multi-sensor information fusion according to claim 1, characterized in that: The plurality of preset image information feature processing functions include a preset first-scale image information feature processing function, a preset second-scale image information feature processing function, and a plurality of preset image information feature vector processing functions; The step of performing feature extraction and analysis processing on the virtual target image information and the scene video frame information based on a plurality of preset image information feature processing functions to obtain the virtual target image feature information and the scene video frame feature information specifically includes: Performing format conversion on the virtual target image information and the scene video frame information to obtain a virtual target pixel matrix and a scene video frame pixel matrix; performing analytical calculations on the virtual target pixel matrix and the scene video frame pixel matrix according to a preset first-scale image information feature processing function and a preset second-scale image information feature processing function to obtain first-scale virtual target pixel matrix variable information, second-scale virtual target pixel matrix variable information, first-scale scene video frame pixel matrix variable information, and second-scale scene video frame pixel matrix variable information; performing a difference calculation on the first-scale virtual target pixel matrix variable information and the second-scale virtual target pixel matrix variable information to obtain scale-difference virtual target pixel matrix variable information; Performing difference calculation on the first-scale scene video frame pixel matrix variable information and the second-scale scene video frame pixel matrix variable information to obtain scale difference scene video frame pixel matrix variable information; Extracting extreme values ​​of the scale-difference virtual target pixel matrix variable information and extreme values ​​of the scale-difference scene video frame pixel matrix variable information to obtain virtual target image feature center information and scene video frame feature center information; The virtual target image feature information and the scene video frame feature information are obtained according to the virtual target image feature center information, the scene video frame feature center information, the preset image information feature extraction radius information and a plurality of preset image information feature vector processing functions.

3. The three-dimensional visualization navigation method based on multi-sensor information fusion according to claim 2, characterized in that: The step of obtaining the virtual target image feature information and the scene video frame feature information according to the virtual target image feature center information, the scene video frame feature center information, the preset image information feature extraction radius information, and a plurality of preset image information feature vector processing functions specifically includes: Based on the virtual target image feature center information and the scene video frame feature center information, and according to preset image information feature extraction radius information, a plurality of virtual target image feature extraction point information and a plurality of scene video frame feature extraction point information are obtained; Generate virtual target image feature extraction vector information and scene video frame feature extraction vector information according to the plurality of virtual target image feature extraction point information and the plurality of scene video frame feature extraction point information; The virtual target image feature information and the scene video frame feature information are obtained according to the virtual target image feature extraction vector information, the scene video frame feature extraction vector information and a plurality of preset image information feature vector processing functions.

4. The three-dimensional visualization navigation method based on multi-sensor information fusion according to claim 3, characterized in that: The plurality of preset image information feature vector processing functions include a preset first image information feature vector processing function, a preset second image information feature vector processing function, and a preset third image information feature vector processing function; The step of obtaining virtual target image feature information and scene video frame feature information based on the virtual target image feature extraction vector information, the scene video frame feature extraction vector information, and a plurality of preset image information feature vector processing functions specifically includes: Obtaining first virtual target image feature vector transformation information and first scene video frame feature vector transformation information according to the virtual target image feature extraction vector information, the scene video frame feature extraction vector information, and a preset first image information feature vector processing function; Obtaining second virtual target image feature vector transformation information and second scene video frame feature vector transformation information according to the first virtual target image feature vector transformation information, the first scene video frame feature vector transformation information, and a preset second image information feature vector processing function; Obtaining third virtual target image feature vector transformation information and third scene video frame feature vector transformation information according to the second virtual target image feature vector transformation information, the second scene video frame feature vector transformation information, and a preset third image information feature vector processing function; Calculating an average value of the third virtual target image feature vector transformation information and the third scene video frame feature vector transformation information to obtain virtual target image feature vector representation information and scene video frame feature vector representation information; According to the virtual target image feature vector representation information and the scene video frame feature vector representation information, virtual target image feature information and scene video frame feature information are obtained.

5. The three-dimensional visualization navigation method based on multi-sensor information fusion according to claim 4, characterized in that: The step of obtaining virtual target image feature information and scene video frame feature information based on the virtual target image feature vector representation information and the scene video frame feature vector representation information specifically includes: Scaling the virtual target image feature vector representation information and the scene video frame feature vector representation information according to a preset feature vector representation information scaling factor to obtain virtual target image feature vector dimensional transformation information and scene video frame feature vector dimensional transformation information; Obtaining a virtual target feature representation variable and a scene feature representation variable according to the virtual target image feature vector dimensional transformation information, the scene video frame feature vector dimensional transformation information, the preset feature vector coefficient information, and the preset feature vector compensation information; Performing dimensionality reduction processing on the virtual target feature representation variable and the scene feature representation variable according to a preset feature vector representation information scaling coefficient to obtain virtual target image feature vector dimensionality reduction information and scene video frame feature vector dimensionality reduction information; The virtual target image feature information and the scene video frame feature information are obtained according to the virtual target image feature vector dimensional restoration information, the scene video frame feature vector dimensional restoration information, the preset dimensional restoration feature vector coefficient information and the preset dimensional restoration feature vector compensation information.

6. The three-dimensional visualization navigation method based on multi-sensor information fusion according to claim 1, characterized in that: The step of performing registration processing according to the virtual target image feature information and the scene video frame feature information to obtain virtual target image registration information and scene target image feature registration information specifically includes: Based on a preset image feature aggregation window, the virtual target image feature information and the scene video frame feature information are aggregated to obtain a plurality of virtual target image feature aggregation information and a plurality of scene video frame feature aggregation information; Calculating similarities of the plurality of virtual target image feature aggregation information and the plurality of scene video frame feature aggregation information to obtain local similarity information of the plurality of image feature aggregation information; When the local similarity information of the image feature aggregation information is greater than the preset image local feature similarity threshold information, the virtual target image registration information and the scene target image feature registration information are obtained according to the virtual target image feature aggregation information corresponding to the local similarity information of the image feature aggregation information and the scene video frame feature aggregation information.

7. The three-dimensional visualization navigation method based on multi-sensor information fusion according to claim 1, characterized in that: The plurality of sensor measurement information includes acceleration measurement information and angular velocity measurement information; The step of performing image fusion processing on the virtual target image registration information and the scene target image feature registration information based on the plurality of sensor measurement information to generate virtual target image transformation information specifically includes: Extracting multi-dimensional coordinate information of the virtual target image registration information to obtain virtual target image registration coordinate range information; Extracting pixel coordinate information of the scene target image feature registration information to obtain a plurality of scene target image feature registration coordinate information and scene target image feature registration coordinate range information; According to the scene target image feature registration coordinate range information, the virtual target image registration coordinate range information is scaled to generate virtual target image registration scaling information; the virtual target image registration scaling information includes a plurality of virtual target image registration scaling coordinate information; generating image feature angle change information and image feature translation change information based on the acceleration measurement information and the angular velocity measurement information; According to the image feature angle change information and the image feature translation change information, coordinate transformation processing is performed on the plurality of virtual target image registration and scaling coordinate information to generate virtual target image transformation information.

8. The three-dimensional visualization navigation method based on multi-sensor information fusion according to claim 1, characterized in that: The step of performing visual tracking calculation on the scene target image feature registration information based on the plurality of sensor measurement information and the virtual target image transformation information to obtain scene target tracking information, so that the virtual target image transformation information is navigated based on the scene target tracking information to generate virtual target visualization information, specifically includes: Performing format conversion and splicing processing on the virtual target image transformation information and the scene target image feature registration information to generate image registration joint feature vector information; encoding and dimensionally transforming the plurality of sensor measurement information according to the dimensional information of the image registration joint feature vector information to obtain sensor measurement feature vector information; According to the image registration joint feature vector information, the sensor measurement feature vector information and the preset feature vector tracking calculation function, the scene target tracking information is obtained, so that the virtual target image transformation information is navigated based on the scene target tracking information to generate virtual target visualization information.

9. A three-dimensional visual navigation system based on multi-sensor information fusion, characterized in that: include: An information acquisition module, used to acquire multiple scene video frame information and multiple sensor measurement information; A virtual target image information generating module, configured to generate at least one virtual target image information; An image feature information generation module is used to perform feature extraction and analysis processing on the virtual target image information and the scene video frame information based on a plurality of preset image information feature processing functions to obtain virtual target image feature information and scene video frame feature information; An image registration information generating module is used to perform registration processing based on the virtual target image feature information and the scene video frame feature information to obtain virtual target image registration information and scene target image feature registration information; a virtual target image transformation information generating module, configured to perform image fusion processing on the virtual target image registration information and the scene target image feature registration information based on the plurality of sensor measurement information, and generate virtual target image transformation information; A virtual target visualization information generation module is used to perform visual tracking calculation on the scene target image feature registration information based on the multiple sensor measurement information and virtual target image transformation information to obtain scene target tracking information, so that the virtual target image transformation information can be navigated based on the scene target tracking information to generate virtual target visualization information.

10. A terminal device, characterized in that: The terminal device includes a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Three-dimensional visual navigation method based on multi-sensor information fusion

    CN101833104A

  • Safe airplane approach method based on multisensor information fusion

    CN101923789A

  • Target positioning and tracking system and method based on video and three-dimensional spatial information registration fusion

    CN106204656A

  • Image navigation and positioning system and image navigation and positioning method

    CN106580471A

  • Multi-target tracking method combined with video scene feature perception

    CN110660083A

Cited By

  • Video picture and virtual information projection registration and fusion method and system based on digital twinborn scene

    CN121458859A