Monocular vision instant positioning method based on digital twin data and semantic information

Through a monocular visual instant positioning method based on digital twin data and semantic information, the problem of real-time positioning drift and resource-limited equipment processing large batches of image data in traditional positioning technology is solved, and high-precision positioning results are achieved.

CN119991788AActive Publication Date: 2025-05-13TONGJI UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411801551.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-09
Publication Date
2025-05-13
Estimated Expiration
2044-12-09

AI Technical Summary

Technical Problem

Traditional visual information positioning technology relies on dense frame image data, and there is a situation where real-time positioning drifts with the cumulative error of motion estimation. In addition, personal mobile devices and small drones have limited computing and storage resources, making it difficult to process large-scale image data.

Method used

A monocular visual instant positioning method based on digital twin data and semantic information is adopted. By constructing a training data set, a positioning model is constructed and adversarial training is carried out. The image feature extraction module, a semantic feature extraction module, a similarity scoring table generation module, a residual connection module, a decoding module and a positioning module are used to generate accurate positioning results.

Benefits of technology

High-precision positioning on equipment with limited resources is achieved, drift problems caused by cumulative errors are avoided, and precise positioning results are generated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991788A_ABST
    Figure CN119991788A_ABST
Patent Text Reader

Abstract

The invention provides a monocular vision instant positioning method based on digital twin data and semantic information, and the method is characterized in that the method comprises the steps: S1, according to real visual data collected in a preset region, virtual visual data collected by a virtual camera in a digital twin model corresponding to the preset region, and a corresponding label, obtaining real visual data; constructing a training data set; s2, constructing a positioning model and a discriminator, and performing adversarial training on the positioning model according to the training data set and the discriminator to obtain a trained positioning model; and S3, inputting the image data into the trained positioning model to obtain a positioning result. In a word, the method can generate an accurate positioning result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of visual positioning, and specifically relates to a monocular visual instant positioning method based on digital twin data and semantic information. Background Art

[0002] Positioning technology refers to the technology or service that obtains the location information of the carrier through a specific method and marks it. The Global Positioning System, or GPS, is the most common positioning technology basis, but in some scenarios, such as bridges, culverts, tunnels, complex indoor environments, underground spaces and other locations where GPS signals are poor, related technologies are difficult to provide continuous, accurate and timely positioning results.

[0003] Simultaneous localization and mapping (SLAM) and other visual information-based real-time positioning technologies provide an effective alternative. SLAM and other technologies make use of the wide availability of cameras on smart devices. By estimating the position of the camera in three-dimensional space, they can achieve real-time and accurate positioning and navigation of mobile devices, which is applicable to virtual reality, robot navigation and other issues.

[0004] However, traditional visual information positioning technology often relies on dense frame monocular, binocular, and multi-camera photography, and solves the camera pose based on geometric constraints. There is a situation where real-time positioning drifts with the accumulated error of motion estimation. In addition, for personal mobile devices and small drones, the computing and storage resources are limited, and it is difficult to handle large batches of image data processing.

[0005] Therefore, there is an urgent need for a high-precision positioning method based on sparse image data. Summary of the invention

[0006] The present invention is made to solve the above-mentioned problems, and its purpose is to provide a monocular vision instant positioning method based on digital twin data and semantic information.

[0007] The present invention provides a monocular visual instant positioning method based on digital twin data and semantic information, which is used to obtain the positioning result of the mobile device according to the image data collected by the monocular camera device set on the mobile device, and has the following characteristics, including the following steps: step S1, according to the real visual data collected in the preset area, and the virtual visual data and corresponding labels collected by the virtual camera in the digital twin model corresponding to the preset area, a training data set is constructed; step S2, a positioning model and a discriminator are constructed, and adversarial training is performed on the positioning model according to the training data set and the discriminator to obtain a trained positioning model; step S3, the image data is input into the trained positioning model to obtain the positioning result, wherein the positioning model includes: an image feature extraction module , including an image encoder, which is used to extract features from image data to obtain multiple image features of different sizes, and perform self-attention calculation on specified image features to obtain representative features; a semantic feature extraction module, including a text encoder and a preset prompt word, which is used to extract features from the prompt word through the text encoder to obtain semantic features; a similarity score table generation module, which is used to calculate similarities based on semantic features and representative features to obtain a similarity score table; a residual connection module, which is used to extract global information from image data to obtain global features; a decoding module, which is used to obtain a decoded image based on global features, the similarity score table and all image features; a positioning module, which includes a regression head, which is used to generate posture information of a monocular camera device as a positioning result based on the decoded image.

[0008] In the monocular visual instant positioning method based on digital twin data and semantic information provided by the present invention, it can also have the following characteristics: wherein, the residual connection module includes: a patch embedding unit, which is used to adjust the data size of the image data according to the maximum size corresponding to the image feature; a large core attention unit, which is used to extract global information from the adjusted image data to obtain global features.

[0009] In the monocular visual instant positioning method based on digital twin data and semantic information provided by the present invention, it can also have the following characteristics: wherein, the large core attention unit includes a first batch of normalization layers, an attention layer, a first connection layer, a second batch of normalization layers, a feedforward neural network and a second connection layer connected in sequence, the input of the first batch of normalization layers is the adjusted image data, the input of the first connection layer is the output of the attention layer and the adjusted image data, the input of the second connection layer is the output of the second batch of normalization layers and the feedforward neural network, and the output is a global feature.

[0010] In the monocular visual instant positioning method based on digital twin data and semantic information provided by the present invention, it can also have the following characteristics: wherein, the attention layer includes a first convolution sublayer, a first GELU activation function sublayer, an MLKA sublayer and a second convolution sublayer connected in sequence.

[0011] In the monocular visual instant positioning method based on digital twin data and semantic information provided by the present invention, it can also have the following characteristics: wherein, the feedforward neural network includes a third convolution sublayer, a fourth convolution sublayer, a second GELU activation function sublayer and a fifth convolution sublayer connected in sequence.

[0012] In the monocular vision instant positioning method based on digital twin data and semantic information provided by the present invention, it can also have the following characteristics: wherein, in the decoding module, the global feature is superimposed with each image feature of size from large to small in turn, and each superimposed feature is convolved, the number of channels of the superimposed feature is compressed to a preset number of channels, and the superimposed feature with the preset number of channels is used as a new global feature and superimposed with the next feature image. Before superposition, all image features are compressed to the preset number of channels through convolution processing, and the superimposed feature with the preset number of channels corresponding to the last image feature is convolved with the similarity scoring table to obtain a decoded image.

[0013] In the monocular vision instant positioning method based on digital twin data and semantic information provided by the present invention, it can also have the following characteristics: wherein, the specified image feature is the image feature with the smallest size among all image features.

[0014] In the monocular visual instant positioning method based on digital twin data and semantic information provided by the present invention, it can also have the following characteristics: wherein, the training data set includes a real data subset constructed based on real visual data and a synthetic data subset constructed based on virtual visual data and corresponding labels. In step S2, the positioning model is adversarially trained based on the synthetic data subset, and tested based on the real data subset until the error on the real data subset is minimized, thereby obtaining a trained positioning model.

[0015] In the monocular visual instant positioning method based on digital twin data and semantic information provided by the present invention, it can also have the following characteristics: wherein the loss function of adversarial training includes regression loss and classification loss The calculation expression of regression loss is: Where x is the coordinate of the virtual camera in the pose information predicted by the positioning model, is the real coordinate of the virtual camera in the label, q is the azimuth of the virtual camera in the pose information predicted by the positioning model, is the true azimuth of the virtual camera in the label, γ is the hyperparameter error ratio, and the classification loss The calculation expression is: Where D is the discriminator, G(I t ) is the positioning model based on the input image data It The resulting decoded image.

[0016] Functions and Effects of the Invention

[0017] According to the monocular visual instant positioning method based on digital twin data and semantic information involved in the present invention, because, on the one hand, a large amount of training data is obtained through the digital twin model, and the positioning model is trained in combination with the training data of the real environment to improve the prediction performance of the positioning model; on the other hand, the residual connection module is used to enhance the expression ability of image features, and the decoder is used to fuse the outputs of the residual connection module, the image feature extraction module and the similarity scoring table generation module, and then the regression head is used to obtain accurate positioning results. Therefore, the monocular visual instant positioning method based on digital twin data and semantic information of the present invention can generate accurate positioning results. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a flow chart of a monocular vision instant positioning method based on digital twin data and semantic information in an embodiment of the present invention;

[0019] Figure 2 is a schematic diagram of a training data set in an embodiment of the present invention

[0020] Figure 3 is a block diagram of a positioning model in an embodiment of the present invention;

[0021] Figure 4 is a block diagram of a large core attention unit in an embodiment of the present invention. DETAILED DESCRIPTION

[0022] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the following embodiments and the accompanying drawings specifically illustrate the monocular vision real-time positioning method based on digital twin data and semantic information of the present invention.

[0023] In this embodiment, a monocular visual instant positioning method based on digital twin data and semantic information is provided, which is used to obtain a positioning result of the mobile device based on image data collected by a monocular camera device arranged on the mobile device.

[0024] Figure 1 It is a flow chart of a monocular vision instant positioning method based on digital twin data and semantic information in an embodiment of the present invention.

[0025] like Figure 1 As shown in FIG. 1 , the monocular vision instant positioning method based on digital twin data and semantic information includes the following steps:

[0026] Step S1, constructing a training data set based on real visual data collected in a preset area, and virtual visual data and corresponding labels collected by a virtual camera in a digital twin model corresponding to the preset area.

[0027] Figure 2 Schematic diagram of a training data set in an embodiment of the present invention.

[0028] like Figure 2 As shown in the figure, for the real preset area, the real visual data is collected by the camera device to construct the real data subset. A digital twin model is built for the preset area, and the virtual visual data and the corresponding posture parameters are collected by the virtual camera with known posture parameters as labels to construct the synthetic data subset. Then, the final training data set includes the real data subset and the synthetic data subset.

[0029] In this embodiment, the real visual data can be fixed or non-fixed interval image sampling or video stream, where the sampling frequency is unlimited and multiple images do not need to contain the same observed object. In this embodiment, the virtual visual data is fixed or non-fixed interval image sampling or video stream collected along a certain trajectory in the digital twin model.

[0030] Step S2, constructing a positioning model and a discriminator, and performing adversarial training on the positioning model according to the training data set and the discriminator to obtain a trained positioning model.

[0031] Figure 3 is a block diagram of a positioning model in an embodiment of the present invention.

[0032] like Figure 3 As shown, the positioning model 100 includes an image feature extraction module 11, a semantic feature extraction module 12, a similarity scoring table generation module 13, a residual connection module 14, a decoding module 15 and a positioning module 16.

[0033] The image feature extraction module 11 includes an image encoder, which is used to extract features from image data to obtain multiple image features of different sizes, and perform self-attention calculation on designated image features to obtain representative features.

[0034] Among them, the designated image feature is the image feature with the smallest size among all image features. In this embodiment, four image features of different sizes are obtained respectively, and they are sorted from large to small in size as image feature A1, image feature A2, image feature A3 and image feature A4, and the designated image feature is image feature A4. In other embodiments, the existing image encoder can be used without modification, and for the multiple image features outputted, an image feature of a specific size rather than the smallest size can be selected as the designated image feature.

[0035] The calculation expression representing the characteristics in this embodiment is:

[0036]

[0037] In the formula is the representative feature, MHSA is the self-attention calculation, x4 is the specified image feature, is the mean of the specified image features,

[0038] The semantic feature extraction module 12 includes a text encoder and a preset prompt word, and is used to extract features from the prompt word through the text encoder to obtain semantic features.

[0039] The text encoder of the semantic feature extraction module 12 and the image encoder of the image feature extraction module 11 of this embodiment are respectively the text encoder and the image encoder of the existing contrastive language-image pre-training model, namely the CLIP model. In other embodiments, the image encoder and the text encoder of other multimodal models can be used as the text encoder of the semantic feature extraction module 12 and the image encoder of the image feature extraction module 11.

[0040] The preset prompt words of this embodiment are obtained in sequence through the prompt word preselection step and the prompt word setting step. The prompt word preselection step is to input the training data in the synthetic data subset into the positioning model 100, and at the same time input the list of candidate prompt words containing multiple candidate prompt words, and then after the positioning model 100 converges according to the training data, the candidate prompt words are preselected according to the corresponding weights of the candidate prompt word list, and multiple preselected prompt words are obtained as the preselected prompt word set. The prompt word setting step is to manually select appropriate prompt words in the preselected prompt word set as the preset prompt words, or directly select appropriate preselected prompt words as the preset prompt words according to the characteristics of the digital twin model.

[0041] The similarity score table generating module 13 is used to calculate the similarity according to the semantic features and the representative features to obtain a similarity score table.

[0042] The residual connection module 14 is used to extract global information from the image data to obtain global features.

[0043] The residual connection module 14 includes a patch embedding unit 141 and a large core attention unit 142.

[0044] The patch embedding unit 141 is used to adjust the data size of the image data according to the maximum size corresponding to the image feature.

[0045] The large core attention unit 142 is used to extract global information from the adjusted image data to obtain global features.

[0046] Figure 4 is a block diagram of a large core attention unit in an embodiment of the present invention.

[0047] like Figure 4 As shown, the large core attention unit 142 includes a first batch of normalization layers 1421, an attention layer 1422, a first connection layer 1423, a second batch of normalization layers 1424, a feedforward neural network 1425 and a second connection layer 1426 connected in sequence.

[0048] The input of the first batch normalization layer 1421 is the adjusted image data. In this embodiment, the first batch normalization layer 1421 and the second batch normalization layer 1424 perform batch normalization on the input so that the mean is 0 and the variance is 1.

[0049] The attention layer 1422 includes a first convolution sublayer 14221, a first GELU activation function sublayer 14222, an MLKA sublayer 14223, and a second convolution sublayer 14224 which are connected sequentially.

[0050] The input of the first connection layer 1423 is the output of the attention layer 1422 and the adjusted image data.

[0051] The feedforward neural network 1425 includes a third convolution sublayer 14251, a fourth convolution sublayer 14252, a second GELU activation function sublayer 14253 and a fifth convolution sublayer 14254 which are connected in sequence.

[0052] The number of channels of the fourth convolution sublayer 14252 is a preset number of channels, and in this embodiment, the preset number of channels is 256.

[0053] The input of the second connection layer 1426 is the output of the second batch normalization layer 1424 and the feedforward neural network 1425, and the output is the global feature.

[0054] The decoding module 15 includes a plurality of convolutional layers, which are used to obtain a decoded image according to the global features, the similarity score table and all image features.

[0055] Among them, in the decoding module 15, the global feature is sequentially superimposed with each image feature of size from large to small, and each superimposed feature is convoluted to compress the number of channels of the superimposed feature to a preset number of channels. The superimposed feature with the preset number of channels is used as a new global feature to be superimposed with the next feature image. Before superimposition, all image features are convoluted to compress the corresponding number of channels to the preset number of channels. The superimposed feature with the preset number of channels corresponding to the last image feature is convoluted with the similarity score table to obtain a decoded image.

[0056] For example, in this embodiment, the image feature A1, the image feature A2, the image feature A3, and the image feature A4 are all compressed to a preset number of channels by 1x1 convolution. Subsequently, the image feature A4 and the global feature are sequentially superimposed and compressed to a preset number of channels by convolution to obtain a new global feature. The global feature is then sequentially superimposed and compressed with the image feature A3, the image feature A2, and the image feature A1. Finally, the obtained global feature is convolved with the similarity score table to obtain a decoded image.

[0057] The positioning module 16 includes a regression head for generating the position and posture information of the monocular camera device as a positioning result according to the decoded image.

[0058] In this embodiment, the positioning model 100 is adversarially trained based on a synthetic data subset and tested based on a real data subset until the error on the real data subset is minimized, thereby obtaining a trained positioning model. In this embodiment, the error corresponding to the synthetic data subset is less than a preset threshold, and adversarial training is performed by the discriminator to minimize the domain difference error, and the error on the real data subset is minimized at the same time. In this embodiment, the parameters of the positioning model 100 and the discriminator are cyclically optimized in sequence through adversarial training, wherein the parameters of the residual connection module 14 and the decoding module 15 of the positioning model 100 are optimized.

[0059] The loss function of adversarial training includes regression loss and classification loss

[0060] The calculation expression of regression loss is:

[0061]

[0062]

[0063] Where x is the coordinate of the virtual camera in the pose information predicted by the positioning model, is the real coordinate of the virtual camera in the label, q is the azimuth of the virtual camera in the pose information predicted by the positioning model, is the true azimuth of the virtual camera in the label, and γ is the hyperparameter error ratio.

[0064] Classification Loss The calculation expression is:

[0065]

[0066] Where D is the discriminator, G(I t ) is the positioning model based on the input image data I t The resulting decoded image.

[0067] Step S3, input the image data into the trained positioning model 100 to obtain the positioning result. In this embodiment, the image data input into the trained positioning model 100 is a single picture. In other embodiments, the input of the model can be converted into multiple pictures by adding a corresponding module for time series processing to the positioning model 100 and setting a corresponding training loss.

[0068] Functions and Effects of the Embodiments

[0069] According to the monocular visual instant positioning method based on digital twin data and semantic information involved in this embodiment, on the one hand, a large amount of training data is obtained through the digital twin model, and the positioning model is trained in combination with the training data of the real environment to improve the prediction performance of the positioning model; on the other hand, the residual connection module is used to enhance the expression ability of image features, and the decoder is used to fuse the outputs of the residual connection module, the image feature extraction module and the similarity score table generation module, and then the regression head is used to obtain accurate positioning results. In short, this method does not rely on the results of the previous estimation, and no cumulative drift will occur during positioning, and it can generate accurate positioning results.

[0070] Those skilled in the art should understand that the present invention is not limited to the above embodiments, and the above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, and these changes and improvements fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.

Claims

1. A monocular visual instant positioning method based on digital twin data and semantic information, which is used to obtain the positioning result of the mobile device according to the image data collected by the monocular camera device set on the mobile device, characterized in that: The following steps are involved: Step S1, constructing a training data set based on real visual data collected in a preset area, and virtual visual data and corresponding labels collected by a virtual camera in a digital twin model corresponding to the preset area; Step S2, constructing a positioning model and a discriminator, and performing adversarial training on the positioning model according to the training data set and the discriminator to obtain a trained positioning model; Step S3, inputting the image data into the trained positioning model to obtain the positioning result, Wherein, the positioning model includes: An image feature extraction module, including an image encoder, is used to extract features from the image data to obtain multiple image features of different sizes, and to perform self-attention calculation on the specified image features to obtain representative features; A semantic feature extraction module, comprising a text encoder and a preset prompt word, for extracting features from the prompt word through the text encoder to obtain semantic features; A similarity score table generating module, used for calculating similarity according to the semantic features and the representative features to obtain a similarity score table; A residual connection module, used for extracting global information from the image data to obtain global features; A decoding module, used for obtaining a decoded image according to the global feature, the similarity scoring table and all the image features; The positioning module includes a regression head, which is used to generate the position information of the monocular camera device as the positioning result according to the decoded image.

2. The monocular vision instant positioning method based on digital twin data and semantic information according to claim 1, Features: Wherein, the residual connection module comprises: a patch embedding unit, configured to adjust the data size of the image data according to the maximum size corresponding to the image feature; The large core attention unit is used to extract global information from the adjusted image data to obtain the global features.

3. The monocular vision instant positioning method based on digital twin data and semantic information according to claim 2 is characterized in that: in, The large core attention unit includes a first batch of normalized layers, an attention layer, a first connection layer, a second batch of normalized layers, a feedforward neural network, and a second connection layer connected in sequence, The input of the first batch of normalization layers is the adjusted image data, The input of the first connection layer is the output of the attention layer and the adjusted image data, The input of the second connection layer is the output of the second batch normalization layer and the feedforward neural network, and the output is the global feature.

4. The monocular vision instant positioning method based on digital twin data and semantic information according to claim 3 is characterized in that: in, The attention layer includes a first convolution sublayer, a first GELU activation function sublayer, an MLKA sublayer and a second convolution sublayer which are connected in sequence.

5. The monocular vision instant positioning method based on digital twin data and semantic information according to claim 3 is characterized in that: in, The feedforward neural network includes a third convolution sublayer, a fourth convolution sublayer, a second GELU activation function sublayer and a fifth convolution sublayer which are connected in sequence.

6. The monocular vision instant positioning method based on digital twin data and semantic information according to claim 1 is characterized in that: in, In the decoding module, the global feature is sequentially superimposed with each of the image features of decreasing size, and convolution processing is performed on each superimposed feature to compress the number of channels of the superimposed feature to a preset number of channels. The superimposed feature with the preset number of channels is used as the new global feature and is superimposed with the next feature image. Before superposition, all the image features are compressed to the preset number of channels by convolution processing. The superimposed feature with the preset number of channels corresponding to the last image feature is convolved with the similarity scoring table to obtain the decoded image.

7. The monocular vision instant positioning method based on digital twin data and semantic information according to claim 1, characterized in that: in, The specified image feature is the image feature with the smallest size among all the image features.

8. The monocular vision instant positioning method based on digital twin data and semantic information according to claim 1, characterized in that: in, The training data set includes a real data subset constructed according to the real visual data and a synthetic data subset constructed according to the virtual visual data and corresponding labels, In step S2, the positioning model is adversarially trained based on the synthetic data subset and tested based on the real data subset until the error on the real data subset is minimized, thereby obtaining the trained positioning model.

9. The monocular vision instant positioning method based on digital twin data and semantic information according to claim 8, characterized in that: in, The loss function of the adversarial training includes regression loss and classification loss The calculation expression of the regression loss is: Where x is the coordinate of the virtual camera in the pose information predicted by the positioning model, is the real coordinate of the virtual camera in the label, q is the azimuth of the virtual camera in the pose information predicted by the positioning model, is the true azimuth of the virtual camera in the label, γ is the hyperparameter error ratio, The classification loss The calculation expression is: Where D is the discriminator, G(I t ) is the positioning model according to the input image data I t The resulting decoded image.

Citation Information

Patent Citations

  • Indoor positioning method based on digital twin building and heterogeneous feature fusion

    CN114742995A

  • Method and device for intelligently generating element cosmic space

    CN116310500A

  • Remote sensing image visual positioning method based on text guidance

    CN116958829A

  • Deep learning monocular vision ground dynamic target three-dimensional reconstruction method

    CN116977566A

  • Visual positioning system

    CN119006749A