Image positioning method and device, storage medium and electronic equipment

By using an adaptive spatial transformation image retrieval and target viewpoint matching model, the problems of viewpoint differences and resolution inconsistencies between low-altitude video images and remote sensing orthophotos were solved, achieving high-precision spatial positioning and improving positioning accuracy.

CN121170618APending Publication Date: 2025-12-19CHINA TOWER CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511323395.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

Existing technologies suffer from low positioning accuracy in matching and locating low-altitude video images and remote sensing orthophotos due to differences in perspective and resolution.

Method used

By using an adaptive spatial transformation image retrieval model and a target viewpoint matching model, accurate retrieval and matching of low-altitude video images and remote sensing orthophotos are achieved, the three-dimensional coordinates of the matching points are recovered, the exterior orientation elements are determined, and the data are converted to geographic coordinates.

Benefits of technology

It improves the matching accuracy between low-altitude video images and remote sensing orthophotos, enabling high-precision spatial positioning and enhancing positioning accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121170618A_ABST
    Figure CN121170618A_ABST
Patent Text Reader

Abstract

The invention discloses an image positioning method and device, a storage medium and electronic equipment. Relates to the field of image processing, and the method comprises the steps: obtaining a low-altitude video image to be positioned, the low-altitude video image being any frame of image in a target low-altitude video; a target remote sensing orthographic image corresponding to the low-altitude video image is determined from the remote sensing orthographic image set through a target image retrieval model, and the target remote sensing orthographic image and the low-altitude video image have the same coverage area; determining a matching coordinate pair between the low-altitude video image and the target remote sensing orthoimage through a target visual angle matching model according to the low-altitude video image and the target remote sensing orthoimage; and determining an exterior orientation element of the low-altitude video image according to the matched coordinate pair, and determining a geographic coordinate of the low-altitude video image according to the exterior orientation element. According to the method and the device, the problem of relatively low positioning accuracy when image positioning is carried out on a low-altitude video image based on a remote sensing image in related technologies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of image processing, in particular to an image positioning method and device, a storage medium and an electronic device. BACKGROUND

[0002] Low-altitude videos, usually captured by unmanned aerial vehicles, high-point cameras and other devices, can capture detailed information in various scenarios due to their flexibility and real-time performance. Remote sensing images provide a macroscopic view of geographical coverage with high spatial resolution and a wide field of view. If the two can be effectively integrated, the spatio-temporal perception capability of the target area will be greatly enhanced. However, the current technology faces severe challenges in achieving accurate matching and positioning of low-altitude video images and remote sensing orthographic images: the image retrieval method in the related art cannot adapt to the differences in viewing angle, affine transformation, and inconsistent resolution between low-altitude oblique images and remote sensing orthographic images, resulting in low accuracy of global feature matching and large spatial positioning error; the feature point matching method in the related art faces the differences in viewing angle, image resolution, terrain perspective distortion, and lighting conditions between low-altitude video images and remote sensing orthographic images, and has problems such as sparse feature point coverage, poor matching robustness, geometric structure mismatch, and sensitivity to scale changes, resulting in low positioning accuracy.

[0003] In view of the low positioning accuracy when the related art uses remote sensing images to position low-altitude video images, there is currently no effective solution. SUMMARY

[0004] The main purpose of the present application is to provide an image positioning method and device, a storage medium and an electronic device to solve the problem of low positioning accuracy when the related art uses remote sensing images to position low-altitude video images.

[0005] In order to achieve the above purpose, according to one aspect of the present application, an image positioning method is provided. The method comprises: acquiring a low-altitude video image to be positioned, wherein the low-altitude video image is any one frame image in a target low-altitude video; determining a target remote sensing orthographic image corresponding to the low-altitude video image from a remote sensing orthographic image set through a target image retrieval model, wherein the target remote sensing orthographic image has the same coverage area as the low-altitude video image; determining a matching coordinate pair between the low-altitude video image and the target remote sensing orthographic image according to the low-altitude video image and the target remote sensing orthographic image through a target viewing angle matching model; determining the exterior orientation elements of the low-altitude video image according to the matching coordinate pair, and determining the geographic coordinates of the low-altitude video image according to the exterior orientation elements.

[0006] Further, the target image retrieval model is obtained by the following steps: obtaining a first sample low-altitude video image set and a first sample remote sensing orthographic image set; processing each first sample low-altitude video image in the first sample low-altitude video image set and each first sample remote sensing orthographic image in the first sample remote sensing orthographic image set by an embedding processing module of the initial image retrieval model, to obtain a first sample low-altitude video image feature of each first sample low-altitude video image and a first sample remote sensing orthographic image feature of each first sample remote sensing orthographic image; performing normalization processing on the first sample low-altitude video image feature by a spatial transformation module of the initial image retrieval model, to obtain a normalized first sample low-altitude video image feature; processing the normalized first sample low-altitude video image feature and each first sample remote sensing orthographic image feature by an image feature enhancement module of the initial image retrieval model, to obtain a processed image feature of each first sample low-altitude video image and a processed image feature of each first sample remote sensing orthographic image; determining a target loss function of the initial image retrieval model according to the processed image feature of each first sample low-altitude video image and the processed image feature of each first sample remote sensing orthographic image by a feature matching module of the initial image retrieval model, and training the initial image retrieval model based on the target loss function until a predetermined training target is reached, to obtain the target image retrieval model.

[0007] Further, the processing of the normalized first sample low-altitude video image feature and each first sample remote sensing orthographic image feature by the image feature enhancement module of the initial image retrieval model to obtain the processed image feature of each first sample low-altitude video image and the processed image feature of each first sample remote sensing orthographic image includes: processing the normalized first sample low-altitude video image feature and each first sample remote sensing orthographic image feature by the image feature enhancement module to obtain an enhanced feature of each first sample low-altitude video image and an enhanced feature of each first sample remote sensing orthographic image; and performing information interaction processing on the enhanced feature of each first sample low-altitude video image and the enhanced feature of each first sample remote sensing orthographic image based on a preset processing operator, to obtain the processed image feature of each first sample low-altitude video image and the processed image feature of each first sample remote sensing orthographic image.

[0008] Further, the target loss function of the initial image retrieval model is determined by the feature matching module of the initial image retrieval model according to the processed image features of each first sample low-altitude video image and the processed image features of each first sample remote sensing orthographic image, comprising: the feature matching module processes the processed image features of each first sample low-altitude video image and the processed image features of each first sample remote sensing orthographic image respectively to obtain global image features of each first sample low-altitude video image and global image features of each first sample remote sensing orthographic image; for each first sample low-altitude video image, the global image features of the first sample low-altitude video image and the global image features of each first sample remote sensing orthographic image are calculated for similarity to obtain global image feature similarities between the first sample low-altitude video image and each first sample remote sensing orthographic image; the processed image features of the first sample low-altitude video image and the processed image features of each first sample remote sensing orthographic image are calculated for similarity to obtain local image feature similarities between the first sample low-altitude video image and each first sample remote sensing orthographic image; the global image feature similarities and the local image feature similarities are summed to obtain target image feature similarities between the first sample low-altitude video image and each first sample remote sensing orthographic image; and the target loss function is determined according to the global image feature similarities, the local image feature similarities and the target image feature similarities.

[0009] Further, the target view matching model is obtained by the following steps: obtaining a training sample set, wherein the training sample set includes paired second sample low-altitude video images and second sample remote sensing orthographic images; the encoding network of the initial view matching model processes the second sample low-altitude video images and the second sample remote sensing orthographic images respectively to obtain first encoding features and second encoding features; the attention processing module of the initial view matching model processes the first encoding features and the second encoding features respectively to obtain a first set of attention feature representations and a second set of attention feature representations; the first set of attention feature representations and the second set of attention feature representations are gradually matched to obtain sample matching coordinate pairs between the second sample low-altitude video images and the second sample remote sensing orthographic images, wherein the gradual matching starts from the lowest resolution features in the first set of attention feature representations and the second set of attention feature representations and iterates gradually to the highest resolution features, and the matching coordinate pairs predicted each time are used as prior information for the next iteration; and the initial view matching model is trained based on the sample matching coordinate pairs until a predetermined training target is reached to obtain the target view matching model.

[0010] Further, the determining the exterior orientation elements of the low-altitude video image according to the matching coordinate pairs comprises: mapping the matching coordinate pairs to a coordinate system of the digital surface model, determining three-dimensional space coordinates corresponding to each matching point in the matching coordinate pairs according to ground elevation information of the digital surface model; and determining the exterior orientation elements of the low-altitude video image according to the three-dimensional space coordinates and a preset camera pose solving algorithm.

[0011] Further, the determining the geographic coordinates of the low-altitude video image according to the exterior orientation elements comprises: determining depth information of the low-altitude video image according to the depth estimation network; and converting pixel points of the low-altitude video image from a camera coordinate system to a world coordinate system to obtain the geographic coordinates of the low-altitude video image according to the depth information and the exterior orientation elements.

[0012] To achieve the above object, according to another aspect of the present application, an image positioning device is provided. The device comprises: a first acquisition unit configured to acquire a low-altitude video image to be positioned, wherein the low-altitude video image is any one frame image in a target low-altitude video; a first determination unit configured to determine a target remote sensing orthographic image corresponding to the low-altitude video image from a remote sensing orthographic image set through a target image retrieval model, wherein the target remote sensing orthographic image has the same coverage area as the low-altitude video image; a second determination unit configured to determine matching coordinate pairs between the low-altitude video image and the target remote sensing orthographic image according to the low-altitude video image and the target remote sensing orthographic image through a target view angle matching model; and a third determination unit configured to determine exterior orientation elements of the low-altitude video image according to the matching coordinate pairs, and determine geographic coordinates of the low-altitude video image according to the exterior orientation elements.

[0013] Further, the apparatus further comprises the following units for obtaining the target image retrieval model by the following steps: a second acquisition unit configured to acquire a first sample low-altitude video image set and a first sample remote sensing orthographic image set; a first processing unit configured to process each first sample low-altitude video image in the first sample low-altitude video image set and each first sample remote sensing orthographic image in the first sample remote sensing orthographic image set respectively by an embedding processing module of an initial image retrieval model to obtain a first sample low-altitude video image feature of each first sample low-altitude video image and a first sample remote sensing orthographic image feature of each first sample remote sensing orthographic image; a second processing unit configured to perform normalization processing on the first sample low-altitude video image feature by a spatial transformation module of the initial image retrieval model to obtain a normalized first sample low-altitude video image feature; a third processing unit configured to process the normalized first sample low-altitude video image feature and each first sample remote sensing orthographic image feature respectively by an image feature enhancement module of the initial image retrieval model to obtain a processed image feature of each first sample low-altitude video image and a processed image feature of each first sample remote sensing orthographic image; and a fourth processing unit configured to determine a target loss function of the initial image retrieval model according to the processed image feature of each first sample low-altitude video image and the processed image feature of each first sample remote sensing orthographic image by a feature matching module of the initial image retrieval model, and train the initial image retrieval model based on the target loss function until a predetermined training target is reached to obtain the target image retrieval model.

[0014] Further, the third processing unit comprises: a first processing sub-unit configured to process the normalized first sample low-altitude video image feature and each first sample remote sensing orthographic image feature respectively by the image feature enhancement module to obtain an enhanced feature of each first sample low-altitude video image and an enhanced feature of each first sample remote sensing orthographic image; and a second processing sub-unit configured to perform information interaction processing on the enhanced feature of each first sample low-altitude video image and the enhanced feature of each first sample remote sensing orthographic image respectively based on a preset processing operator to obtain the processed image feature of each first sample low-altitude video image and the processed image feature of each first sample remote sensing orthographic image.

[0015] Further, the fourth processing unit comprises: a third processing subunit, configured to process the processed image features of each first sample low-altitude video image and the processed image features of each first sample remote sensing orthographic image respectively by the feature matching module, to obtain global image features of each first sample low-altitude video image and global image features of each first sample remote sensing orthographic image; a fourth processing subunit, configured to, for each first sample low-altitude video image, perform similarity calculation on the global image features of the first sample low-altitude video image and the global image features of each first sample remote sensing orthographic image respectively, to obtain global image feature similarities between the first sample low-altitude video image and each first sample remote sensing orthographic image; a fifth processing subunit, configured to perform similarity calculation on the processed image features of the first sample low-altitude video image and the processed image features of each first sample remote sensing orthographic image respectively, to obtain local image feature similarities between the first sample low-altitude video image and each first sample remote sensing orthographic image; a sixth processing subunit, configured to perform summation calculation on the global image feature similarities and the local image feature similarities respectively, to obtain target image feature similarities between the first sample low-altitude video image and each first sample remote sensing orthographic image; and a seventh processing subunit, configured to determine a target loss function according to the global image feature similarities, the local image feature similarities and the target image feature similarities.

[0016] Further, the device further comprises the following units for obtaining the target view angle matching model by the following steps: a third acquisition unit, configured to acquire a training sample set, wherein the training sample set comprises paired second sample low-altitude video images and second sample remote sensing orthographic images; a fifth processing unit, configured to process the second sample low-altitude video images and the second sample remote sensing orthographic images respectively by an encoding network of the initial view angle matching model, to obtain first encoding features and second encoding features; a sixth processing unit, configured to process the first encoding features and the second encoding features respectively by an attention processing module of the initial view angle matching model, to obtain a first set of attention feature representations and a second set of attention feature representations; a seventh processing unit, configured to perform progressive matching based on the first set of attention feature representations and the second set of attention feature representations, to obtain sample matching coordinate pairs between the second sample low-altitude video images and the second sample remote sensing orthographic images, wherein the progressive matching starts from the lowest resolution features in the first set of attention feature representations and the second set of attention feature representations, and iterates to the highest resolution features gradually, and the matching coordinate pairs predicted each time are used as prior information for the next iteration; and an eighth processing unit, configured to train the initial view angle matching model based on the sample matching coordinate pairs until a predetermined training target is reached, to obtain the target view angle matching model.

[0017] Further, the third determining unit comprises: a first determining subunit, configured to map the matching coordinate pair to a coordinate system of the digital surface model, and determine a three-dimensional space coordinate corresponding to each matching point in the matching coordinate pair according to ground elevation information of the digital surface model; and a second determining subunit, configured to determine the exterior orientation elements of the low-altitude video image according to the three-dimensional space coordinate and a preset camera pose solving algorithm.

[0018] Further, the third determining unit further comprises: a third determining subunit, configured to determine depth information of the low-altitude video image according to the depth estimation network; and a conversion subunit, configured to convert a pixel point of the low-altitude video image from a camera coordinate system to a world coordinate system according to the depth information and the exterior orientation elements, to obtain geographic coordinates of the low-altitude video image.

[0019] According to another aspect of the embodiments of the present application, an electronic device is further provided, comprising: a memory storing an executable program; and a processor configured to run the program, wherein the program is configured to execute the image positioning method of any one of the above aspects when running.

[0020] According to another aspect of the embodiments of the present application, a computer readable storage medium is further provided, the storage medium storing a program, wherein the program is configured to control a device where the storage medium is located to execute the image positioning method of any one of the above aspects when running.

[0021] In the embodiments of the present application, the following steps are adopted: obtaining a low-altitude video image to be positioned, wherein the low-altitude video image is any one frame image in a target low-altitude video; determining a target remote sensing orthographic image corresponding to the low-altitude video image from a remote sensing orthographic image set through a target image retrieval model, wherein the target remote sensing orthographic image has the same coverage area as the low-altitude video image; determining a matching coordinate pair between the low-altitude video image and the target remote sensing orthographic image according to the low-altitude video image and the target remote sensing orthographic image through a target view angle matching model; determining exterior orientation elements of the low-altitude video image according to the matching coordinate pair, and determining geographic coordinates of the low-altitude video image according to the exterior orientation elements. The technical problem of low positioning accuracy in the related art when positioning the low-altitude video image based on remote sensing images is solved.

[0022] In the present solution, adaptive spatial transformation image retrieval is used to achieve accurate retrieval of remote sensing orthographic images under large parallax, accurate pixel matching between remote sensing orthographic images and low-altitude video images is completed based on the retrieval results, and the exterior orientation elements are determined by restoring the three-dimensional coordinates of the matching points, so as to realize pixel-level geographic coordinate back calculation, effectively solve the problems of view angle difference and insufficient matching accuracy between the low-altitude video image and the remote sensing orthographic image, achieve high-precision spatial positioning, and improve the positioning accuracy. BRIEF DESCRIPTION OF DRAWINGS

[0023] The accompanying drawings, which form a part of this application, are included to provide a further understanding of the application and are incorporated in and constitute a part of this application. The illustrations, together with the description, serve to explain the application, but are not intended to limit the application.

[0024] Figure 1 A hardware structure block diagram of a computer terminal for implementing the image positioning method is shown;

[0025] Figure 2 A flow chart of the image positioning method according to an embodiment of the application is shown;

[0026] Figure 3 A schematic diagram of the image positioning flow according to an embodiment of the application is shown;

[0027] Figure 4 A schematic diagram of the image positioning device according to an embodiment of the application is shown;

[0028] Figure 5 A structure block diagram of an electronic device according to an embodiment of the application is shown. DETAILED DESCRIPTION

[0029] In order to make the personnel in the art better understand the scheme of the present application, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should belong to the scope of protection of the present application.

[0030] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily mean a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0031] It should be noted that the information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for display, analyzed data, etc.) involved in the present application are information and data authorized by the user or authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of related data comply with relevant laws, regulations and standards, necessary security measures are taken, do not violate public order and good customs, and provide corresponding operation portal for user to choose authorization or refusal. For example, the system and the interface between the related users or institutions are provided with the corresponding operation portal for the user to choose to agree or refuse the automatic decision result; if the user chooses to refuse, the expert decision process is entered.

[0032] Embodiment 1

[0033] According to the embodiments of the present application, a method embodiment of image positioning is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.

[0034] The method embodiment provided by the embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing the image positioning method is shown. As shown in Figure 1 The computer terminal 10 (or mobile device) can include one or more processors 102 (the processor 102 can include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it can also include a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. Those skilled in the art can understand that Figure 1 The structure shown is only schematic, which does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 can also include more or less components than those shown in Figure 1 or have a different configuration than that shown in Figure 1 .

[0035] It should be noted that the one or more processors 102 and / or other data processing circuitry described above can be referred to herein generally as "data processing circuitry". The data processing circuitry can be embodied in whole or in part as software, hardware, firmware, or any combination thereof. Furthermore, the data processing circuitry can be a single standalone processing module, or incorporated in whole or in part within any of the other elements of the computer terminal 10 (or mobile device). As referred to in the embodiments herein, the data processing circuitry acts as a processor to control, for example, the selection of the variable resistance terminal path in connection with the interface.

[0036] The memory 104 can be used to store software programs of application software and modules, such as program instructions / data storage means corresponding to the image positioning method in the embodiments herein. The processor 102 can execute various functional applications and data processing, i.e. implement the image positioning method described above, by running the software programs and modules stored in the memory 104. The memory 104 can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 104 can further include memories disposed remotely with respect to the processor 102, which can be connected to the computer terminal 10 through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0037] The transmission device 106 is configured to receive or send data via a network. Examples of the network include, but are not limited to, a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network interface controller (NIC) that can be connected to other network devices through a base station to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module configured to communicate with the Internet in a wireless manner.

[0038] The display can be a touch screen liquid crystal display (LCD) that enables a user to interact with the user interface of the computer terminal 10 (or mobile device).

[0039] In the above-described operating environment, the embodiments herein provide an image positioning method as shown in Figure 2 Figure 2 is a flowchart of the image positioning method according to the first embodiment herein. The image positioning method includes the following steps:

[0040] In step S201, a low-altitude video image to be positioned is acquired, wherein the low-altitude video image is any frame image in a target low-altitude video. ​

[0041] Optionally, the image positioning method provided by the embodiments of the present application can be widely applied to smart city management, farmland protection supervision, ecological environment monitoring, emergency rescue command and the like. For example, when it is necessary to monitor an emergency in a city, a low-altitude video (i.e., a target low-altitude video) can be captured by using a low-altitude flying or hanging device such as a drone, a helicopter, a camera and the like, and an image positioning is started from any one frame of the image, and the target is to locate the position of the event in the video according to the remote sensing orthographic image.

[0042] In step S202, a target remote sensing orthographic image corresponding to the low-altitude video image is determined from the remote sensing orthographic image set by using the target image retrieval model, wherein the target remote sensing orthographic image has the same coverage area as the low-altitude video image.

[0043] Optionally, first, adaptive spatial transformation image retrieval is performed, and one or more target remote sensing orthographic images having the same coverage area as the low-altitude video image are retrieved from the remote sensing orthographic image set by using the target image retrieval model.

[0044] Optionally, in the image positioning method provided by the embodiments of the present application, the target image retrieval model is obtained by the following steps: obtaining a first sample low-altitude video image set and a first sample remote sensing orthographic image set; processing each first sample low-altitude video image in the first sample low-altitude video image set and each first sample remote sensing orthographic image in the first sample remote sensing orthographic image set by using an embedding processing module of an initial image retrieval model, to obtain a first sample low-altitude video image feature of each first sample low-altitude video image and a first sample remote sensing orthographic image feature of each first sample remote sensing orthographic image; performing normalization processing on the first sample low-altitude video image feature by using a spatial transformation module of the initial image retrieval model, to obtain a normalized first sample low-altitude video image feature; processing the normalized first sample low-altitude video image feature and each first sample remote sensing orthographic image feature by using an image feature enhancement module of the initial image retrieval model, to obtain a processed image feature of each first sample low-altitude video image and a processed image feature of each first sample remote sensing orthographic image; determining a target loss function of the initial image retrieval model according to the processed image feature of each first sample low-altitude video image and the processed image feature of each first sample remote sensing orthographic image by using a feature matching module of the initial image retrieval model, and training the initial image retrieval model based on the target loss function until a predetermined training target is reached, to obtain the target image retrieval model.

[0045] In an optional embodiment, model training is first needed, and a first sample low-altitude video image set D={d ii = 1, 2…, n}, n represents the number of first sample low-altitude video images, the first sample remote sensing orthographic image set S = {s j j = 1, 2…, m}, m represents the number of first sample remote sensing orthographic images. For a given low-altitude video image d i (i.e. the first sample low-altitude video image) and a remote sensing orthographic image s j (i.e. the first sample remote sensing orthographic image), the basic features are extracted by the embedding processing module (patch embedding module) of the initial image retrieval model, represented as follows:

[0046] f di = Pe(d i )

[0047] f sj = Pe(s j )

[0048] wherein Pe() represents the patch embedding module, f di , f sj ∈ R C×H×W C×H×W respectively represent the basic features extracted from d i and s j , i.e. the first sample low-altitude video image features and the first sample remote sensing orthographic image features, C represents the number of feature channels, and H and W represent the height and width of the features.

[0049] Since the low-altitude video images have diverse perspectives and significant differences in resolution, in order to make f di and f sj similar in feature distribution, the spatial transformation module of the initial image retrieval model is used to perform spatial transformation, i.e. normalization, on f di , represented as follows:

[0050] T = STN(f di )

[0051]

[0052] wherein STN() represents the spatial transformation network, T ∈ R 2×3 is an affine transformation matrix, SpatialTransform() represents the spatial set transformation function, is the normalized low-altitude image feature, i.e. the normalized first sample low-altitude video image feature.

[0053] Optionally, in the image positioning method provided in the embodiments of the present application, the image feature enhancement module of the initial image retrieval model is used to process the normalized first sample low-altitude video image features and each first sample remote sensing orthographic image feature respectively, to obtain the processed image features of each first sample low-altitude video image and the processed image features of each first sample remote sensing orthographic image, including: the image feature enhancement module is used to process the normalized first sample low-altitude video image features and each first sample remote sensing orthographic image feature respectively, to obtain the enhanced features of each first sample low-altitude video image and the enhanced features of each first sample remote sensing orthographic image; and the preset processing operator is used to perform information interaction processing on the enhanced features of each first sample low-altitude video image and the enhanced features of each first sample remote sensing orthographic image respectively, to obtain the processed image features of each first sample low-altitude video image and the processed image features of each first sample remote sensing orthographic image.

[0054] In an optional embodiment, in image retrieval, landmark pair matching low-altitude video and remote sensing image plays an important role, so in obtaining After that, the image feature enhancement module (i.e., landmark enhancement type self-interaction module) of the initial image retrieval model is used to process and f sj to extract features, which are represented as follows:

[0055]

[0056] M sj = Sigmoid(Conv(f sj ))

[0057]

[0058] wherein Conv() represents a convolution network function, Sigmoid() represents a sigmoid function, and represents point-by-point multiplication, M di , M sj respectively represent the feature attention maps extracted from and f sj , and f respectively represent the landmark enhanced features extracted in the low-altitude video image and the remote sensing image, i.e., the enhanced features of the first sample low-altitude video image and the enhanced features of the first sample remote sensing orthographic image.

[0059] After obtaining , a landmark-guided transformer operator (i.e., a preset processing operator) is used to realize information interaction between features, which is represented as follows:

[0060] k di = TopKL(M di )

[0061]

[0062] k sj = TopKL(M sj )

[0063]

[0064] where TopKL() denotes the position index of K maximum values, k di and k sj denote the position index of K maximum values obtained from M di and M sj , respectively, the Grid() function denotes sampling according to the index position, key di and key sj denote the features sampled from and according to the position index of K maximum values. Cross_attn() denotes the cross attention function, which takes or as query, key di or key sj as both key and value, and denote the final output of the landmark-oriented transformer, that is, the processed image features of each first sample low-altitude video image and the processed image features of each first sample remote sensing orthographic image.

[0065] Optionally, in the image positioning method provided by the embodiment of the present application, the target loss function of the initial image retrieval model is determined by the feature matching module of the initial image retrieval model according to the processed image features of each first sample low-altitude video image and the processed image features of each first sample remote sensing orthographic image, comprising: the feature matching module processes the processed image features of each first sample low-altitude video image and the processed image features of each first sample remote sensing orthographic image respectively to obtain global image features of each first sample low-altitude video image and global image features of each first sample remote sensing orthographic image; for each first sample low-altitude video image, similarity calculation is performed on the global image features of the first sample low-altitude video image and the global image features of each first sample remote sensing orthographic image respectively to obtain global image feature similarities between the first sample low-altitude video image and each first sample remote sensing orthographic image; similarity calculation is performed on the processed image features of the first sample low-altitude video image and the processed image features of each first sample remote sensing orthographic image respectively to obtain local image feature similarities between the first sample low-altitude video image and each first sample remote sensing orthographic image; sum calculation is performed on the global image feature similarities and the local image feature similarities respectively to obtain target image feature similarities between the first sample low-altitude video image and each first sample remote sensing orthographic image; and the target loss function is determined according to the global image feature similarities, the local image feature similarities and the target image feature similarities.

[0066] In an optional embodiment, the global features and the local features are used to compare the similarities of the low-altitude video images and the remote sensing images, the feature matching module (i.e. the global-local feature similarity matching module) processes the processed image features of each first sample low-altitude video image and the processed image features of each first sample remote sensing orthographic image respectively to obtain global image features of each first sample low-altitude video image and global image features of each first sample remote sensing orthographic image, for each first sample low-altitude video image, similarity calculation is performed on the global image features of the first sample low-altitude video image and the global image features of each first sample remote sensing orthographic image respectively to obtain global image feature similarities between the first sample low-altitude video image and each first sample remote sensing orthographic image, similarity calculation is performed on the processed image features of the first sample low-altitude video image and the processed image features of each first sample remote sensing orthographic image respectively to obtain local image feature similarities between the first sample low-altitude video image and each first sample remote sensing orthographic image, sum calculation is performed on the global image feature similarities and the local image feature similarities respectively to obtain target image feature similarities between the first sample low-altitude video image and each first sample remote sensing orthographic image, which are represented as follows:

[0067]

[0068] wherein, maxp() represents a maxpooling function for obtaining global features, GSim() represents a global feature similarity calculation function, represents the global feature similarity between the low-altitude video image d i and the remote sensing image s j , LSim() represents a local feature similarity calculation function, represents the local feature similarity between the low-altitude video image d i and the remote sensing image s j , represents the final global-local feature similarity, i.e., the target image feature similarity.

[0069] Further, a target loss function is determined according to the global image feature similarity, the local image feature similarity, and the target image feature similarity. In the training phase, a multi-modal collaborative consistency loss is adopted to constrain the image features of the same modality and the same coverage area, the image features of the cross-modality and the same coverage area to be close to each other, and the image features of the same modality and different coverage areas, the image features of the cross-modality and different coverage areas to be far away from each other. The specific loss function is:

[0070] loss l = Cr(Ls ds ) + Cr(Ls dd ) + Cr(Ls ss )

[0071] loss g = Cr(Gs ds ) + Cr(Gs dd ) + Cr(Gs ss )

[0072] loss gl = Cr(Fs ds ) + Cr(Fs dd ) + Cr(Fs ss )

[0073] loss = loss l + loss g + loss gl

[0074] wherein, Cr() represents a cross-entropy function, Ls ds , Gs ds , Fs ds represent the local feature similarity matrix, the global feature similarity matrix, and the global-local feature similarity matrix between the low-altitude video image and the remote sensing image, respectively, Ls dd , Gs dd , Fs ddrespectively represent the local feature similarity matrix, the global feature similarity matrix, the global-local feature similarity matrix between low-altitude video images, Ls ss ss ss respectively represent the local feature similarity matrix, the global feature similarity matrix, the global-local feature similarity matrix between remote sensing images, loss l loss represents the local feature matching loss function, loss g loss represents the global feature matching loss function, loss gl loss represents the global-local feature matching loss function, loss is the final model training loss function, that is, the target loss function.

[0075] Based on the target loss function, the initial image retrieval model is trained until a predetermined training target is reached, for example, the loss reaches a set threshold or a predetermined training round is completed, to obtain an adaptive spatial transformation image retrieval network model (i.e., a target image retrieval model). After obtaining the adaptive spatial transformation image retrieval network model, for a given low-altitude video image, a remote sensing image with the same coverage range can be retrieved.

[0076] In step S203, the matching coordinate pairs between the low-altitude video image and the target remote sensing orthographic image are determined by the target view angle matching model according to the low-altitude video image and the target remote sensing orthographic image.

[0077] Optionally, in the matching stage, the progressive multi-order cross-view angle matching technology is used to accurately match the pixels in the low-altitude video image with the pixels in the target remote sensing orthographic image through the progressive multi-order cross-view angle matching network model (i.e., the target view angle matching model), to obtain a series of matching point coordinates, i.e., matching coordinate pairs.

[0078] ​​Optionally, in the image positioning method provided in the embodiments of the present application, the target view matching model is obtained through the following steps: a training sample set is obtained, wherein the training sample set includes paired second sample low-altitude video images and second sample remote sensing orthographic images; the second sample low-altitude video images and the second sample remote sensing orthographic images are processed through the encoding network of the initial view matching model respectively to obtain first encoding features and second encoding features; the first encoding features and the second encoding features are processed through the attention processing module of the initial view matching model respectively to obtain a first set of attention feature representations and a second set of attention feature representations; sample matching coordinate pairs between the second sample low-altitude video images and the second sample remote sensing orthographic images are obtained based on the gradual matching of the first set of attention feature representations and the second set of attention feature representations, wherein the gradual matching starts from the lowest resolution features in the first set of attention feature representations and the second set of attention feature representations and iterates gradually to the highest resolution features, and the matching coordinate pairs predicted each time are used as prior information for the next iteration; the initial view matching model is trained based on the sample matching coordinate pairs until a predetermined training target is reached to obtain the target view matching model.

[0079] In an optional embodiment, to achieve pixel-level image matching, a progressive multi-order cross-view matching network is proposed, including a multi-scale window-cross window attention feature extraction stage, a progressive matching stage, and a symmetry progressive training stage. In the multi-scale window-cross window attention feature extraction stage, let the source image be I src , the reference image be I ref , I src and I ref be input into the encoding network to obtain the encoding features F src and F ref , i.e., the first encoding features and the second encoding features. Then input the multi-scale window-cross window center attention module (i.e., the attention processing module), wherein the multi-scale window-cross window center attention module includes a multi-scale window attention extraction module and a cross window center attention module. The multi-scale window attention extraction module calculates the window attention features in each scale by dividing the feature map into multiple scale windows, and the cross window center attention module takes the feature vector at the geometric center of each window as the center token, constructs the key and value based on the center token, and takes the feature map as the query, so that the information of different windows can be transmitted to the entire feature map through the center token, realizing approximate global information interaction. Let there be k multi-scale window-cross window center attention modules, then the output of I src and I ref at each multi-scale window-cross window center attention module is and i.e. the first set of attention feature representations and the second set of attention feature representations. Wherein, the feature resolution of is twice that of .

[0080] In the progressive matching stage, first, and are input into the matching network to predict the corresponding matching point position of the pixel point in in , and this point position is taken as a priori, together with the high-resolution feature , to input into the matching network to predict the corresponding matching points between , and as a priori, to predict the corresponding matching points between . Repeat the above process until the corresponding matching points between are predicted as the final matching points, completing the pixel-level matching between the source image I src and the reference image I ref , i.e. progressive matching based on the first set of attention feature representations and the second set of attention feature representations, to obtain the sample matching coordinate pair between the second sample low-altitude video image and the second sample remote sensing orthographic image.

[0081] Further, based on the sample matching coordinate pair, the initial view angle matching model is trained until a predetermined training target is reached, e.g. the loss reaches a set threshold or a predetermined number of training rounds is completed, to obtain the target view angle matching model. In the symmetric progressive training stage, the progressive multi-order cross-view matching network is trained using a symmetric progressive training method, using two modes of training: taking the sample low-altitude video image as the source image and the sample remote sensing orthographic image as the reference image, and taking the sample remote sensing orthographic image as the source image and the sample low-altitude video image as the reference image.

[0082] Step S204, determine the exterior orientation elements of the low-altitude video image according to the matching coordinate pair, and determine the geographic coordinates of the low-altitude video image according to the exterior orientation elements.

[0083] Optionally, after obtaining the cross-view matching points (i.e. matching coordinate pairs), the exterior orientation elements (i.e. external pose: position and orientation) of the camera that took the low-altitude video image are determined through the three-dimensional coordinate information of the matching points combined with the digital surface model data, and then the geographic coordinates of the pixel points in the low-altitude video image are calculated based on the depth map of the low-altitude video image and the exterior orientation elements, achieving precise geographic positioning of the event, so that resources can be quickly dispatched for response.

[0084] ​In summary, the adaptive spatial transformation image retrieval is used to realize accurate retrieval of remote sensing orthographic images under large parallax, accurate pixel matching of the remote sensing orthographic images and the low-altitude video images is completed on the basis of the retrieval results, the exterior orientation elements are determined by restoring the three-dimensional coordinates of the matching points, the pixel-level geographic coordinate back calculation is realized, the problems of the view angle difference and the insufficient matching accuracy between the low-altitude video images and the remote sensing orthographic images are effectively solved, the high-precision spatial positioning is realized, and the positioning accuracy is improved.

[0085] Optionally, in the image positioning method provided in the embodiments of the present application, determining the exterior orientation elements of the low-altitude video image according to the matching coordinate pairs comprises: mapping the matching coordinate pairs to the coordinate system of the digital surface model, determining the three-dimensional space coordinates corresponding to each matching point in the matching coordinate pairs according to the ground elevation information of the digital surface model; and determining the exterior orientation elements of the low-altitude video image according to the three-dimensional space coordinates and a preset camera pose solving algorithm.

[0086] In an optional embodiment, the matching coordinate pairs are mapped to the coordinate system of the digital surface model, for each matching point, the three-dimensional space coordinates thereof are restored by using the ground elevation provided by the digital surface model and combining the known geographic coordinates of the remote sensing orthographic image, a set of two-dimensional-three-dimensional correspondence relationships are obtained, then the exterior orientation elements (the three-dimensional coordinates of the camera center and the three-axis rotation angles) of the low-altitude video image are estimated by using the two-dimensional-three-dimensional correspondence relationships and the efficient perspective n-point and random sample consensus algorithm (i.e., the preset camera pose solving algorithm).

[0087] Optionally, in the image positioning method provided in the embodiments of the present application, determining the geographic coordinates of the low-altitude video image according to the exterior orientation elements comprises: determining the depth information of the low-altitude video image according to a depth estimation network; and converting the pixel points of the low-altitude video image from the camera coordinate system to the world coordinate system according to the depth information and the exterior orientation elements, to obtain the geographic coordinates of the low-altitude video image.

[0088] In an optional embodiment, the depth estimation network is used on the low-altitude video image sequence, and a dense depth map (i.e., depth information) consistent with the original resolution is output, then based on the known low-altitude video camera intrinsic parameters, the exterior orientation elements and the corresponding depth values, the pixels are first back-projected to the camera coordinate system, then are transformed to the world coordinate system according to the exterior orientation elements, to obtain the geographic coordinates of the low-altitude video image, and realize visual positioning.

[0089] In an optional embodiment, Figure 3 is a schematic diagram of the image positioning process provided in the embodiments of the present application, as Figure 3As shown, four core modules of adaptive spatial transformation image retrieval, progressive multi-order cross-view matching, exterior orientation element estimation and visual positioning estimation based on depth map are integrated to realize coarse retrieval to precise positioning. Among them, adaptive spatial transformation image retrieval includes adaptive spatial transformation image retrieval network model training and remote sensing orthographic image retrieval. A cross-modal image retrieval method is adopted, which fuses adaptive spatial transformation, landmark enhanced self-interaction and global-local similarity fusion. Through adaptive spatial transformation normalization feature distribution, explicit mining and interaction of key landmark information, and collaborative use of global semantic and local detail features, the retrieval accuracy and positioning range accuracy of low-altitude video images to remote sensing orthographic images under complex parallax conditions are significantly improved. Based on the multi-scale window-cross window center attention and the progressive multi-order strategy cross-view matching method, combined with efficient multi-scale global context modeling, coarse-to-fine iterative matching optimization and symmetry training mechanism, the low-altitude image and remote sensing image are precisely matched under the conditions of large parallax, terrain fluctuation and scale change.

[0090] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown.

[0091] Embodiment 2

[0092] The embodiment of the present application also provides an image positioning device. It should be noted that the image positioning device of the embodiment of the present application can be used to execute the image positioning method provided by the embodiment of the present application. The image positioning device provided by the embodiment of the present application is introduced as follows.

[0093] According to the embodiment of the present application, an image positioning device for implementing the above-mentioned image positioning method is also provided, as shown in the figure, the device comprises: a first acquisition unit 401, a first determination unit 402, a second determination unit 403 and a third determination unit 404. Figure 4

[0094] The first acquisition unit 401 is configured to acquire a low-altitude video image to be positioned, wherein the low-altitude video image is any frame image in a target low-altitude video.

[0095] The first determination unit 402 is configured to determine a target remote sensing orthographic image corresponding to the low-altitude video image from a remote sensing orthographic image set through a target image retrieval model, wherein the target remote sensing orthographic image has the same coverage area as the low-altitude video image.

[0096] The second determination unit 403 is configured to determine a matching coordinate pair between the low-altitude video image and the target remote sensing orthographic image according to the low-altitude video image and the target remote sensing orthographic image through a target view matching model.​

[0097] The third determining unit 404 is configured to determine the exterior orientation elements of the low-altitude video image according to the matching coordinate pair, and determine the geographic coordinates of the low-altitude video image according to the exterior orientation elements.

[0098] The image positioning apparatus provided by the embodiments of the present application comprises a first obtaining unit 401 configured to obtain a low-altitude video image to be positioned, wherein the low-altitude video image is any frame image in a target low-altitude video; a first determining unit 402 configured to determine a target remote sensing orthographic image corresponding to the low-altitude video image from a remote sensing orthographic image set through a target image retrieval model, wherein the target remote sensing orthographic image has the same coverage area as the low-altitude video image; a second determining unit 403 configured to determine a matching coordinate pair between the low-altitude video image and the target remote sensing orthographic image according to the low-altitude video image and the target remote sensing orthographic image through a target view angle matching model; and a third determining unit 404 configured to determine the exterior orientation elements of the low-altitude video image according to the matching coordinate pair, and determine the geographic coordinates of the low-altitude video image according to the exterior orientation elements.

[0099] Optionally, in the image positioning apparatus provided by the embodiments of the present application, the apparatus further comprises the following units for obtaining the target image retrieval model through the following steps: a second obtaining unit configured to obtain a first sample low-altitude video image set and a first sample remote sensing orthographic image set; a first processing unit configured to process each first sample low-altitude video image in the first sample low-altitude video image set and each first sample remote sensing orthographic image in the first sample remote sensing orthographic image set through an embedding processing module of an initial image retrieval model, to obtain a first sample low-altitude video image feature of each first sample low-altitude video image and a first sample remote sensing orthographic image feature of each first sample remote sensing orthographic image; a second processing unit configured to perform normalization processing on the first sample low-altitude video image feature through a spatial transformation module of the initial image retrieval model, to obtain a normalized first sample low-altitude video image feature; a third processing unit configured to process the normalized first sample low-altitude video image feature and each first sample remote sensing orthographic image feature through an image feature enhancement module of the initial image retrieval model, to obtain a processed image feature of each first sample low-altitude video image and a processed image feature of each first sample remote sensing orthographic image; and a fourth processing unit configured to determine a target loss function of the initial image retrieval model according to the processed image feature of each first sample low-altitude video image and the processed image feature of each first sample remote sensing orthographic image through a feature matching module of the initial image retrieval model, and train the initial image retrieval model based on the target loss function until a predetermined training target is reached, to obtain the target image retrieval model.

[0100] Optionally, in the image positioning apparatus provided by the embodiment of the present application, the third processing unit comprises: a first processing subunit, configured to process the normalized first sample low-altitude video image features and the first sample remote sensing orthographic image features respectively by using the image feature enhancement module, to obtain enhanced features of each first sample low-altitude video image and enhanced features of each first sample remote sensing orthographic image; and a second processing subunit, configured to perform information interaction processing on the enhanced features of each first sample low-altitude video image and the enhanced features of each first sample remote sensing orthographic image respectively based on a preset processing operator, to obtain processed image features of each first sample low-altitude video image and processed image features of each first sample remote sensing orthographic image.

[0101] Optionally, in the image positioning apparatus provided by the embodiment of the present application, the fourth processing unit comprises: a third processing subunit, configured to process the processed image features of each first sample low-altitude video image and the processed image features of each first sample remote sensing orthographic image respectively by using a feature matching module, to obtain global image features of each first sample low-altitude video image and global image features of each first sample remote sensing orthographic image; a fourth processing subunit, configured to, for each first sample low-altitude video image, perform similarity calculation on the global image features of the first sample low-altitude video image and the global image features of each first sample remote sensing orthographic image respectively, to obtain global image feature similarities between the first sample low-altitude video image and each first sample remote sensing orthographic image; a fifth processing subunit, configured to perform similarity calculation on the processed image features of the first sample low-altitude video image and the processed image features of each first sample remote sensing orthographic image respectively, to obtain local image feature similarities between the first sample low-altitude video image and each first sample remote sensing orthographic image; a sixth processing subunit, configured to perform summation calculation on the global image feature similarities and the local image feature similarities respectively, to obtain target image feature similarities between the first sample low-altitude video image and each first sample remote sensing orthographic image; and a seventh processing subunit, configured to determine a target loss function according to the global image feature similarities, the local image feature similarities and the target image feature similarities.

[0102] Optionally, in the image positioning apparatus provided by the embodiment of the present application, the apparatus further comprises the following units for obtaining the target view matching model by the following steps: a third obtaining unit is configured to obtain a training sample set, wherein the training sample set comprises paired second sample low-altitude video images and second sample remote sensing orthographic images; a fifth processing unit is configured to process the second sample low-altitude video images and the second sample remote sensing orthographic images respectively through the encoding network of the initial view matching model to obtain first encoding features and second encoding features; a sixth processing unit is configured to process the first encoding features and the second encoding features respectively through the attention processing module of the initial view matching model to obtain a first set of attention feature representations and a second set of attention feature representations; a seventh processing unit is configured to perform progressive matching based on the first set of attention feature representations and the second set of attention feature representations to obtain sample matching coordinate pairs between the second sample low-altitude video images and the second sample remote sensing orthographic images, wherein the progressive matching starts from the lowest resolution features in the first set of attention feature representations and the second set of attention feature representations and iterates to the highest resolution features gradually, and the matching coordinate pairs predicted each time are used as prior information for the next iteration; and an eighth processing unit is configured to train the initial view matching model based on the sample matching coordinate pairs until a predetermined training target is reached to obtain the target view matching model.

[0103] Optionally, in the image positioning apparatus provided by the embodiment of the present application, the third determining unit 404 comprises: a first determining subunit configured to map the matching coordinate pairs to the coordinate system of the digital surface model, and determine the three-dimensional space coordinates corresponding to each matching point in the matching coordinate pairs according to the ground elevation information of the digital surface model; and a second determining subunit configured to determine the exterior orientation elements of the low-altitude video images according to the three-dimensional space coordinates and a preset camera pose solving algorithm.

[0104] Optionally, in the image positioning apparatus provided by the embodiment of the present application, the third determining unit 404 further comprises: a third determining subunit configured to determine the depth information of the low-altitude video images according to the depth estimation network; and a conversion subunit configured to convert the pixel points of the low-altitude video images from the camera coordinate system to the world coordinate system according to the depth information and the exterior orientation elements to obtain the geographic coordinates of the low-altitude video images.

[0105] It should be noted that the first acquisition unit 401, the first determination unit 402, the second determination unit 403, and the third determination unit 404 mentioned above correspond to steps S201 to S204 in Embodiment 1. The four units and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules or units can be hardware or software components stored in memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The above units can also be part of a device and run in the computer terminal 10 provided in Embodiment 1.

[0106] Example 3

[0107] Embodiments of this application may provide an electronic device. Figure 5 This is a structural block diagram of an electronic device according to an embodiment of this application. Figure 5 As shown, the electronic device may include: one or more ( Figure 5 Only one of the components is shown: processor 502, memory 504, memory controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module, and display.

[0108] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and apparatus in the embodiments of this application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby implementing the above-described methods. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0109] The processor can access information and applications stored in memory via a transmission device to perform the following steps: acquiring a low-altitude video image to be located, wherein the low-altitude video image is any frame from the target low-altitude video; determining the target remote sensing orthophoto image corresponding to the low-altitude video image from a remote sensing orthophoto image set using a target image retrieval model, wherein the target remote sensing orthophoto image and the low-altitude video image have the same coverage area; determining the matching coordinate pair between the low-altitude video image and the target remote sensing orthophoto image using a target viewpoint matching model based on the low-altitude video image and the target remote sensing orthophoto image; determining the exterior orientation elements of the low-altitude video image based on the matching coordinate pair, and determining the geographic coordinates of the low-altitude video image based on the exterior orientation elements.

[0110] The processor can call information and application programs stored in the memory through the transmission device to perform the following steps: obtaining a first sample low-altitude video image set and a first sample remote sensing orthographic image set; processing each first sample low-altitude video image in the first sample low-altitude video image set and each first sample remote sensing orthographic image in the first sample remote sensing orthographic image set respectively through an embedding processing module of an initial image retrieval model to obtain a first sample low-altitude video image feature of each first sample low-altitude video image and a first sample remote sensing orthographic image feature of each first sample remote sensing orthographic image; performing normalization processing on the first sample low-altitude video image feature through a spatial transformation module of the initial image retrieval model to obtain a normalized first sample low-altitude video image feature; processing the normalized first sample low-altitude video image feature and each first sample remote sensing orthographic image feature respectively through an image feature enhancement module of the initial image retrieval model to obtain a processed image feature of each first sample low-altitude video image and a processed image feature of each first sample remote sensing orthographic image; determining a target loss function of the initial image retrieval model according to the processed image feature of each first sample low-altitude video image and the processed image feature of each first sample remote sensing orthographic image through a feature matching module of the initial image retrieval model, and training the initial image retrieval model based on the target loss function until a predetermined training target is reached to obtain a target image retrieval model.

[0111] The processor can call information and application programs stored in the memory through the transmission device to perform the following steps: processing the normalized first sample low-altitude video image feature and each first sample remote sensing orthographic image feature respectively through an image feature enhancement module to obtain an enhanced feature of each first sample low-altitude video image and an enhanced feature of each first sample remote sensing orthographic image; performing information interaction processing on the enhanced feature of each first sample low-altitude video image and the enhanced feature of each first sample remote sensing orthographic image respectively based on a preset processing operator to obtain a processed image feature of each first sample low-altitude video image and a processed image feature of each first sample remote sensing orthographic image.

[0112] The processor can call information and application programs stored in the memory through the transmission device to perform the following steps: processing the processed image features of each first sample low-altitude video image and the processed image features of each first sample remote sensing orthographic image respectively through the feature matching module to obtain global image features of each first sample low-altitude video image and global image features of each first sample remote sensing orthographic image; for each first sample low-altitude video image, the global image features of the first sample low-altitude video image and the global image features of each first sample remote sensing orthographic image are calculated respectively to obtain the global image feature similarity between the first sample low-altitude video image and each first sample remote sensing orthographic image; the processed image features of the first sample low-altitude video image and the processed image features of each first sample remote sensing orthographic image are calculated respectively to obtain the local image feature similarity between the first sample low-altitude video image and each first sample remote sensing orthographic image; the global image feature similarity and the local image feature similarity are summed respectively to obtain the target image feature similarity between the first sample low-altitude video image and each first sample remote sensing orthographic image; the target loss function is determined according to the global image feature similarity, the local image feature similarity and the target image feature similarity.

[0113] The processor can call information and application programs stored in the memory through the transmission device to perform the following steps: obtaining a training sample set, wherein the training sample set includes paired second sample low-altitude video images and second sample remote sensing orthographic images; processing the second sample low-altitude video images and the second sample remote sensing orthographic images respectively through the encoding network of the initial view angle matching model to obtain first encoding features and second encoding features; processing the first encoding features and the second encoding features respectively through the attention processing module of the initial view angle matching model to obtain a first set of attention feature representations and a second set of attention feature representations; performing progressive matching based on the first set of attention feature representations and the second set of attention feature representations to obtain a sample matching coordinate pair between the second sample low-altitude video images and the second sample remote sensing orthographic images, wherein the progressive matching starts from the lowest resolution features in the first set of attention feature representations and the second set of attention feature representations, and iterates gradually to the highest resolution features, and the matching coordinate pair predicted each time is used as prior information for the next iteration; training the initial view angle matching model based on the sample matching coordinate pair until a predetermined training target is reached to obtain a target view angle matching model.

[0114] The processor can call information and application programs stored in the memory through the transmission device to perform the following steps: mapping the matching coordinate pair to a coordinate system of the digital surface model, determining the three-dimensional space coordinates corresponding to each matching point in the matching coordinate pair according to the ground elevation information of the digital surface model; and determining the exterior orientation elements of the low-altitude video image according to the three-dimensional space coordinates and a preset camera pose solving algorithm.

[0115] The processor can call information and application programs stored in the memory through the transmission device to perform the following steps: determining the depth information of the low-altitude video image according to the depth estimation network; and converting the pixel points of the low-altitude video image from the camera coordinate system to the world coordinate system according to the depth information and the exterior orientation elements to obtain the geographic coordinates of the low-altitude video image.

[0116] Those skilled in the art can understand that, Figure 5 The structure shown is only schematic, and the electronic device can also be a smart phone, a tablet computer, a palm computer, a Mobile Internet Device (MID), a PAD, or the like. Figure 5 It does not limit the structure of the electronic device. For example, the electronic device can include more or fewer components (such as a network interface, a display device, etc.) than Figure 5 or have a different configuration than Figure 5 shown.

[0117] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by programs instructing the related hardware of the terminal device, and the programs can be stored in a computer readable storage medium, which can include a flash disk, a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk or an optical disk, etc.

[0118] Embodiment 4

[0119] The embodiments of the present application also provide a computer readable storage medium. Optionally, in the present embodiment, the above storage medium can be used to save the program code executed by the image positioning method provided in Embodiment 1.

[0120] Optionally, in the present embodiment, the above storage medium can be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.

[0121] The present application also provides a computer program product adapted to execute the steps of the image positioning method when executed on a data processing device.

[0122] The above-mentioned sequence numbers of the embodiments of the present application are only for description, and do not represent the advantages or disadvantages of the embodiments.

[0123] In the above-mentioned embodiments of the present application, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.

[0124] In the several embodiments of the present application, it should be understood that the disclosed technology can be implemented in other ways. Of course, the above-described device embodiments are only illustrative, and the division of units is only a logical function division. There can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, units or modules, and can be electrical or other forms.

[0125] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., they can be located in one place or distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiments.

[0126] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The above integrated unit can be realized in the form of hardware or in the form of software functional unit.

[0127] When the integrated unit is realized in the form of software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.

[0128] The above merely describes the preferred embodiments of the present application, and it should be pointed out that, for those skilled in the art, some improvements and refinements can be made without departing from the principles of the present application, and these improvements and refinements should also be considered as the protection scope of the present application.

Claims

1. An image positioning method characterized by, The method comprises the following steps: acquiring a low-altitude video image to be positioned, wherein the low-altitude video image is any frame image in a target low-altitude video; determining a target remote sensing orthographic image corresponding to the low-altitude video image from a remote sensing orthographic image set through a target image retrieval model, wherein the target remote sensing orthographic image has the same coverage area as the low-altitude video image; determining a matching coordinate pair between the low-altitude video image and the target remote sensing orthographic image through a target view angle matching model according to the low-altitude video image and the target remote sensing orthographic image; determining an exterior orientation element of the low-altitude video image according to the matching coordinate pair, and determining a geographic coordinate of the low-altitude video image according to the exterior orientation element.

2. The method of claim 1, wherein, The target image retrieval model is obtained through the following steps: acquiring a first sample low-altitude video image set and a first sample remote sensing orthographic image set; processing each first sample low-altitude video image in the first sample low-altitude video image set and each first sample remote sensing orthographic image in the first sample remote sensing orthographic image set through an embedding processing module of an initial image retrieval model to obtain a first sample low-altitude video image feature of each first sample low-altitude video image and a first sample remote sensing orthographic image feature of each first sample remote sensing orthographic image; normalizing the first sample low-altitude video image feature through a spatial transformation module of the initial image retrieval model to obtain a normalized first sample low-altitude video image feature; processing the normalized first sample low-altitude video image feature and each first sample remote sensing orthographic image feature through an image feature enhancement module of the initial image retrieval model to obtain a processed image feature of each first sample low-altitude video image and a processed image feature of each first sample remote sensing orthographic image; determining a target loss function of the initial image retrieval model according to the processed image feature of each first sample low-altitude video image and the processed image feature of each first sample remote sensing orthographic image through a feature matching module of the initial image retrieval model, and training the initial image retrieval model based on the target loss function until a predetermined training target is reached to obtain the target image retrieval model.

3. The method of claim 2, wherein, The processing of the normalized first sample low-altitude video image feature and each first sample remote sensing orthographic image feature through the image feature enhancement module of the initial image retrieval model to obtain the processed image feature of each first sample low-altitude video image and the processed image feature of each first sample remote sensing orthographic image comprises: processing the normalized first sample low-altitude video image feature and each first sample remote sensing orthographic image feature through the image feature enhancement module to obtain an enhanced feature of each first sample low-altitude video image and an enhanced feature of each first sample remote sensing orthographic image; The enhanced features of each first sample low-altitude video image and the enhanced features of each first sample remote sensing orthographic image are interactively processed based on preset processing operators to obtain processed image features of each first sample low-altitude video image and processed image features of each first sample remote sensing orthographic image.

4. The method of claim 2, wherein, The target loss function of the initial image retrieval model is determined by the feature matching module of the initial image retrieval model according to the processed image features of each first sample low-altitude video image and the processed image features of each first sample remote sensing orthographic image, and the target loss function of the initial image retrieval model includes: The processed image features of each first sample low-altitude video image and the processed image features of each first sample remote sensing orthographic image are processed by the feature matching module to obtain global image features of each first sample low-altitude video image and global image features of each first sample remote sensing orthographic image; For each first sample low-altitude video image, the global image features of the first sample low-altitude video image and the global image features of each first sample remote sensing orthographic image are calculated for similarity to obtain global image feature similarity between the first sample low-altitude video image and each first sample remote sensing orthographic image; The processed image features of the first sample low-altitude video image and the processed image features of each first sample remote sensing orthographic image are calculated for similarity to obtain local image feature similarity between the first sample low-altitude video image and each first sample remote sensing orthographic image; The global image feature similarity and the local image feature similarity are summed to obtain target image feature similarity between the first sample low-altitude video image and each first sample remote sensing orthographic image. The target loss function is determined according to the global image feature similarity, the local image feature similarity, and the target image feature similarity.

5. The method of claim 1, wherein, The target view angle matching model is obtained by the following steps: Obtain a training sample set, wherein the training sample set includes paired second sample low-altitude video images and second sample remote sensing orthographic images; The first encoding features and the second encoding features are obtained by processing the second sample low-altitude video images and the second sample remote sensing orthographic images through the encoding network of the initial view angle matching model; The first attention feature representation set and the second attention feature representation set are obtained by processing the first encoding features and the second encoding features through the attention processing module of the initial view angle matching model; progressively match the first set of attention feature representations and the second set of attention feature representations to obtain a sample matching coordinate pair between the second sample low-altitude video image and the second sample remote sensing orthographic image, wherein the progressive matching starts from the lowest resolution feature in the first set of attention feature representations and the second set of attention feature representations, and iterates to the highest resolution feature, and a matching coordinate pair predicted each time is used as prior information for the next iteration; train the initial view angle matching model based on the sample matching coordinate pair until a predetermined training target is reached to obtain the target view angle matching model.

6. The method of claim 1, wherein, determining, according to the matching coordinate pair, an exterior orientation element of the low-altitude video image includes: mapping the matching coordinate pair to a coordinate system of a digital surface model, and determining, according to ground elevation information of the digital surface model, a three-dimensional space coordinate corresponding to each matching point in the matching coordinate pair; determining, according to the three-dimensional space coordinate and a preset camera pose solving algorithm, the exterior orientation element of the low-altitude video image.

7. The method of claim 1, wherein, determining, according to the exterior orientation element, a geographic coordinate of the low-altitude video image includes: determining, according to a depth estimation network, depth information of the low-altitude video image; converting, according to the depth information and the exterior orientation element, a pixel point of the low-altitude video image from a camera coordinate system to a world coordinate system to obtain the geographic coordinate of the low-altitude video image.

8. An image positioning apparatus characterized by comprising: includes: a first obtaining unit configured to obtain a low-altitude video image to be positioned, wherein the low-altitude video image is any one frame of image in a target low-altitude video; a first determining unit configured to determine, by a target image retrieval model, a target remote sensing orthographic image corresponding to the low-altitude video image from a remote sensing orthographic image set, wherein the target remote sensing orthographic image has a same coverage area as the low-altitude video image; a second determining unit configured to determine, by a target view angle matching model, a matching coordinate pair between the low-altitude video image and the target remote sensing orthographic image according to the low-altitude video image and the target remote sensing orthographic image; a third determining unit configured to determine, according to the matching coordinate pair, an exterior orientation element of the low-altitude video image, and determine, according to the exterior orientation element, a geographic coordinate of the low-altitude video image.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium includes a stored executable program, wherein the executable program, when executed, controls a device in which the computer readable storage medium is located to perform the image positioning method of any one of claims 1 to 7.

10. An electronic device, comprising: includes: a memory storing an executable program; a processor configured to execute the program, wherein the program, when executed, performs the image positioning method of any one of claims 1 to 7.