Image information recognition method and apparatus, and computing device

By constructing a multimodal feature vector model of the target region and using the matching of images with regional viewpoints, the problem of geographic location recognition of image scenes without initial location information was solved, achieving efficient and accurate positioning.

WO2025222922A1PCT designated stage Publication Date: 2025-10-30HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/141849
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-26
Filing Date
2024-12-24
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Existing technologies cannot accurately identify the geographical location of a scene in an image without initial location information.

Method used

By constructing a multimodal feature vector model of the target region, and matching the multimodal feature vector of the image with the multimodal feature vector of each viewpoint in the target region, the geographical location of the scene in the image can be identified.

Benefits of technology

In the absence of prior information, it improves the accuracy and efficiency of image retrieval, and can accurately locate the scene in the image at its geographical location.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024141849_30102025_PF_FP_ABST
    Figure CN2024141849_30102025_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides an image information recognition method. The method comprises: receiving an image to be recognized input by a user; acquiring multi-modal feature vectors of said image; and on the basis of the multi-modal feature vectors of said image and multi-modal feature vectors of each viewpoint in a target region, obtaining at least one target viewpoint matching said image, and displaying the at least one target viewpoint to the user. The target region is a region where a scene of said image is located; the multi-modal feature vectors of said image include at least two of the following feature vectors: a scene depth feature vector, a semantic feature vector, and a contour feature vector of said image; and the target viewpoint is used for indicating the geographic position of the scene of said image in the target region. In the method, the geographic position of the scene in an image to be recognized can be recognized on the basis of an existing map of a target region corresponding to said image.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, apparatus and computing devices for image information recognition

[0001] This application claims priority to Chinese Patent Application No. 202410529217.5, filed on April 26, 2024, with the China National Intellectual Property Administration, entitled “Method, Apparatus and Computing Device for Image Information Recognition”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of image processing, and more specifically, to a method, apparatus, and computing device for image information recognition. Background Technology

[0003] Currently, identifying the spatial location (also known as geographic location) of a scene in an image is a common pursuit in many scenarios, including but not limited to: urban governance, disaster prevention and mitigation, internet applications, criminal investigation applications, etc.

[0004] One related image information recognition technology is web search technology, which involves uploading images to the internet and identifying landmarks, objects, people, clothing, etc., within the captured images to find similar images or similar products within images. However, this technology cannot pinpoint the geographical location of the scene within the image.

[0005] Another related image information recognition technology is augmented reality (AR) visual positioning technology. This AR visual positioning technology requires obtaining the initial location information of the image and then identifying the geographical location of the scene in the image based on the initial location information. When identifying the geographical location of a scene in a captured image, this technology relies not only on existing maps of the target region (e.g., the target city) but also on the initial location information of the image. However, most images encountered in daily life or work lack initial location information; therefore, this technology has certain limitations in identifying the geographical location of scenes in captured images.

[0006] Therefore, how to identify the geographical location of a scene in an image based on an existing map of the target region where the image is located, without any other prior information, has become an urgent technical problem to be solved. Summary of the Invention

[0007] This application provides a method for image information recognition, which can identify the geographical location of a scene in an image based on an existing map of the target region where the image to be identified is located, without any other prior information.

[0008] Firstly, a method for image information recognition is provided. This method includes: receiving an image to be recognized input by a user; obtaining a multimodal feature vector of the image to be recognized; and, based on the multimodal feature vector of the image to be recognized and the multimodal feature vector of each viewpoint in the target region, obtaining and displaying to the user at least one target viewpoint matching the image to be recognized. The target region is the region where the scene in the image to be recognized is located. The multimodal feature vector of the image to be recognized includes at least two of the following feature vectors: a scene depth feature vector, a semantic feature vector, and a contour feature vector of the image to be recognized. The target viewpoint is used to indicate the geographical location of the scene in the image to be recognized within the target region.

[0009] The target regions mentioned above can include, but are not limited to: rural areas, counties, cities, provinces, countries, and even the world.

[0010] In the above technical solution, by using the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region, at least one target viewpoint matching the image to be identified is determined, thereby identifying the geographical location of the scene in the image to be identified in the target region. In this way, without any other prior information, the geographical location of the scene in the image to be identified can be identified based on the existing map of the target region where the image to be identified is located, thereby effectively improving the accuracy and efficiency of image retrieval and avoiding some limitations of existing methods for identifying the geographical location of images.

[0011] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: constructing a model of the target region; obtaining each viewpoint of the target region from the model of the target region; obtaining and storing the multimodal feature vector of each viewpoint, wherein the multimodal feature vector of each viewpoint includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of each viewpoint.

[0012] In the above technical solution, a model of the target region can be established, thereby obtaining viewpoints on the model of the target region, and performing panoramic visual rendering on each available viewpoint to extract features such as depth, semantics, and texture. A unified feature expression for features such as depth, semantics, and texture can be established to form a multimodal feature vector. In this way, the map features of the target region can be effectively extracted, eliminating the dependence on the initial location information, so that the geographical location of the image to be identified in the target region can be determined in the absence of other prior knowledge.

[0013] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: receiving information about the target region input by a user, the information about the target region being used to indicate the target retrieval range of the target region, and the scene in the image to be identified being located within the target retrieval range; obtaining at least one target viewpoint matching the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region includes: obtaining at least one target viewpoint matching the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target retrieval range.

[0014] In the above technical solution, users can also input or import a search range, which is used to indicate a search range in the target region. In this way, the geographical location of the image to be identified can be determined from the search range, avoiding the need to directly search from the entire target region to determine the geographical location of the image to be identified, thus improving the efficiency of retrieval.

[0015] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: encoding each viewpoint based on the multimodal feature vector of each viewpoint in the target region to obtain the encoding value of each viewpoint, wherein the encoding value of each viewpoint is used to determine at least one target viewpoint that matches the image to be identified.

[0016] In conjunction with the first aspect, in some implementations of the first aspect, obtaining at least one target viewpoint matching the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region includes: obtaining multiple recommended target viewpoints matching the image to be identified based on the multimodal feature vector of the image to be identified and the encoding value of each viewpoint; and selecting at least one target viewpoint matching the image to be identified from the multiple recommended target viewpoints based on the multimodal feature vector of the image to be identified and the multimodal feature vectors of the multiple recommended target viewpoints.

[0017] In the above technical solution, multiple recommended target viewpoints that match the image to be identified can be quickly obtained based on the encoding value of each viewpoint. Then, based on the multimodal feature vector of the image to be identified and the multimodal feature vectors of the multiple recommended target viewpoints, at least one target viewpoint that matches the image to be identified can be selected from the multiple recommended target viewpoints. In this way, the efficiency of image retrieval can be further improved.

[0018] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: displaying the plurality of recommended target viewpoints to the user; receiving a set of viewpoints input by the user, the set of viewpoints including a plurality of viewpoints selected by the user from the plurality of recommended target viewpoints; and filtering at least one target viewpoint matching the image to be identified from the plurality of recommended target viewpoints based on the multimodal feature vector of the image to be identified and the multimodal feature vectors of the plurality of recommended target viewpoints, including: filtering at least one target viewpoint matching the image to be identified from the set of viewpoints based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the set of viewpoints.

[0019] In the above technical solution, multiple recommended target viewpoints can be displayed to the user, who can then select a set of viewpoints from the multiple recommended target viewpoints. Based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the user-selected viewpoint set, at least one target viewpoint that matches the image to be identified can be filtered from the viewpoint set. In this way, at least one target viewpoint that matches the image to be identified can be obtained quickly, thereby further improving the efficiency and accuracy of image retrieval.

[0020] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: obtaining information about a building in the image to be identified, the information about the building including at least one of the following: the name of the building, an identification ID; and displaying the information about the building in the image to be identified to the user when displaying the at least one target viewpoint to the user.

[0021] The above technical solution can also acquire and display information about buildings in the image to be identified to the user, so as to be applied to various scenarios where buildings in the image to be identified need to be identified.

[0022] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: when displaying the at least one target viewpoint to the user, displaying to the user the degree of matching between each target viewpoint and the image to be identified, the degree of matching being used to indicate the similarity between the scene within the field of view of each target viewpoint and the scene of the image to be identified.

[0023] In the above technical solution, the matching degree between each target viewpoint and the image to be identified can also be displayed to the user. This allows the user to quickly determine the geographical location of the scene in the image to be identified based on the similarity between the scene in the field of view of each target viewpoint and the scene in the image to be identified.

[0024] In conjunction with the first aspect, in some implementations of the first aspect, the viewpoint is an available viewpoint adaptively obtained from the feasible ground area and low-altitude area on the model of the target region.

[0025] In the above technical solution, available viewpoints can be adaptively obtained from feasible ground areas and low-altitude areas on the map of the target region, avoiding obtaining all viewpoints on the map of the target region without filtering, which can further improve the efficiency of retrieval.

[0026] Secondly, a method for constructing a corresponding feature index library for a target region is provided. The method includes: obtaining each viewpoint in the target region, obtaining the multimodal feature vector of each viewpoint in the target region, and storing the multimodal feature vector of each viewpoint in the feature index library. The multimodal feature vector of each viewpoint includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of each viewpoint.

[0027] In conjunction with the second aspect, in some implementations of the second aspect, the method further includes: fusing at least one feature vector of each available viewpoint to obtain a multimodal feature index, and storing it in a feature index library.

[0028] In conjunction with the second aspect, in some implementations of the second aspect, the method further includes: constructing a model of the target region; obtaining each viewpoint of the target region includes: obtaining each viewpoint of the target region from the model of the target region.

[0029] In conjunction with the second aspect, in some implementations of the second aspect, the method includes: receiving an image to be identified input by a user; obtaining a multimodal feature vector of the image to be identified; and, based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region, obtaining and displaying to the user at least one target viewpoint matching the image to be identified. Here, the target region is the region where the scene in the image to be identified is located, and the multimodal feature vector of the image to be identified includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of the image to be identified. The target viewpoint is used to indicate the geographical location of the scene in the image to be identified within the target region.

[0030] The target regions mentioned above can include, but are not limited to: rural areas, counties, cities, provinces, countries, and even the world.

[0031] In conjunction with the second aspect, in some implementations of the second aspect, the method further includes: constructing a model of the target region; obtaining each viewpoint of the target region from the model of the target region; obtaining and storing the multimodal feature vector of each viewpoint, wherein the multimodal feature vector of each viewpoint includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of each viewpoint.

[0032] In conjunction with the second aspect, in some implementations of the second aspect, the method further includes: receiving information about the target region input by a user, the information about the target region being used to indicate the target retrieval range of the target region, and the scene in the image to be identified being located within the target retrieval range; obtaining at least one target viewpoint matching the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region includes: obtaining at least one target viewpoint matching the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target retrieval range.

[0033] In conjunction with the second aspect, in some implementations of the second aspect, the method further includes: encoding each viewpoint based on the multimodal feature vector of each viewpoint in the target region to obtain the encoding value of each viewpoint, wherein the encoding value of each viewpoint is used to determine at least one target viewpoint that matches the image to be identified.

[0034] In conjunction with the second aspect, in some implementations of the second aspect, obtaining at least one target viewpoint matching the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region includes: obtaining multiple recommended target viewpoints matching the image to be identified based on the multimodal feature vector of the image to be identified and the encoding value of each viewpoint; and selecting at least one target viewpoint matching the image to be identified from the multiple recommended target viewpoints based on the multimodal feature vector of the image to be identified and the multimodal feature vectors of the multiple recommended target viewpoints.

[0035] In conjunction with the second aspect, in some implementations of the second aspect, the method further includes: displaying the plurality of recommended target viewpoints to the user; receiving a set of viewpoints input by the user, the set of viewpoints including a plurality of viewpoints selected by the user from the plurality of recommended target viewpoints; and filtering at least one target viewpoint matching the image to be identified from the plurality of recommended target viewpoints based on the multimodal feature vector of the image to be identified and the multimodal feature vectors of the plurality of recommended target viewpoints, including: filtering at least one target viewpoint matching the image to be identified from the set of viewpoints based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the set of viewpoints.

[0036] In conjunction with the second aspect, in some implementations of the second aspect, the method further includes: obtaining information about a building in the image to be identified, the information about the building including at least one of the following: the name of the building, an identification ID; and displaying the information about the building in the image to be identified to the user when displaying the at least one target viewpoint to the user.

[0037] In conjunction with the second aspect, in some implementations of the second aspect, the method further includes: when displaying the at least one target viewpoint to the user, displaying to the user the degree of matching between each target viewpoint and the image to be identified, the degree of matching being used to indicate the similarity between the scene within the field of view of each target viewpoint and the scene of the image to be identified.

[0038] In conjunction with the second aspect, in some implementations of the second aspect, the viewpoint is an available viewpoint adaptively obtained from the feasible ground area and low-altitude area on the model of the target region.

[0039] It should be understood that for the beneficial effects of the second aspect and its various implementations, please refer to the beneficial effects of the first aspect and its various implementations; they will not be elaborated here.

[0040] Thirdly, an image information recognition apparatus is provided, comprising: a receiving module, a processing module, and a display module. The receiving module receives an image to be recognized input by a user; the processing module acquires the multimodal feature vector of the image to be recognized, and obtains at least one target viewpoint matching the image to be recognized based on the multimodal feature vector of the image to be recognized and the multimodal feature vector of each viewpoint in the target region; the display module displays the at least one target viewpoint matching the image to be recognized to the user. The target region is the region where the scene in the image to be recognized is located, and the multimodal feature vector of the image to be recognized includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of the image to be recognized. The target viewpoint indicates the geographical location of the scene in the image to be recognized within the target region.

[0041] In conjunction with the third aspect, in some implementations of the third aspect, the processing module is also used to: construct a model of the target region; obtain each viewpoint of the target region from the model of the target region; obtain and store the multimodal feature vector of each viewpoint, wherein the multimodal feature vector of each viewpoint includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of each viewpoint.

[0042] In conjunction with the third aspect, in some implementations of the third aspect, the receiving module is further configured to receive information about the target region input by the user, the information about the target region being used to indicate the target retrieval range of the target region, and the scene in the image to be identified being located within the target retrieval range; the processing module is specifically configured to: obtain at least one target viewpoint that matches the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target retrieval range.

[0043] In conjunction with the third aspect, in some implementations of the third aspect, the processing module is further configured to: encode each viewpoint based on the multimodal feature vector of each viewpoint in the target region to obtain the encoded value of each viewpoint, wherein the encoded value of each viewpoint is used to determine at least one target viewpoint that matches the image to be identified.

[0044] In conjunction with the third aspect, in some implementations of the third aspect, the processing module is specifically used to: obtain multiple recommended target viewpoints that match the image to be identified based on the multimodal feature vector of the image to be identified and the encoding value of each viewpoint; and select at least one target viewpoint that matches the image to be identified from the multiple recommended target viewpoints based on the multimodal feature vector of the image to be identified and the multimodal feature vectors of the multiple recommended target viewpoints.

[0045] In conjunction with the third aspect, in some implementations of the third aspect, the display module is further used to display the multiple recommended target viewpoints to the user; the receiving module is further used to receive the viewpoint set input by the user, the viewpoint set including multiple viewpoints selected by the user from the multiple recommended target viewpoints; the processing module is specifically used to: based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the viewpoint set, filter at least one target viewpoint that matches the image to be identified from the viewpoint set.

[0046] In conjunction with the third aspect, in some implementations of the third aspect, the processing module is further configured to obtain information about buildings in the image to be identified, the information about buildings including at least one of the following: the name of the building, identification ID; the display module is further configured to display the information about buildings in the image to be identified to the user when displaying the at least one target viewpoint to the user.

[0047] In conjunction with the third aspect, in some implementations of the third aspect, the display module is further configured to, when displaying the at least one target viewpoint to the user, display to the user the degree of matching between each target viewpoint and the image to be identified, the degree of matching being used to indicate the similarity between the scene within the field of view of each target viewpoint and the scene of the image to be identified.

[0048] In conjunction with the third aspect, in some implementations of the third aspect, the viewpoint is an available viewpoint adaptively obtained from the feasible ground area and low-altitude area on the model of the target region.

[0049] It should be understood that the beneficial effects of the third aspect and its various implementations are similar to those of the first aspect and its various implementations, and will not be elaborated upon here.

[0050] Fourthly, an apparatus for constructing a corresponding feature index library for a target region is provided. The apparatus includes a processing module and a storage module. The processing module is used to acquire each viewpoint in the target region and acquire the multimodal feature vector of each viewpoint in the target region. The storage module is used to store the multimodal feature vector of each viewpoint in the feature index library. The multimodal feature vector of each viewpoint includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of each viewpoint.

[0051] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the processing module is also used to fuse at least two feature vectors for each viewpoint to obtain a multimodal feature index for each viewpoint, and the storage module is also used to store the multimodal feature index for each viewpoint in a feature index library.

[0052] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the processing module is specifically used to: construct a model of the target region; and obtain each viewpoint of the target region from the model of the target region.

[0053] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the device further includes: a processing module and a storage module, wherein the processing module is used to acquire each viewpoint in the target region and acquire the multimodal feature vector of each viewpoint in the target region; the storage module is used to store the multimodal feature vector of each viewpoint in the target region in a feature index library, wherein the multimodal feature vector of each viewpoint includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of each viewpoint.

[0054] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the processing module is also used to: construct a model of the target region; obtain each viewpoint of the target region from the model of the target region; obtain and store the multimodal feature vector of each viewpoint, wherein the multimodal feature vector of each viewpoint includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of each viewpoint.

[0055] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the device further includes: a receiving module and a display module, wherein the receiving module is used to receive an image to be identified input by a user; the processing module is further used to acquire the multimodal feature vector of the image to be identified, and based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region, obtain at least one target viewpoint matching the image to be identified; the display module is used to display to the user the at least one target viewpoint matching the image to be identified. The target region is the region where the scene in the image to be identified is located, the multimodal feature vector of the image to be identified includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of the image to be identified, and the target viewpoint is used to indicate the geographical location of the scene in the image to be identified in the target region.

[0056] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the receiving module is further configured to receive information about the target region input by the user, the information about the target region being used to indicate the target retrieval range of the target region, and the scene in the image to be identified being located within the target retrieval range; the processing module is specifically configured to: obtain at least one target viewpoint that matches the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target retrieval range.

[0057] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the processing module is further configured to: encode each viewpoint based on the multimodal feature vector of each viewpoint in the target region to obtain the encoded value of each viewpoint, wherein the encoded value of each viewpoint is used to determine at least one target viewpoint that matches the image to be identified.

[0058] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the processing module is specifically used to: obtain multiple recommended target viewpoints that match the image to be identified based on the multimodal feature vector of the image to be identified and the encoding value of each viewpoint; and select at least one target viewpoint that matches the image to be identified from the multiple recommended target viewpoints based on the multimodal feature vector of the image to be identified and the multimodal feature vectors of the multiple recommended target viewpoints.

[0059] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the display module is further used to display the multiple recommended target viewpoints to the user; the receiving module is further used to receive the viewpoint set input by the user, the viewpoint set including multiple viewpoints selected by the user from the multiple recommended target viewpoints; the processing module is specifically used to: based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the viewpoint set, filter at least one target viewpoint that matches the image to be identified from the viewpoint set.

[0060] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the processing module is further configured to obtain information about buildings in the image to be identified, the information about buildings including at least one of the following: the name of the building, identification ID; the display module is further configured to display the information about buildings in the image to be identified to the user when displaying the at least one target viewpoint to the user.

[0061] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the display module is further configured to, when displaying the at least one target viewpoint to the user, display to the user the degree of matching between each target viewpoint and the image to be identified, the degree of matching being used to indicate the similarity between the scene within the field of view of each target viewpoint and the scene of the image to be identified.

[0062] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the viewpoint is an available viewpoint adaptively obtained from the feasible ground area and low-altitude area on the model of the target region.

[0063] It should be understood that the beneficial effects of the fourth aspect and its various implementations are similar to those of the first aspect and its various implementations, and will not be elaborated upon here.

[0064] Fifthly, a computing device is provided, including a processor and a memory, and optionally, an input / output interface. The processor controls the input / output interface to send and receive information, the memory stores a computer program, and the processor retrieves and runs the computer program from the memory, causing the computing device to execute the method of the first aspect or any possible implementation thereof, or to execute the method of the second aspect or any possible implementation thereof.

[0065] Optionally, the processor can be a general-purpose processor, which can be implemented in hardware or software. When implemented in hardware, the processor can be a logic circuit, integrated circuit, etc.; when implemented in software, the processor can be a general-purpose processor that reads software code stored in memory. This memory can be integrated into the processor or located outside the processor and exist independently.

[0066] In a sixth aspect, a computing device cluster is provided, including at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, such that the computing device cluster performs a method in the first aspect or any possible implementation thereof, or performs a method in the second aspect or any possible implementation thereof.

[0067] In a seventh aspect, a chip is provided that acquires and executes instructions to implement the methods in the first aspect and any implementation thereof.

[0068] Optionally, as one implementation, the chip includes a processor and a data interface, through which the processor reads instructions stored in the memory and executes the methods in the first aspect and any implementation thereof.

[0069] Optionally, as one implementation, the chip may further include a memory storing instructions, and the processor is used to execute the instructions stored in the memory. When the instructions are executed, the processor is used to execute the method in the first aspect and any implementation thereof, or to execute the method in the second aspect or any possible implementation thereof.

[0070] Eighthly, a computer program product comprising instructions is provided, which, when executed by a computing device, causes the computing device to perform the method as described in the first aspect and any implementation thereof, or to perform the method as described in the second aspect or any possible implementation thereof.

[0071] In a ninth aspect, a computer program product containing instructions is provided, which, when executed by a cluster of computing devices, cause the cluster of computing devices to perform the methods described in the first aspect and any implementation thereof, or to perform the methods described in the second aspect or any possible implementation thereof.

[0072] In a tenth aspect, a computer-readable storage medium is provided, including computer program instructions that, when executed by a computing device, perform a method as described in the first aspect and any implementation thereof, or perform a method as described in the second aspect or any possible implementation thereof.

[0073] As examples, these computer-readable storage devices include, but are not limited to, one or more of the following: read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), flash memory, electrically EPROM (EEPROM), and hard drive.

[0074] Alternatively, as one implementation method, the aforementioned storage medium can specifically be a non-volatile storage medium.

[0075] Eleventhly, a computer-readable storage medium is provided, including computer program instructions that, when executed by a cluster of computing devices, perform the method as described in the first aspect and any implementation thereof, or perform the method as described in the second aspect or any possible implementation thereof.

[0076] As examples, these computer-readable storage devices include, but are not limited to, one or more of the following: read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), flash memory, electrically EPROM (EEPROM), and hard drive.

[0077] Alternatively, as one implementation method, the aforementioned storage medium can specifically be a non-volatile storage medium. Attached Figure Description

[0078] Figure 1 is a schematic block diagram of a system architecture applied in this application.

[0079] Figure 2 is a schematic flowchart of a method for constructing a feature index library corresponding to a target region according to an embodiment of this application.

[0080] Figure 3 is a schematic diagram of a 3D white model of a target region provided in an embodiment of this application.

[0081] Figure 4 is a schematic diagram of at least one feature corresponding to a viewpoint provided in an embodiment of this application.

[0082] Figure 5 is a schematic diagram of an embodiment of this application that uses a neural network to extract at least one feature vector for each available viewpoint.

[0083] Figure 6 is a schematic flowchart of an image information recognition method provided in an embodiment of this application.

[0084] Figure 7 is a schematic diagram of an image to be recognized input by a user in an embodiment of this application.

[0085] Figure 8 is a schematic diagram of another image to be recognized input by the user in an embodiment of this application.

[0086] Figure 9 is a schematic diagram of at least one feature extracted from the image to be identified shown in Figure 7, according to an embodiment of this application.

[0087] Figure 10 is a schematic diagram of the output for the image to be identified shown in Figure 7, provided by an embodiment of this application.

[0088] Figure 11 is a schematic diagram of the output for the image to be identified shown in Figure 8, provided by an embodiment of this application.

[0089] Figure 12 is a schematic block diagram of an image information recognition device 1200 provided in an embodiment of this application.

[0090] Figure 13 is a schematic block diagram of an apparatus 1300 for constructing a feature index library corresponding to a target region, according to an embodiment of this application.

[0091] Figure 14 is a schematic diagram of the architecture of a computing device 1500 provided in an embodiment of this application.

[0092] Figure 15 is a schematic diagram of the architecture of a computing device cluster provided in an embodiment of this application.

[0093] Figure 16 is a schematic diagram of the connection between computing devices 1500A and 1500B via a network provided in an embodiment of this application. Detailed Implementation

[0094] The technical solutions in this application will now be described with reference to the accompanying drawings.

[0095] This application will present various aspects, embodiments, or features relating to systems comprising multiple devices, components, modules, etc. It should be understood and appreciated that individual systems may include additional devices, components, modules, etc., and / or may not include all devices, components, modules, etc. discussed in conjunction with the accompanying drawings. Furthermore, combinations of these approaches are also possible.

[0096] Furthermore, in the embodiments of this application, the words "exemplary," "for example," etc., are used to indicate that they are examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" in this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of the term "exemplary" is intended to present the concept in a concrete manner.

[0097] In the embodiments of this application, "corresponding" and "corresponding" can sometimes be used interchangeably. It should be noted that when the distinction is not emphasized, their intended meanings are consistent.

[0098] The business scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0099] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0100] In this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A alone, A and B simultaneously, and B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0101] Currently, identifying the spatial location (also known as geographic location) of a scene in an image is a common need, given the lack of or limited prior information. Several possible application scenarios are introduced below.

[0102] 1. Urban governance scenarios:

[0103] For example, based on photos taken by patrol officers / grid workers using their mobile phones, detailed information about the patrol targets in the photos can be confirmed, such as the geographical location of the patrol targets.

[0104] 2. Disaster prevention and mitigation scenarios:

[0105] For example, it can quickly locate suspected fire spots captured by surveillance cameras, find the geographical location of the suspected fire spot in the city, and initiate rapid rescue.

[0106] 3. Scenarios for Internet applications:

[0107] For example, for photos that interest you in your WeChat Moments, you can search for information such as the location where the photo was taken and the scene.

[0108] 4. Scenarios for criminal investigation applications:

[0109] For example, a photograph or a surveillance video can be used to identify the location of a crime scene, thereby strengthening the chain of evidence.

[0110] One related image information recognition technology is web search technology. This involves uploading images to the internet and identifying landmarks, objects, people, clothing, etc., within the captured images to find similar images or similar products within those images. However, this technology cannot pinpoint the geographical location of the scene within the image.

[0111] Another related image information recognition technology is augmented reality (AR) visual positioning technology. This technology obtains the initial location information of an image through hardware sensors (e.g., GPS, IMU). Based on this initial location information, the search range is narrowed. A small search area is obtained from an existing map of the target region (e.g., the target city), and the geographical location of the scene in the image is identified within this small search area. This technology relies on both the existing map of the target region and the initial location information of the image when identifying the geographical location of a scene in a captured image. However, most images encountered in daily life or work lack initial location information; therefore, this technology has certain limitations in identifying the geographical location of scenes in captured images.

[0112] In view of this, embodiments of this application provide a method for image information recognition, which can identify the geographical location of a scene in an image based on an existing map of the target region where the image is located, without any other prior information.

[0113] For ease of description, a system architecture of this application will be described in detail below with reference to Figure 1.

[0114] As an example, Figure 1 is a schematic block diagram of a system architecture according to this application. As shown in Figure 1, the system may include two parts: the first part is used to construct a feature index library corresponding to the target region (also known as the region to be searched), and the second part is used to obtain the geographical location of the scene in the user-input image (an image taken in the target region, also known as an image to be identified) from the feature index library. The functions of the modules involved in the above two parts are described in detail below.

[0115] 1. First part: Used to build a feature index library corresponding to the target region.

[0116] As an example, this first part includes a map acquisition module, a map construction module, and a feature index library. The map acquisition module is used to acquire map information of the target region using at least one method; the map construction module is used to construct a map of the target region based on the map information acquired using at least one method; and the feature index library is used to store feature vectors of various viewpoints in the map of the target region. The specific implementation process of this part will be described in detail below with reference to Figure 2, and will not be elaborated upon here.

[0117] 2. Second part: Used to obtain the geographical location of the scene in the image to be identified from the feature index library.

[0118] As an example, this second part includes a visualization interface, an image feature extraction module, an image location calculation / recommendation module, and an application platform. The visualization interface is used for user input of the image to be identified and outputs the geographical location information of the scene within the image. The image feature extraction module extracts the feature vector of the user-input image. The image location calculation / recommendation module compares the feature vector of the image to be identified with feature vectors stored in a feature index library to obtain and display the possible geographical location information of the scene within the image. The application platform includes, but is not limited to: a government affairs integrated query platform, a disaster automatic identification platform, an emergency response command platform, a public security clue acquisition platform, a consumer application platform, and an internet application platform. The specific implementation process of this part will be described in detail below with reference to Figure 5; it will not be elaborated here.

[0119] Figure 2 is a schematic flowchart of a method for constructing a feature index library corresponding to a target region according to an embodiment of this application. As shown in Figure 2, the method includes steps 210-250, which will be described in detail below.

[0120] Step 210: Collect map information of the target region.

[0121] As an example, embodiments of this application can collect map information of a target area. There are various methods of collection, and this application does not specifically limit them. For example, map information of a target area can be obtained through satellite images collected by satellites. Alternatively, map information of a target area can be obtained through images taken by aircraft. Furthermore, map information of a target area can be obtained through panoramic streetscapes collected by sensors installed on vehicles.

[0122] Step 220: Construct a 3D model of the target region based on the map information of the target region.

[0123] As an example, let's consider obtaining map information of a target area using satellite imagery. As an example, a 3D white model of the target area can be obtained from satellite imagery. For instance, data related to buildings, ground surfaces, roads, and waterways in the target area can be extracted from the satellite imagery. Image processing techniques and algorithms can then be used to identify and extract feature points, edges, and textures from the satellite imagery. Then, using professional 3D modeling software, the extracted feature points and edges can be converted into geometric shapes and vertices in three-dimensional space, thereby constructing a 3D white model of the target area. During the modeling process, factors such as lighting, materials, and textures can also be considered to make the model more realistic and lifelike.

[0124] For example, Figure 3 is a schematic diagram of a 3D white model of the target region constructed based on the above method.

[0125] It should be understood that the aforementioned 3D white model of the target area is an uncolored model widely used in computer graphics. It typically uses geometric shapes, vertices, wireframes, etc., to represent objects in the real world (e.g., terrain information in a region), and is a highly abstract and simplified model. Through preprocessing, the 3D white model can be semantically divided into different categories such as buildings, ground, roads, and water systems, and the outlines of features in different categories can be further extracted.

[0126] It should also be understood that the 3D white model of the aforementioned target area can also be manipulated (such as rotated, scaled, and moved), rendered, etc. on a computer.

[0127] Step 230: Obtain available viewpoints from the 3D model of the target region.

[0128] For example, one can obtain all viewpoints from a 3D model of a target region, and then filter out usable viewpoints from that set. The specific implementation processes for obtaining all viewpoints and usable viewpoints are described in detail below.

[0129] The aforementioned available viewpoints can also be understood as viewpoints in passable areas. That is, viewpoints in passable areas are selected from all viewpoints and used as available viewpoints.

[0130] It should be noted that, in the embodiments of this application, the viewpoint corresponding to a certain position in the three-dimensional model of the target region can be described by six degrees of freedom (6DOF). 6DOF includes the translational degrees of freedom along the three rectangular coordinate axes x, y, and z, and the rotational degrees of freedom around the three coordinate axes x, y, and z.

[0131] The following details how to obtain all viewpoints from a 3D model of the target region.

[0132] In one possible implementation, the spatial extent of the three-dimensional model of the target region can be divided into multiple grids according to a fixed size, and multiple viewpoints can be obtained based on the multiple grids according to sampling rules. These multiple viewpoints are also the full set of viewpoints mentioned above.

[0133] The following details how to filter out usable viewpoints from the full set of viewpoints.

[0134] In one possible implementation, for each viewpoint in the full set of viewpoints, the three-dimensional coordinates x, y, z of the spatial location of each viewpoint can be obtained by vertical ray tracing, as well as the semantic type of the scene within the field of view of each viewpoint. The semantic type may include, but is not limited to: road, ground, building, water system, etc.

[0135] Step 240: Generate a multimodal feature index for each available viewpoint and store the multimodal feature index of each available viewpoint in the feature index library.

[0136] As an example, at least one feature can be extracted for each available viewpoint. As shown in Figure 4, this at least one feature may include, but is not limited to, depth features, semantic features, and contour features of the scene within the viewpoint's field of view.

[0137] In one possible implementation, for each available viewpoint, based on the principle of panoramic spherical projection, the depth, semantics, contour and other features of the scene within the field of view of each available viewpoint are extracted within a certain range in the horizontal and vertical directions.

[0138] In this embodiment of the application, after obtaining at least one feature of each available viewpoint, at least one feature vector of each available viewpoint can be extracted, such as scene depth feature vector, semantic feature vector, contour feature vector, and at least one feature vector of each available viewpoint can be fused to obtain a multimodal feature index, which is then stored in the feature index library.

[0139] In one possible implementation, as shown in Figure 5, a neural network can be used to extract at least one feature vector for each available viewpoint, thereby obtaining a multimodal feature vector for each viewpoint. An adaptive spatial density distribution multimodal feature index is then constructed for the multimodal feature vectors of each viewpoint and stored in a feature index library.

[0140] Step 250: Encode each available viewpoint and store the encoded value of each available viewpoint in the feature index library.

[0141] As an example, in this embodiment of the application, hash encoding can be performed based on the 6DOF (e.g., x / y value) and multimodal feature index of each available viewpoint to obtain the encoded value of each available viewpoint, and the encoded value of each available viewpoint can be stored in the feature index library.

[0142] It should be understood that the encoded value of each available viewpoint can reflect the multimodal features of the scene within the field of view of each available viewpoint, such as depth, semantics, and contour.

[0143] Optionally, the blocks can be divided into squares of a certain length, and the viewpoints in each block can be assigned to a feature extraction task. Multiple feature extraction tasks can form a task list.

[0144] In the above technical solution, available viewpoints are adaptively acquired in the feasible ground area and low-altitude area on the map of the target region. Panoramic visual rendering is performed on each available viewpoint to extract features such as depth, semantics, and texture. A unified feature expression for features such as depth, semantics, and texture is established to form a multimodal feature vector and store it in the feature index library. In this way, map features of the target region can be effectively extracted, eliminating the dependence on the initial location information, so that the geographical location of the image to be identified in the target region can be determined in the absence of other prior knowledge.

[0145] Figure 6 is a schematic flowchart of an image information recognition method provided in an embodiment of this application. As shown in Figure 6, the method includes steps 610-640, which will be described in detail below.

[0146] Step 610: The user imports the image to be recognized.

[0147] For example, a user can input or import an image to be identified, which can also be referred to as the image to be retrieved.

[0148] Optionally, the user can also input information about the region where the image to be identified is located, which can be referred to as the region to be searched or the target region.

[0149] The geographical information of the image to be identified may include, but is not limited to: the region where the image was taken, and the target search range of the target region. Specifically, the scene in the image to be identified must fall within the target search range of the target region.

[0150] The aforementioned target retrieval range is used to indicate a search range within the target region. This allows the geographical location of the image to be identified to be determined from the search range, avoiding the need to directly search the entire target region to determine the geographical location of the image to be identified, thus improving retrieval efficiency.

[0151] For example, the files in the search scope mentioned above can be in geojson format, supporting multiple polygons.

[0152] Example 1, taking a command and inspection scenario as an example, shows an image to be identified as shown in Figure 7, captured by an inspector / grid member. The inspector / grid member needs to confirm the information of the inspection object in the image (e.g., the name or ID of the building in the upper left corner of the image). This allows for the construction of an urban inspection and governance application based on spatial search and positioning services, linked with a city information modeling (CIM) platform. Through digital empowerment, the efficiency and precision of business processing are improved. The inspector / grid member can input the image to be identified as shown in Figure 7. If the inspector / grid member knows that the image was taken in Xi'an, the target region entered by the user is Xi'an.

[0153] Example 2, taking a disaster monitoring scenario as an example, shows a frame from a Changsha city surveillance video in Figure 8. However, due to inadequate urban data governance, the camera installation location information may be lost, inaccurate, or unreliable, making it impossible to identify suspected fire points in the image and thus failing to meet disaster prevention and mitigation needs. Therefore, it is necessary to determine the information of suspected fire points in the image (e.g., the name or ID of the building where the suspected fire point is located, or which unit and floor of the building the suspected fire point is located in), in order to carry out disaster prevention and mitigation work. After obtaining the image to be identified shown in Figure 8, urban management staff can input the image shown in Figure 8. If the urban management staff knows that the image to be identified was taken in Changsha, the target region entered by the user is Changsha.

[0154] Step 620: Extract multimodal feature vectors from the image to be identified.

[0155] For example, at least one feature can be extracted from the image to be identified input by the user. The at least two features of the image to be identified may include, but are not limited to, scene depth features, semantic features, and contour features of the image to be identified.

[0156] For example, taking the image to be identified as shown in Figure 7, the scene depth features, semantic features, and contour features extracted from the image to be identified are shown in Figure 9.

[0157] In this embodiment of the application, after obtaining at least two features of the image to be identified, a multimodal feature vector of the image to be identified can also be extracted, such as the scene depth feature vector, semantic feature vector, and contour feature vector of the image to be identified.

[0158] In one possible implementation, based on the principle of panoramic spherical projection, the scene depth, semantics, contour and other features of the image to be identified can be extracted in the horizontal and vertical directions within a certain range. Then, a neural network is used to extract at least two feature vectors for each available viewpoint to obtain the multimodal feature vector of the image to be identified.

[0159] Step 630: Based on the multimodal feature vector of the image to be identified, search for target viewpoints in the target region that match the image to be identified from the feature index library.

[0160] In this embodiment of the application, after obtaining the multimodal feature vector of the image to be identified, the target viewpoint that matches the image to be identified can be found in the feature index library corresponding to the target region.

[0161] There are multiple ways to find viewpoints in the target region that match the image to be identified from the feature index library. This application does not make any specific limitations on this. Several possible implementation methods are introduced below.

[0162] In one possible implementation, the multimodal feature vector of the image to be identified is compared with the multimodal feature indexes of each viewpoint stored in the feature index library to find the target viewpoint that matches the image to be identified.

[0163] In another possible implementation, based on the multimodal feature vector of the image to be identified and the encoded values ​​of each viewpoint stored in the feature index library, multiple recommended target viewpoints that match the multimodal feature vector of the image to be identified are determined. Then, the multimodal feature vector of the image to be identified is compared with the multimodal feature indices of each recommended target viewpoint stored in the feature index library, and a target viewpoint matching the image to be identified is found from these multiple recommended target viewpoints.

[0164] In another possible implementation, multiple recommended target viewpoints can be displayed to the user, who then selects a set of viewpoints from these recommendations. This set includes the viewpoints chosen by the user from the multiple recommended viewpoints. Then, based on the multimodal feature vectors of the image to be identified and the multimodal feature vectors of the multiple recommended viewpoints, at least one target viewpoint that matches the image to be identified is selected from the multiple recommended viewpoints. This further improves the efficiency of image retrieval.

[0165] Optionally, in some embodiments, if the target viewpoint includes multiple target viewpoints, the matching degree between each target viewpoint and the image to be identified can also be obtained. The matching degree is used to indicate the similarity between the scene in the field of view of each target viewpoint and the scene of the image to be identified.

[0166] Optionally, in some embodiments, if the image to be identified includes buildings, the name or ID of the buildings included in the image to be identified can be determined based on the data of the 3D model corresponding to the target region and the 6DOF of the target viewpoint.

[0167] Step 640: Show the user the target viewpoint.

[0168] In this embodiment of the application, after obtaining the viewpoint that matches the image to be identified, the target viewpoint can be displayed to the user. For example, the target viewpoint can be output to the user through a visual interface.

[0169] It should be understood that since the viewpoint corresponding to a certain location in the 3D model of the target region is represented by 6DOF, the translational degrees of freedom along the three Cartesian coordinate axes (x, y, z) of the target viewpoint matched with the image to be identified are used to indicate the shooting point of the image to be identified, and the rotational degrees of freedom around the three coordinate axes (x, y, z) are used to indicate the shooting direction of the image to be identified. Therefore, the user can determine the geographical location of the scene in the image to be identified in the target region by using the target viewpoint.

[0170] Optionally, in some embodiments, if the target viewpoint includes multiple target viewpoints, the matching degree between each target viewpoint and the image to be identified can be displayed to the user when showing the user at least one target viewpoint. For example, the matching degree between each target viewpoint and the image to be identified can be output to the user through a visual interface. This allows the user to determine the geographical location of the scene in the image to be identified within the target region based on the matching degree between each target viewpoint and the image to be identified.

[0171] In some embodiments, if the image to be identified includes buildings, the name or ID of the buildings included in the image to be identified may be displayed to the user when the at least one target viewpoint is shown to the user, for example, by outputting the name or ID of the buildings included in the image to be identified to the user through a visual interface.

[0172] Example 1: Taking the image to be recognized input by the user as shown in Figure 7 as an example, the output of this embodiment through the visual interface is shown in Figure 10. For example, it can output the building in the upper left corner of the image to be recognized, which is located on ×× Road, ×× District, Xi'an City, with the ID 2926-038743.

[0173] Example 2, taking the image to be identified input by the user as shown in Figure 8 as an example, the output of this embodiment through the visual interface is shown in Figure 11. For example, it can output that the building suspected of being the ignition point of the fire in the image to be identified is the building with ID 038496 located on XX Road, XX District, Changsha City.

[0174] In the above technical solution, by using the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region, at least one target viewpoint matching the image to be identified is determined, thereby identifying the geographical location of the scene in the image to be identified in the target region. In this way, without any other prior information, the geographical location of the scene in the image to be identified can be identified based on the existing map of the target region where the image to be identified is located, thereby effectively improving the accuracy and efficiency of image retrieval and avoiding some limitations of existing methods for identifying the geographical location of images.

[0175] The methods provided by the embodiments of this application have been described in detail above with reference to Figures 1 to 11. The embodiments of the apparatus of this application will be described in detail below with reference to Figures 12 to 16. It should be understood that the descriptions of the method embodiments correspond to the descriptions of the apparatus embodiments; therefore, any parts not described in detail can be referred to the preceding method embodiments.

[0176] Figure 12 is a schematic block diagram of an image information recognition device 1200 provided in an embodiment of this application. The device 1200 can be implemented by software, hardware, or a combination of both. The device 1200 provided in this embodiment can implement the method flow shown in Figure 6 of this embodiment. The device 1200 includes: a receiving module 1210, a processing module 1220, and a display module 1230. The receiving module 1210 is used to receive an image to be recognized input by a user; the processing module 1220 is used to obtain the multimodal feature vector of the image to be recognized, and based on the multimodal feature vector of the image to be recognized and the multimodal feature vector of each viewpoint in the target region, obtain at least one target viewpoint matching the image to be recognized; the display module 1230 is used to display the at least one target viewpoint matching the image to be recognized to the user. The target region is the region where the scene in the image to be identified is located. The multimodal feature vector of the image to be identified includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of the image to be identified. The target viewpoint is used to indicate the geographical location of the scene in the image to be identified in the target region.

[0177] Optionally, the processing module 1220 is further configured to: construct a model of the target region; obtain each viewpoint of the target region from the model of the target region; obtain and store the multimodal feature vector of each viewpoint, wherein the multimodal feature vector of each viewpoint includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of each viewpoint.

[0178] Optionally, the receiving module 1210 is further configured to receive information about the target region input by the user, the information about the target region being used to indicate the target retrieval range of the target region, and the scene in the image to be identified being located within the target retrieval range; the processing module 1220 is specifically configured to: obtain at least one target viewpoint that matches the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target retrieval range.

[0179] Optionally, the processing module 1220 is further configured to: encode each viewpoint based on the multimodal feature vector of each viewpoint in the target region to obtain the encoded value of each viewpoint, wherein the encoded value of each viewpoint is used to determine at least one target viewpoint that matches the image to be identified.

[0180] Optionally, the processing module 1220 is specifically used to: obtain multiple recommended target viewpoints that match the image to be identified based on the multimodal feature vector of the image to be identified and the encoding value of each viewpoint; and select at least one target viewpoint that matches the image to be identified from the multiple recommended target viewpoints based on the multimodal feature vector of the image to be identified and the multimodal feature vectors of the multiple recommended target viewpoints.

[0181] Optionally, the display module 1230 is further configured to display the plurality of recommended target viewpoints to the user; the receiving module is further configured to receive a set of viewpoints input by the user, the set of viewpoints including a plurality of viewpoints selected by the user from the plurality of recommended target viewpoints; the processing module 1220 is specifically configured to: based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the set of viewpoints, filter at least one target viewpoint that matches the image to be identified from the set of viewpoints.

[0182] Optionally, the processing module 1220 is further configured to obtain information about the building in the image to be identified, the information of the building including at least one of the following: the name of the building, identification ID; the display module 1230 is further configured to display the information about the building in the image to be identified to the user when displaying the at least one target viewpoint to the user.

[0183] Optionally, the display module 1230 is further configured to, when displaying the at least one target viewpoint to the user, display to the user the degree of matching between each target viewpoint and the image to be identified, the degree of matching being used to indicate the similarity between the scene within the field of view of each target viewpoint and the scene of the image to be identified.

[0184] Optionally, the viewpoint can be an available viewpoint adaptively obtained from the feasible ground area and low-altitude area on the model of the target region.

[0185] Figure 13 is a schematic block diagram of an apparatus 1300 for constructing a feature index library corresponding to a target region according to an embodiment of this application. The apparatus 1300 can be implemented by software, hardware, or a combination of both. The apparatus 1300 provided in this embodiment can implement the method flow shown in Figure 2 of this embodiment. The apparatus 1300 includes: a processing module 1310 and a storage module 1320. The processing module 1310 is used to acquire each viewpoint in the target region and acquire the multimodal feature vector of each viewpoint in the target region. The storage module 1320 is used to store the multimodal feature vector of each viewpoint in the target region in the feature index library. The multimodal feature vector of each viewpoint includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of each viewpoint.

[0186] Optionally, the processing module 1310 is further configured to fuse at least two feature vectors of each viewpoint to obtain a multimodal feature index of each viewpoint, and the storage module 1320 is further configured to store the multimodal feature index of each viewpoint in a feature index library.

[0187] Optionally, the processing module 1310 is specifically used to: construct a model of the target region; and obtain each viewpoint of the target region from the model of the target region.

[0188] Optionally, the device 1300 further includes: a receiving module and a display module, wherein the receiving module is used to receive an image to be identified input by a user; the processing module 1310 is further used to obtain the multimodal feature vector of the image to be identified, and based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region, obtain at least one target viewpoint that matches the image to be identified; the display module is used to display the at least one target viewpoint that matches the image to be identified to the user. The target region is the region where the scene in the image to be identified is located, and the multimodal feature vector of the image to be identified includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of the image to be identified. The target viewpoint is used to indicate the geographical location of the scene in the image to be identified within the target region.

[0189] Optionally, the receiving module is further configured to receive information about the target region input by the user, the information about the target region being used to indicate the target retrieval range of the target region, and the scene in the image to be identified being located within the target retrieval range; the processing module 1310 is specifically configured to: obtain at least one target viewpoint that matches the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target retrieval range.

[0190] Optionally, the processing module 1310 is further configured to: encode each viewpoint based on the multimodal feature vector of each viewpoint in the target region to obtain the encoded value of each viewpoint, wherein the encoded value of each viewpoint is used to determine at least one target viewpoint that matches the image to be identified.

[0191] Optionally, the processing module 1310 is specifically configured to: obtain multiple recommended target viewpoints that match the image to be identified based on the multimodal feature vector of the image to be identified and the encoding value of each viewpoint; and select at least one target viewpoint that matches the image to be identified from the multiple recommended target viewpoints based on the multimodal feature vector of the image to be identified and the multimodal feature vectors of the multiple recommended target viewpoints.

[0192] Optionally, the display module is further configured to display the plurality of recommended target viewpoints to the user; the receiving module is further configured to receive the viewpoint set input by the user, the viewpoint set including the plurality of viewpoints selected by the user from the plurality of recommended target viewpoints; the processing module 1310 is specifically configured to: based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the viewpoint set, filter at least one target viewpoint that matches the image to be identified from the viewpoint set.

[0193] Optionally, the processing module 1310 is further configured to obtain information about buildings in the image to be identified, the information about buildings including at least one of the following: the name of the building, identification ID; the display module is further configured to display the information about buildings in the image to be identified to the user when displaying the at least one target viewpoint to the user.

[0194] Optionally, the display module is further configured to, when displaying the at least one target viewpoint to the user, show the user the degree of matching between each target viewpoint and the image to be identified, wherein the degree of matching is used to indicate the similarity between the scene within the field of view of each target viewpoint and the scene of the image to be identified.

[0195] Optionally, the viewpoint can be an available viewpoint adaptively obtained from the feasible ground area and low-altitude area on the model of the target region.

[0196] The device 1200 or device 1300 here may be embodied in the form of a functional module. The term "module" here may be implemented in software and / or hardware, without specific limitation.

[0197] For example, a "module" can be a software program, a hardware circuit, or a combination of both that implements the above functions. For instance, the implementation of the receiving module 1210 will be described below using device 1200 as an example. Similarly, the implementation of other modules, such as processing module 1220 and demonstration module 1230, can refer to the implementation of the receiving module 1210.

[0198] As an example of a software functional unit, the receiving module 1210 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Further, the aforementioned computing instance may be one or more. For example, the receiving module 1210 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.

[0199] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.

[0200] As an example of a hardware functional unit, the receiving module 1210 may include at least one computing device, such as a server. Alternatively, the receiving module 1210 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.

[0201] The multiple computing devices included in the receiving module 1210 can be distributed in the same region or in different regions. Similarly, the multiple computing devices included in the receiving module 1210 can be distributed in the same Availability Zone (AZ) or in different AZs. Likewise, the multiple computing devices included in the receiving module 1210 can be distributed in the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0202] Therefore, the modules of the various examples described in the embodiments of this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0203] It should be noted that the above embodiments of the device, when executing the above methods, are only illustrative examples of the division of functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. For example, the receiving module 1210 can be used to execute any step in the above methods, the processing module 1220 can be used to execute any step in the above methods, and the display module 1230 can be used to execute any step in the above methods. The steps implemented by the receiving module 1210, the processing module 1220, and the display module 1230 can be specified as needed. By implementing different steps in the above methods through the receiving module 1210, the processing module 1220, and the display module 1230, all the functions of the above device can be realized.

[0204] Furthermore, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments above, which will not be repeated here.

[0205] The method provided in this application can be executed by a computing device, which can also be referred to as a computer system. It includes a hardware layer, an operating system layer running on top of the hardware layer, and an application layer running on the operating system layer. The hardware layer includes hardware such as processing units, memory, and memory control units; the functions and structure of this hardware will be described in detail later. The operating system can be any one or more computer operating systems that implement business processing through processes, such as Linux, Unix, Android, iOS, or Windows. The application layer includes applications such as browsers, address books, word processing software, and instant messaging software. Optionally, the computer system can be a handheld device such as a smartphone, or a terminal device such as a personal computer; this application does not particularly limit this, as long as the method provided in this application can be used. The executing entity of the method provided in this application can be a computing device, or a functional module within the computing device capable of calling and executing programs.

[0206] The following describes in detail, with reference to Figure 14, a computing device provided in an embodiment of this application.

[0207] Figure 14 is a schematic diagram of the architecture of a computing device 1500 provided in an embodiment of this application. The computing device 1500 may be a server, a computer, or other device with computing capabilities. The computing device 1500 shown in Figure 14 includes at least one processor 1510 and a memory 1520.

[0208] It should be understood that this application does not limit the number of processors and memories in the computing device 1500.

[0209] The processor 1510 executes instructions in the memory 1520, causing the computing device 1500 to implement the method provided in this application. Alternatively, the processor 1510 executes instructions in the memory 1520, causing the computing device 1500 to implement the various functional modules provided in this application, thereby implementing the method provided in this application.

[0210] Optionally, the computing device 1500 also includes a communication interface 1530. The communication interface 1530 uses a transceiver module, such as, but not limited to, a network interface card or a transceiver, to enable communication between the computing device 1500 and other devices or communication networks.

[0211] Optionally, the computing device 1500 also includes a system bus 1540, wherein the processor 1510, memory 1520, and communication interface 1530 are respectively connected to the system bus 1540. The processor 1510 can access the memory 1520 through the system bus 1540; for example, the processor 1510 can perform data read / write or code execution in the memory 1520 through the system bus 1540. The system bus 1540 is a peripheral component interconnect express (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The system bus 1540 is divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used in Figure 14, but this does not mean that there is only one bus or one type of bus.

[0212] In one possible implementation, the processor 1510 primarily functions to interpret the instructions (or code) of a computer program and process data within the computer software. The instructions of the computer program and the data within the computer software can be stored in memory 1520 or cache 1516.

[0213] Optionally, processor 1510 may be an integrated circuit chip with signal processing capabilities. By way of example and not limitation, processor 1510 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. Among these, a general-purpose processor is a microprocessor, etc. For example, processor 1510 may be a central processing unit (CPU).

[0214] Optionally, each processor 1510 includes at least one processing unit 1512 and a memory control unit 1514.

[0215] Optionally, the processing unit 1512, also known as the core, is the most important component of the processor. The processing unit 1512 is manufactured from single-crystal silicon using a specific production process. All calculations, command reception, command storage, and data processing are performed by the core. Each processing unit independently executes program instructions, utilizing parallel computing capabilities to accelerate program execution. Various processing units have fixed logical structures; for example, a processing unit includes logical units such as a Level 1 cache, a Level 2 cache, an execution unit, an instruction-level unit, and a bus interface.

[0216] In one implementation example, the memory control unit 1514 controls the data interaction between the memory 1520 and the processing unit 1512. Specifically, the memory control unit 1514 receives memory access requests from the processing unit 1512 and controls access to memory based on the memory access requests. By way of example and not limitation, the memory control unit is a device such as a memory management unit (MMU).

[0217] In one implementation example, each memory control unit 1514 addresses the memory 1520 via the system bus. An arbitrator (not shown in Figure 14) is configured on the system bus to handle and coordinate contention for access by the multiple processing units 1512.

[0218] In one implementation example, the processing unit 1512 and the memory control unit 1514 are connected via internal chip connection lines, such as address lines, thereby enabling communication between the processing unit 1512 and the memory control unit 1514.

[0219] Optionally, each processor 1510 also includes a cache 1516, which is a buffer for data exchange (called a cache). When the processing unit 1512 needs to read data, it first looks for the required data in the cache. If the data is found, it is executed directly; otherwise, it looks for the data in memory. Since the cache operates much faster than memory, its purpose is to help the processing unit 1512 run faster.

[0220] The memory 1520 provides runtime space for processes in the computing device 1500. For example, the memory 1520 stores the computer program (specifically, the program code) used to generate the process. After the computer program is run by the processor to generate a process, the processor allocates corresponding storage space for the process in the memory 1520. Furthermore, the aforementioned storage space further includes text segments, initialized data segments, bit initialized data segments, stack segments, heap segments, etc. The memory 1520 stores data generated during the process's execution, such as intermediate data or process data, in the aforementioned process-specific storage space.

[0221] Optionally, the memory, also known as RAM, is used to temporarily store the data processed by the processor 1510, as well as data exchanged with external storage devices such as hard disks. As long as the computer is running, the processor 1510 will load the data that needs to be processed into RAM for processing, and after the processing is completed, the processing unit 1512 will send the result out.

[0222] By way of example and not limitation, memory 1520 is volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory is read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory is random access memory (RAM) used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory 1520 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0223] The structure of the computing device 1500 listed above is merely illustrative and is not limited thereto. The computing device 1500 in this application includes various hardware components in existing computer systems. For example, the computing device 1500 also includes other memories besides memory 1520, such as disk storage. Those skilled in the art should understand that the computing device 1500 may also include other devices necessary for normal operation. Furthermore, depending on specific needs, those skilled in the art should understand that the computing device 1500 may also include hardware devices for implementing other additional functions. In addition, those skilled in the art should understand that the computing device 1500 may only include the devices necessary for implementing the embodiments of this application, and not necessarily all the devices shown in FIG. 14.

[0224] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device may be a server. In some embodiments, the computing device may also be a desktop computer, laptop computer, or smartphone, or other terminal device.

[0225] As shown in Figure 15, the computing device cluster includes at least one computing device 1500. The memory 1520 of one or more computing devices 1500 in the computing device cluster may store the same instructions for performing the above-described methods.

[0226] In some possible implementations, the memory 1520 of one or more computing devices 1500 in the computing device cluster may also each store a portion of the instructions for executing the above-described methods. In other words, a combination of one or more computing devices 1500 can jointly execute the instructions of the above-described methods.

[0227] It should be noted that the memory 1520 in different computing devices 1500 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the aforementioned device. That is, the instructions stored in the memory 1520 of different computing devices 1500 can implement the functions of one or more modules within the aforementioned device.

[0228] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 16 illustrates one possible implementation. As shown in Figure 16, two computing devices, 1500A and 1500B, are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device.

[0229] It should be understood that the functions of computing device 1500A shown in Figure 16 can also be performed by multiple computing devices 1500. Similarly, the functions of computing device 1500B can also be performed by multiple computing devices 1500.

[0230] In this embodiment, a computer program product containing instructions is also provided. The computer program product may be a software or program product containing instructions capable of running on a computing device or stored on any usable medium. When run on a computing device, it causes the computing device to perform the methods provided above, or causes the computing device to perform the functions of the apparatus provided above.

[0231] In this embodiment, a computer-readable storage medium is also provided. This computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that, when executed on a computing device, cause the computing device to perform the method described above.

[0232] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0233] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0234] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0235] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0236] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0237] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0238] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0239] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for image information recognition, characterized in that, The method includes: Receives the image to be recognized from the user input; The multimodal feature vector of the image to be identified is obtained, and the multimodal feature vector of the image to be identified includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of the image to be identified; Based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region, at least one target viewpoint matching the image to be identified is obtained. The target region is the region where the scene in the image to be identified is located, and the target viewpoint is used to indicate the geographical location of the scene in the image to be identified in the target region. The user is shown at least one target viewpoint.

2. The method according to claim 1, characterized in that, The method further includes: Construct a model of the target region; Obtain each viewpoint of the target region from the model of the target region, and each viewpoint is used to indicate each location in the target region; Obtain the multimodal feature vector of each viewpoint, wherein the multimodal feature vector of each viewpoint includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of each viewpoint.

3. The method according to claim 1 or 2, characterized in that, The method further includes: The system receives information about the target region input by the user, which is used to indicate the target retrieval range of the target region. The scene in the image to be identified is located within the target retrieval range of the target region. The step of obtaining at least one target viewpoint matching the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region includes: Based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target retrieval range, at least one target viewpoint that matches the image to be identified is obtained.

4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Based on the multimodal feature vector of each viewpoint in the target region, the encoding value of each viewpoint is obtained, and the encoding value of each viewpoint is used to determine at least one target viewpoint that matches the image to be identified.

5. The method according to claim 4, characterized in that, The step of obtaining at least one target viewpoint matching the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region includes: Based on the multimodal feature vector of the image to be identified and the encoding value of each viewpoint, multiple recommended target viewpoints matching the image to be identified are obtained; Based on the multimodal feature vector of the image to be identified and the multimodal feature vectors of the plurality of recommended target viewpoints, at least one target viewpoint that matches the image to be identified is selected from the plurality of recommended target viewpoints.

6. The method according to claim 5, characterized in that, The method further includes: Show the user the multiple recommended target viewpoints; Receive the set of viewpoints input by the user, the set of viewpoints including multiple viewpoints selected by the user from the multiple recommended target viewpoints; The step of selecting at least one target viewpoint matching the image to be identified from the multiple recommended target viewpoints based on the multimodal feature vector of the image to be identified and the multimodal feature vectors of the multiple recommended target viewpoints includes: Based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the viewpoint set, at least one target viewpoint that matches the image to be identified is selected from the viewpoint set.

7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: Obtain information about buildings in the image to be identified, wherein the building information includes at least one of the following: the name of the building, and its identification ID; When displaying the at least one target viewpoint to the user, information about the buildings in the image to be identified is displayed to the user.

8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: When displaying the at least one target viewpoint to the user, the user is shown the degree of matching between each target viewpoint and the image to be identified, the degree of matching indicating the similarity between the scene within the field of view of each target viewpoint and the scene of the image to be identified.

9. An image information recognition device, characterized in that, The device includes: The receiving module is used to receive the image to be recognized input by the user; The processing module is used to obtain the multimodal feature vector of the image to be identified, wherein the multimodal feature vector of the image to be identified includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of the image to be identified; The processing module is further configured to obtain at least one target viewpoint that matches the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region, wherein the target region is the region where the scene in the image to be identified is located, and the target viewpoint is used to indicate the geographical location of the scene in the image to be identified in the target region. The display module is used to display the at least one target viewpoint to the user.

10. The apparatus according to claim 9, characterized in that, The processing module is also used for: Construct a model of the target region; Obtain each viewpoint of the target region from the model of the target region, and each viewpoint is used to indicate each location in the target region; Obtain the multimodal feature vector of each viewpoint, wherein the multimodal feature vector of each viewpoint includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of each viewpoint.

11. The apparatus according to claim 9 or 10, characterized in that, The receiving module is further configured to receive information about the target region input by the user, the information about the target region being used to indicate the target retrieval range of the target region, and the scene in the image to be identified being located within the target retrieval range of the target region; The processing module is specifically used for: Based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target retrieval range, at least one target viewpoint that matches the image to be identified is obtained.

12. The apparatus according to any one of claims 9 to 11, characterized in that, The processing module is further configured to obtain the encoding value of each viewpoint based on the multimodal feature vector of each viewpoint in the target region, and the encoding value of each viewpoint is used to determine at least one target viewpoint that matches the image to be identified.

13. The apparatus according to claim 12, characterized in that, The processing module is specifically used for: Based on the multimodal feature vector of the image to be identified and the encoding value of each viewpoint, multiple recommended target viewpoints matching the image to be identified are obtained; Based on the multimodal feature vector of the image to be identified and the multimodal feature vectors of the plurality of recommended target viewpoints, at least one target viewpoint that matches the image to be identified is selected from the plurality of recommended target viewpoints.

14. The apparatus according to claim 13, characterized in that, The display module is also used to display the multiple recommended target viewpoints to the user; The receiving module is further configured to receive the set of viewpoints input by the user, wherein the set of viewpoints includes multiple viewpoints selected by the user from the multiple recommended target viewpoints; The processing module is specifically used for: Based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the viewpoint set, at least one target viewpoint that matches the image to be identified is selected from the viewpoint set.

15. The apparatus according to any one of claims 9 to 14, characterized in that, The processing module is further configured to obtain information about buildings in the image to be identified, wherein the building information includes at least one of the following: the name of the building and its identifier ID; The display module is also used to display information about buildings in the image to be identified to the user when displaying the at least one target viewpoint to the user.

16. The apparatus according to any one of claims 9 to 15, characterized in that, The display module is further configured to, when displaying the at least one target viewpoint to the user, display the degree of matching between each target viewpoint and the image to be identified, wherein the degree of matching is used to indicate the similarity between the scene within the field of view of each target viewpoint and the scene of the image to be identified.

17. A computing device cluster, characterized in that, The system includes at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1 to 8.

18. A computer program product containing instructions, characterized in that, When the instruction is executed by the computing device cluster, the computing device cluster causes the computing device cluster to perform the method as described in any one of claims 1 to 8.

19. A computer-readable storage medium, characterized in that, It includes computer program instructions, which, when executed by a cluster of computing devices, perform the method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Indoor positioning method, server and system

    CN104936283A

  • Visual positioning method, terminal and server

    CN112348886A

  • Scene recognition method and device

    CN115049909A

  • Visual positioning method, medium and electronic equipment

    CN117710451A