Method and apparatus for image information recognition, and computing device
The method employs multimodal feature vectors to identify the geographical location of a scene in an image using existing maps, addressing the limitations of existing technologies by enhancing retrieval accuracy and efficiency without requiring initial location information.
Patent Information
- Authority / Receiving Office
- HK · HK
- Patent Type
- Applications
- Current Assignee / Owner
- HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
- Filing Date
- 2026-04-26
- Publication Date
- 2026-07-17
AI Technical Summary
Existing image recognition technologies struggle to accurately identify the geographical location of a scene in an image without relying on initial location information, limiting their effectiveness in various applications.
A method and apparatus that utilize multimodal feature vectors, including scene depth, semantic, and contour vectors, to determine the geographical location of a scene in an image based on an existing map of the target region, enabling identification without prior information.
Enhances the accuracy and efficiency of image retrieval by determining geographical location using existing maps, eliminating the need for initial location information and improving search precision and speed.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
(19) State Intellectual Property Office (12) Invention Patent Application (10) Application Publication Number (43) Application Publication Date (21) Application Number 202410529217.5 (22) Application Date 2024.04.26 (71) Applicant Huawei Cloud Computing Technology Co., Ltd. Address 550025, Guizhou Province, Guiyang City, Gui'an New District, Qianzhong Avenue, Xinggong Road, Huawei Cloud Data Center (72) Inventors Kang Yifei, Chen Gang, Liu Liu (74) Patent Agency Beijing Longshuang Lida Intellectual Property Agency Co., Ltd. 11329 Patent Attorney Zhou Qiao, Wang Jun (51) Int.Cl. G06V 20 / 00 (2022.01) G06V 10 / 44 (2022.01) G06V 30 / 41 (2022.01) G06F 16 / 9537 (2019.01) G06F 16 / 29 (2019.01) (54) Invention Title: Method, Apparatus, and Computing Device for Image Information Recognition (57) Abstract: This application provides a method for image information recognition, comprising: receiving an image to be recognized input by a user; obtaining a multimodal feature vector of the image to be recognized; and obtaining and displaying to the user at least one target viewpoint matching the image to be recognized based on the multimodal feature vector of the image to be recognized and the multimodal feature vector of each viewpoint in the target region. The target region is the region where the scene in the image to be recognized is located. The multimodal feature vector of the image to be recognized includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of the image to be recognized. The target viewpoint is used to indicate the geographical location of the scene in the image to be recognized within the target region. This method can identify the geographical location of the scene in the image to be recognized based on an existing map of the target region where the image to be recognized is located. Claims 3 pages, Description 22 pages, Drawings 9 pages, CN 120852972 A 2025.10.28 CN 1 20 85 29 72 A 1. A method for image information recognition, characterized in that the method comprises: receiving an image to be recognized input by a user; acquiring a multimodal feature vector of the image to be recognized, the multimodal feature vector of the image to be recognized including at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of the image to be recognized; obtaining at least one target viewpoint matching the image to be recognized based on the multimodal feature vector of the image to be recognized and the multimodal feature vector of each viewpoint in the target region, the target region being the region where the scene in the image to be recognized is located, the target viewpoint being used to indicate the geographical location of the scene in the image to be recognized in the target region; and displaying the at least one target viewpoint to the user.2. The method according to claim 1, characterized in that the method further comprises: constructing a model of the target region; obtaining each viewpoint of the target region from the model of the target region, the viewpoint being used to indicate each location in the target region; obtaining a multimodal feature vector of each viewpoint, the multimodal feature vector of each viewpoint including at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of each viewpoint. 3. The method according to claim 1 or 2, characterized in that the method further comprises: receiving information of the target region input by the user, the information of the target region being used to indicate the target retrieval range of the target region, the scene in the image to be identified being located within the target retrieval range of the target region; obtaining at least one target viewpoint matching the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region, comprising: obtaining at least one target viewpoint matching the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target retrieval range. 4. The method according to any one of claims 1 to 3, characterized in that the method further comprises: obtaining an encoding value for each viewpoint based on the multimodal feature vector of each viewpoint in the target region, wherein the encoding value of each viewpoint is used to determine at least one target viewpoint matching the image to be identified. 5. The method according to claim 4, characterized in that obtaining at least one target viewpoint matching the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region comprises: obtaining a plurality of recommended target viewpoints matching the image to be identified based on the multimodal feature vector of the image to be identified and the encoding value of each viewpoint; and selecting at least one target viewpoint matching the image to be identified from the plurality of recommended target viewpoints based on the multimodal feature vector of the image to be identified and the multimodal feature vectors of the plurality of recommended target viewpoints. 6. The method according to claim 5, characterized in that the method further comprises: displaying the plurality of recommended target viewpoints to the user; receiving a set of viewpoints input by the user, the set of viewpoints including a plurality of viewpoints selected by the user from the plurality of recommended target viewpoints; the step of filtering at least one target viewpoint matching the image to be identified from the plurality of recommended target viewpoints based on the multimodal feature vector of the image to be identified and the multimodal feature vector of the plurality of recommended target viewpoints, comprising:Based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the viewpoint set, at least one target viewpoint matching the image to be identified is selected from the viewpoint set. 7. The method according to any one of claims 1 to 6, characterized in that the method further comprises: obtaining information about buildings in the image to be identified, the information about buildings including at least one of the following: the name of the building, an identification ID; and, when displaying the at least one target viewpoint to the user, displaying the information about buildings in the image to be identified to the user. 8. The method according to any one of claims 1 to 7, characterized in that the method further comprises: when displaying the at least one target viewpoint to the user, displaying to the user the matching degree between each target viewpoint and the image to be identified, the matching degree indicating the similarity between the scene within the field of view of each target viewpoint and the scene of the image to be identified. 9. An image information recognition apparatus, characterized in that the apparatus comprises: a receiving module for receiving an image to be recognized input by a user; a processing module for acquiring a multimodal feature vector of the image to be recognized, wherein the multimodal feature vector of the image to be recognized includes at least two of the following feature vectors: a scene depth feature vector, a semantic feature vector, and a contour feature vector of the image to be recognized; the processing module is further configured to obtain at least one target viewpoint matching the image to be recognized based on the multimodal feature vector of the image to be recognized and the multimodal feature vector of each viewpoint in a target region, wherein the target region is the region where the scene in the image to be recognized is located, and the target viewpoint is used to indicate the geographical location of the scene in the image to be recognized in the target region; and a display module for displaying the at least one target viewpoint to the user. 10. The apparatus according to claim 9, wherein the processing module is further configured to: construct a model of the target region; obtain each viewpoint of the target region from the model of the target region, wherein each viewpoint is used to indicate each location in the target region; obtain a multimodal feature vector of each viewpoint, wherein the multimodal feature vector of each viewpoint includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of each viewpoint. 11. The apparatus according to claim 9 or 10, wherein the receiving module is further configured to receive information about the target region input by the user, wherein the information about the target region is used to indicate the target retrieval range of the target region, and the scene in the image to be identified is located within the target retrieval range of the target region; the processing module is specifically configured to:Based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target retrieval range, at least one target viewpoint matching the image to be identified is obtained. 12. The apparatus according to any one of claims 9 to 11, wherein the processing module is further configured to obtain an encoding value for each viewpoint based on the multimodal feature vector of each viewpoint in the target region, the encoding value of each viewpoint being used to determine at least one target viewpoint matching the image to be identified. 13. The apparatus according to claim 12, wherein the processing module is specifically configured to: obtain a plurality of recommended target viewpoints matching the image to be identified based on the multimodal feature vector of the image to be identified and the encoding value of each viewpoint; and filter at least one target viewpoint matching the image to be identified from the plurality of recommended target viewpoints based on the multimodal feature vector of the image to be identified and the multimodal feature vectors of the plurality of recommended target viewpoints. 14. The apparatus according to claim 13, wherein the display module is further configured to display the plurality of recommended target viewpoints to the user; the receiving module is further configured to receive a set of viewpoints input by the user, the set of viewpoints including a plurality of viewpoints selected by the user from the plurality of recommended target viewpoints; and the processing module is specifically configured to: filter at least one target viewpoint matching the image to be identified from the set of viewpoints based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the set of viewpoints. 15. The apparatus according to any one of claims 9 to 14, wherein the processing module is further configured to obtain information about buildings in the image to be identified, the information about buildings including at least one of the following: the name of the building, an identification ID; and the display module is further configured to display the information about buildings in the image to be identified to the user when displaying the at least one target viewpoint. 16. The apparatus according to any one of claims 9 to 15, wherein the display module is further configured to, while displaying the at least one target viewpoint to the user, display to the user the degree of matching between each target viewpoint and the image to be identified, the degree of matching indicating the similarity between the scene within the field of view of each target viewpoint and the scene of the image to be identified. 17. A computing device cluster, comprising at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device being configured to execute data stored in the memory of the at least one computing device.18. A computer program product comprising instructions, characterized in that, when the instructions are executed by a computing device cluster, the computing device cluster performs the method as described in any one of claims 1 to 8. 19. A computer-readable storage medium, characterized in that it comprises computer program instructions, which, when executed by a computing device cluster, cause the computing device cluster to perform the method as described in any one of claims 1 to 8. Claims 3 / 3 Page 4 CN 120852972 A Method, Apparatus and Computing Device for Image Information Recognition Technical Field
[0001] This application relates to the field of image processing, and more specifically, to a method, apparatus and computing device for image information recognition. Background Art
[0002] Currently, in many scenarios, recognizing the spatial location (also known as geographical location) of a scene in an image is a common pursuit, which may include, but is not limited to: urban governance, disaster prevention and mitigation, Internet applications, criminal investigation applications, etc.
[0003] One related technology for image information recognition is web search technology, which identifies landmarks, objects, people, clothing, etc. in the captured image by uploading it to the network, thereby finding similar images or similar products in the image. This technology cannot identify the geographical location of the scene in the image.
[0004] Another related technology for image information recognition is augmented reality (AR) visual positioning technology. This AR visual positioning technology requires obtaining the initial location information of the image and identifying the geographical location of the scene in the image based on the initial location information. When identifying the geographical location of the scene in the captured image, this technology relies not only on the existing map of the target region (e.g., the target city where the image is located) but also on the initial location information of the image. However, most images that people encounter in daily life or work do not have initial location information, so this technology has certain limitations in identifying the geographical location of the scene in the captured image.
[0005] Therefore, how to identify the geographical location of the scene in the image based on the existing map of the target region where the image is located without other prior information has become an urgent technical problem to be solved.
[0006] The present application provides a method for image information recognition, which can identify the geographical location of a scene in an image based on an existing map of the target region where the image to be recognized is located, without any other prior information.
[0007] In a first aspect, an image information recognition method is provided, the method comprising: receiving a user input of an image to be recognized.The image to be identified is used to obtain the multimodal feature vector of the image to be identified. Based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region, at least one target viewpoint matching the image to be identified is obtained and displayed to the user. The target region is the region where the scene in the image to be identified is located. The multimodal feature vector of the image to be identified includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of the image to be identified. The target viewpoint is used to indicate the geographical location of the scene in the image to be identified within the target region.
[0008] The target region may include, but is not limited to: rural areas, counties, cities, provinces, countries, and even the entire world.
[0009] In the above technical solution, by using the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region, at least one target viewpoint matching the image to be identified is determined, thereby identifying the geographical location of the scene in the image to be identified in the target region. In this way, without other prior information, the geographical location of the scene in the image to be identified can be identified based on the existing map of the target region where the image to be identified is located, thereby effectively improving the accuracy and efficiency of image retrieval and avoiding some limitations of existing methods for identifying the geographical location of images.
[0010] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: constructing a model of the target region; obtaining each viewpoint of the target region from the model of the target region; obtaining and storing the multimodal feature vector of each viewpoint, wherein the multimodal feature vector of each viewpoint includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of each viewpoint.
[0011] In the above technical solution, a model of the target region can be established, thereby obtaining viewpoints on the model of the target region, and performing panoramic visual rendering on each available viewpoint to extract features such as depth, semantics, and texture, and establishing a unified feature expression for features such as depth, semantics, and texture to form a multimodal feature vector. In this way, the map features of the target region can be effectively extracted, eliminating the dependence on the initial location information, so that the geographical location of the image to be identified in the target region can be determined in the absence of other prior knowledge.
[0012] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: receiving information of the target region input by the user, the information of the target region being used to indicate the target retrieval range of the target region, and the scene in the image to be identified being located within the target retrieval range; obtaining at least one target viewpoint matching the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region, including:Based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target retrieval range, at least one target viewpoint matching the image to be identified is obtained.
[0013] In the above technical solution, the user can also input or import a retrieval range, which is used to indicate a search range in the target region. In this way, the geographical location of the image to be identified can be determined from the search range, avoiding direct search from the entire target region to determine the geographical location of the image to be identified, thus improving the retrieval efficiency.
[0014] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: encoding each viewpoint based on the multimodal feature vector of each viewpoint in the target region to obtain the encoding value of each viewpoint, and the encoding value of each viewpoint is used to determine at least one target viewpoint matching the image to be identified.
[0015] In conjunction with the first aspect, in some implementations of the first aspect, obtaining at least one target viewpoint matching the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region includes: obtaining multiple recommended target viewpoints matching the image to be identified based on the multimodal feature vector of the image to be identified and the encoding value of each viewpoint; and selecting at least one target viewpoint matching the image to be identified from the multiple recommended target viewpoints based on the multimodal feature vector of the image to be identified and the multimodal feature vectors of the multiple recommended target viewpoints.
[0016] In the above technical solution, multiple recommended target viewpoints matching the image to be identified can be quickly obtained based on the encoding value of each viewpoint, and then at least one target viewpoint matching the image to be identified can be selected from the multiple recommended target viewpoints based on the multimodal feature vector of the image to be identified and the multimodal feature vectors of the multiple recommended target viewpoints. This can further improve the efficiency of image retrieval.
[0017] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: displaying the plurality of recommended target viewpoints to the user; receiving a set of viewpoints input by the user, the set of viewpoints including a plurality of viewpoints selected by the user from the plurality of recommended target viewpoints; the step of filtering at least one target viewpoint matching the image to be identified from the plurality of recommended target viewpoints based on the multimodal feature vector of the image to be identified and the multimodal feature vectors of the plurality of recommended target viewpoints, including: filtering at least one target viewpoint matching the image to be identified from the set of viewpoints based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the set of viewpoints.
[0018] In the above technical solution, multiple recommended target viewpoints can be displayed to the user, who can then select a set of viewpoints from the multiple recommended target viewpoints. Based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the user-selected viewpoint set, at least one target viewpoint matching the image to be identified is selected from the viewpoint set. In this way, at least one target viewpoint matching the image to be identified can be obtained quickly, thereby further improving the efficiency and accuracy of image retrieval.
[0019] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: obtaining information about buildings in the image to be identified, the information about buildings including at least one of the following: the name of the building, identification ID; and displaying the information about buildings in the image to be identified to the user when displaying the at least one target viewpoint.
[0020] In the above technical solution, information about buildings in the image to be identified can also be obtained and displayed to the user, so as to be applicable to various scenarios where buildings in the image to be identified need to be identified.
[0021] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: when displaying the at least one target viewpoint to the user, displaying to the user the matching degree between each target viewpoint and the image to be identified, the matching degree being used to indicate the similarity between the scene within the field of view of each target viewpoint and the scene of the image to be identified.
[0022] In the above technical solution, the matching degree between each target viewpoint and the image to be identified can also be displayed to the user, so that the user can quickly determine the geographical location of the scene in the image to be identified based on the similarity between the scene within the field of view of each target viewpoint and the scene of the image to be identified.
[0023] In conjunction with the first aspect, in some implementations of the first aspect, the viewpoint is an available viewpoint adaptively obtained from the feasible ground area and low-altitude area on the model of the target region.
[0024] In the above technical solution, available viewpoints can be adaptively obtained from the feasible ground area and low-altitude area on the map of the target region, avoiding obtaining all viewpoints on the map of the target region without filtering, which can further improve the efficiency of retrieval.
[0025] In a second aspect, a method for constructing a feature index library corresponding to a target region is provided. The method includes: obtaining each viewpoint in the target region; obtaining a multimodal feature vector for each viewpoint in the target region; and storing the multimodal feature vector for each viewpoint in the feature index library. The multimodal feature vector for each viewpoint includes at least two of the following feature vectors: a scene depth feature vector, a semantic feature vector, and a contour feature vector for each viewpoint.
[0026] In conjunction with the second aspect, in some implementations of the second aspect, the method further includes: fusing at least one feature vector of each available viewpoint to obtain a multimodal feature index, and storing it in a feature index library.
[0027] In conjunction with the second aspect, in some implementations of the second aspect, the method further includes: constructing a model of the target region; obtaining each viewpoint of the target region includes: obtaining each viewpoint of the target region from the model of the target region.
[0028] In conjunction with the second aspect, in some implementations of the second aspect, the method includes: receiving an image to be identified input by a user, obtaining the multimodal feature vector of the image to be identified, and obtaining and displaying to the user at least one target viewpoint matching the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region. Wherein, the target region is the region where the scene in the image to be identified is located, and the multimodal feature vector of the image to be identified includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of the image to be identified. The target viewpoint is used to indicate the geographical location of the scene in the image to be identified in the target region.
[0029] The above-mentioned target region may include, but is not limited to: rural areas, counties, cities, provinces, countries, and even the world.
[0030] In conjunction with the second aspect, in some implementations of the second aspect, the method further includes: constructing a model of the target region; obtaining each viewpoint of the target region from the model of the target region; obtaining and storing the multimodal feature vector of each viewpoint, wherein the multimodal feature vector of each viewpoint includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of each viewpoint.
[0031] In conjunction with the second aspect, in some implementations of the second aspect, the method further includes: receiving information about the target region input by a user, the information about the target region being used to indicate the target retrieval range of the target region, and the scene in the image to be identified being located within the target retrieval range; obtaining at least one target viewpoint matching the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region includes: obtaining at least one target viewpoint matching the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target retrieval range.
[0032] In conjunction with the second aspect, in some implementations of the second aspect, the method further includes: encoding each viewpoint based on the multimodal feature vector of each viewpoint in the target region to obtain an encoded value for each viewpoint, the encoded value of each viewpoint being used to determine at least one target viewpoint matching the image to be identified.
[0033] In conjunction with the second aspect, in some implementations of the second aspect, obtaining at least one target viewpoint matching the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region includes: obtaining a plurality of recommended target viewpoints matching the image to be identified based on the multimodal feature vector of the image to be identified and the encoding value of each viewpoint; and selecting at least one target viewpoint matching the image to be identified from the plurality of recommended target viewpoints based on the multimodal feature vector of the image to be identified and the multimodal feature vectors of the plurality of recommended target viewpoints.
[0034] In conjunction with the second aspect, in some implementations of the second aspect, the method further includes: displaying the plurality of recommended target viewpoints to the user; receiving a set of viewpoints input by the user, the set of viewpoints including a plurality of viewpoints selected by the user from the plurality of recommended target viewpoints; the step of filtering at least one target viewpoint matching the image to be identified from the plurality of recommended target viewpoints based on the multimodal feature vector of the image to be identified and the multimodal feature vectors of the plurality of recommended target viewpoints includes: filtering at least one target viewpoint matching the image to be identified from the set of viewpoints based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the set of viewpoints.
[0035] In conjunction with the second aspect, in some implementations of the second aspect, the method further includes: obtaining information about buildings in the image to be identified, the information about buildings including at least one of the following: the name of the building, an identification ID; and, when displaying the at least one target viewpoint to the user, displaying the information about buildings in the image to be identified to the user.
[0036] In conjunction with the second aspect, in some implementations of the second aspect, the method further includes: when displaying the at least one target viewpoint to the user, displaying to the user the matching degree between each target viewpoint and the image to be identified, the matching degree being used to indicate the similarity between the scene within the field of view of each target viewpoint and the scene of the image to be identified.
[0037] In conjunction with the second aspect, in some implementations of the second aspect, the viewpoint is an available viewpoint adaptively obtained from the ground feasible area and low-altitude area on the model of the target region. Specification 4 / 22 pages 8 CN 120852972 A
[0038] It should be understood that for the beneficial effects of the second aspect and various implementations of the second aspect, please refer to the beneficial effects of the first aspect and various implementations of the first aspect, which will not be repeated here.
[0039] In a third aspect, an image information recognition apparatus is provided, the apparatus comprising: a receiving module, a processing module, and a display module. The receiving module is used to receive an image to be identified input by a user; the processing module is used to acquire the image to be identified...The image's multimodal feature vector is used to obtain at least one target viewpoint that matches the image to be identified, based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region. The display module is used to display the at least one target viewpoint that matches the image to be identified to the user. Wherein, the target region is the region where the scene in the image to be identified is located, and the multimodal feature vector of the image to be identified includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of the image to be identified. The target viewpoint is used to indicate the geographical location of the scene in the image to be identified in the target region.
[0040] In conjunction with the third aspect, in some implementations of the third aspect, the processing module is further used to: construct a model of the target region; obtain each viewpoint of the target region from the model of the target region; obtain and store the multimodal feature vector of each viewpoint, and the multimodal feature vector of each viewpoint includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of each viewpoint.
[0041] In conjunction with the third aspect, in some implementations of the third aspect, the receiving module is further configured to receive information about the target region input by the user, the information about the target region being used to indicate the target retrieval range of the target region, and the scene in the image to be identified being located within the target retrieval range; the processing module is specifically configured to: obtain at least one target viewpoint matching the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target retrieval range.
[0042] In conjunction with the third aspect, in some implementations of the third aspect, the processing module is further configured to: encode each viewpoint based on the multimodal feature vector of each viewpoint in the target region to obtain the encoded value of each viewpoint, the encoded value of each viewpoint being used to determine at least one target viewpoint matching the image to be identified.
[0043] In conjunction with the third aspect, in some implementations of the third aspect, the processing module is specifically configured to: obtain multiple recommended target viewpoints matching the image to be identified based on the multimodal feature vector of the image to be identified and the encoding value of each viewpoint; and select at least one target viewpoint matching the image to be identified from the multiple recommended target viewpoints based on the multimodal feature vector of the image to be identified and the multimodal feature vectors of the multiple recommended target viewpoints.
[0044] In conjunction with the third aspect, in some implementations of the third aspect, the display module is further configured to display the multiple recommended target viewpoints to the user; the receiving module is further configured to receive a set of viewpoints input by the user, the set of viewpoints including multiple viewpoints selected by the user from the multiple recommended target viewpoints; the processing module is specifically configured to: obtain multiple recommended target viewpoints matching the image to be identified based on the multimodal feature vector of the image to be identified and the encoding value of each viewpoint; and select at least one target viewpoint matching the image to be identified from the multiple recommended target viewpoints based on the multimodal feature vector of the image to be identified and the multimodal feature vectors of the multiple recommended target viewpoints.The multimodal feature vectors of other images and the multimodal feature vectors of each viewpoint in the viewpoint set are used to filter at least one target viewpoint that matches the image to be identified from the viewpoint set.
[0045] In conjunction with the third aspect, in some implementations of the third aspect, the processing module is further configured to obtain information about buildings in the image to be identified, the information about buildings including at least one of the following: the name of the building, identification ID; the display module is further configured to display the information about buildings in the image to be identified to the user when displaying the at least one target viewpoint to the user.
[0046] In conjunction with the third aspect, in some implementations of the third aspect, the display module is further configured to display to the user the matching degree between each target viewpoint and the image to be identified when displaying the at least one target viewpoint to the user, the matching degree being used to indicate the similarity between the scene within the field of view of each target viewpoint and the scene of the image to be identified.
[0047] In conjunction with the third aspect, in some implementations of the third aspect, the viewpoint is an available viewpoint adaptively obtained from the feasible ground area and low-altitude area on the model of the target region.
[0048] It should be understood that for the beneficial effects of the third aspect and its various implementations, please refer to the beneficial effects of the first aspect and its various implementations, which will not be repeated here.
[0049] In a fourth aspect, an apparatus for constructing a corresponding feature index library for a target region is provided. The apparatus includes: a processing module and a storage module, wherein the processing module is used to obtain each viewpoint in the target region and obtain a multimodal feature vector of each viewpoint in the target region, and the storage module is used to store the multimodal feature vector of each viewpoint in the target region in the feature index library, wherein the multimodal feature vector of each viewpoint includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of each viewpoint.
[0050] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the processing module is further configured to fuse at least two feature vectors of each viewpoint to obtain a multimodal feature index for each viewpoint, and the storage module is further configured to store the multimodal feature index of each viewpoint in a feature index library.
[0051] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the processing module is specifically configured to: construct a model of the target region; and obtain each viewpoint of the target region from the model of the target region.
[0052] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the device further includes: a processing module and a storage module, wherein the processing module is configured to obtain each viewpoint in the target region and obtain the multimodal feature index of each viewpoint in the target region.The storage module is used to store the multimodal feature vector of each viewpoint in the target region in the feature index library, wherein the multimodal feature vector of each viewpoint includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of each viewpoint.
[0053] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the processing module is further used to: construct a model of the target region; obtain each viewpoint of the target region from the model of the target region; obtain and store the multimodal feature vector of each viewpoint, wherein the multimodal feature vector of each viewpoint includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of each viewpoint.
[0054] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the apparatus further includes: a receiving module and a display module, wherein the receiving module is configured to receive an image to be identified input by a user; the processing module is further configured to acquire the multimodal feature vector of the image to be identified, and obtain at least one target viewpoint matching the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region; the display module is configured to display to the user at least one target viewpoint matching the image to be identified. The target region is the region where the scene in the image to be identified is located, and the multimodal feature vector of the image to be identified includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of the image to be identified; the target viewpoint is used to indicate the geographical location of the scene in the image to be identified in the target region.
[0055] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the receiving module is further configured to receive information about the target region input by the user, the information about the target region being used to indicate the target retrieval range of the target region, and the scene in the image to be identified being located within the target retrieval range; the processing module is specifically configured to: obtain at least one target viewpoint matching the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target retrieval range.
[0056] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the processing module is further configured to: encode each viewpoint based on the multimodal feature vector of each viewpoint in the target region to obtain the encoded value of each viewpoint, the encoded value of each viewpoint being used to determine at least one target viewpoint matching the image to be identified.
[0057] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the processing module is specifically used to: obtain multiple features matching the image to be identified based on the multimodal feature vector of the image to be identified and the encoding value of each viewpoint.Recommended target viewpoints; based on the multimodal feature vector of the image to be identified and the multimodal feature vectors of the multiple recommended target viewpoints, at least one target viewpoint matching the image to be identified is selected from the multiple recommended target viewpoints.
[0058] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the display module is further configured to display the multiple recommended target viewpoints to the user; the receiving module is further configured to receive a set of viewpoints input by the user, the set of viewpoints including multiple viewpoints selected by the user from the multiple recommended target viewpoints; the processing module is specifically configured to: based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the set of viewpoints, at least one target viewpoint matching the image to be identified is selected from the set of viewpoints.
[0059] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the processing module is further configured to obtain information about buildings in the image to be identified, the information about buildings including at least one of the following: the name of the building, identification ID; the display module is further configured to display the information about buildings in the image to be identified to the user when displaying the at least one target viewpoint to the user.
[0060] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the display module is further configured to, when displaying the at least one target viewpoint to the user, display to the user the degree of matching between each target viewpoint and the image to be identified, the degree of matching being used to indicate the similarity between the scene within the field of view of each target viewpoint and the scene of the image to be identified.
[0061] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the viewpoint is an available viewpoint adaptively obtained from the feasible ground area and low-altitude area on the model of the target region.
[0062] It should be understood that for the beneficial effects of the fourth aspect and its various implementations, please refer to the beneficial effects of the first aspect and its various implementations, which will not be repeated here.
[0063] In a fifth aspect, a computing device is provided, including a processor and a memory, and optionally, an input / output interface. The processor controls the input / output interface to send and receive information, the memory stores computer programs, and the processor retrieves and runs the computer programs from the memory, causing the computing device to execute the method in the first aspect or any possible implementation of the first aspect, or to execute the method in the second aspect or any possible implementation of the second aspect.
[0064] Optionally, the processor can be a general-purpose processor, which can be implemented in hardware or software. When implemented in hardware, the processor can be a logic circuit, integrated circuit, etc.; when implemented in software, the processor can be a general-purpose processor, implemented by reading software code stored in the memory, which can be an integrated circuit.The processor is located within the processor and can be located outside the processor, existing independently.
[0065] In a sixth aspect, a computing device cluster is provided, including at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method in the first aspect or any possible implementation of the first aspect, or executes the method in the second aspect or any possible implementation of the second aspect.
[0066] In a seventh aspect, a chip is provided, which acquires instructions and executes the instructions to implement the methods in the first aspect and any implementation of the first aspect.
[0067] Optionally, as one implementation, the chip includes a processor and a data interface, the processor reads instructions stored in the memory through the interface of the data specification 7 / 22 page 11 CN 120852972 A, and executes the methods in the first aspect and any implementation of the first aspect.
[0068] Optionally, as one implementation, the chip may further include a memory storing instructions, and the processor is configured to execute the instructions stored in the memory. When the instructions are executed, the processor is configured to execute the method of the first aspect and any implementation thereof, or execute the method of the second aspect or any possible implementation thereof.
[0069] In an eighth aspect, a computer program product including instructions is provided, which, when executed by a computing device, causes the computing device to execute the method of the first aspect and any implementation thereof, or execute the method of the second aspect or any possible implementation thereof.
[0070] In a ninth aspect, a computer program product including instructions is provided, which, when executed by a cluster of computing devices, causes the cluster of computing devices to execute the method of the first aspect and any implementation thereof, or execute the method of the second aspect or any possible implementation thereof.
[0071] In a tenth aspect, a computer-readable storage medium is provided, comprising computer program instructions that, when executed by a computing device, perform a method as described in the first aspect and any implementation thereof, or perform a method as described in the second aspect or any possible implementation thereof.
[0072] As examples, such computer-readable storage includes, but is not limited to, one or more of the following: read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), flash memory, electrically EPROM.EPROM, EEPROM) and hard drive.
[0073] Optionally, as an implementation, the above-mentioned storage medium may specifically be a non-volatile storage medium.
[0074] In an eleventh aspect, a computer-readable storage medium is provided, including computer program instructions, which, when executed by a computing device cluster, execute the method as described in the first aspect and any implementation thereof, or execute the method in the second aspect or any possible implementation thereof.
[0075] As an example, these computer-readable storages include, but are not limited to, one or more of the following: read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), flash memory, electrically EPROM (EEPROM), and hard drive.
[0076] Optionally, as an implementation, the above-mentioned storage medium may specifically be a non-volatile storage medium. Brief Description of the Drawings
[0077] FIG1 is a schematic block diagram of a system architecture applied to this application.
[0078] Figure 2 is a schematic flowchart of a method for constructing a feature index library corresponding to a target region according to an embodiment of this application.
[0079] Figure 3 is a schematic diagram of a 3D white model of a target region according to an embodiment of this application.
[0080] Figure 4 is a schematic diagram of at least one feature corresponding to a viewpoint according to an embodiment of this application.
[0081] Figure 5 is a schematic diagram of using a neural network to extract at least one feature vector for each available viewpoint according to an embodiment of this application.
[0082] Figure 6 is a schematic flowchart of an image information recognition method according to an embodiment of this application. Specification 8 / 22 pages 12 CN 120852972 A
[0083] Figure 7 is a schematic diagram of an image to be recognized input by a user according to an embodiment of this application.
[0084] Figure 8 is a schematic diagram of another image to be recognized input by a user according to an embodiment of this application.
[0085] Figure 9 is a schematic diagram of at least one feature extracted from the image to be recognized shown in Figure 7 according to an embodiment of this application.
[0086] Figure 10 is a schematic diagram of the output of the image to be recognized shown in Figure 7 according to an embodiment of this application.
[0087] FIG11 is a schematic diagram of the output for the image to be recognized shown in FIG8 provided in an embodiment of this application.
[0088] FIG12 is a schematic block diagram of an image information recognition device 1200 provided in an embodiment of this application.
[0089] Figure 13 is a schematic block diagram of an apparatus 1300 for constructing a feature index library corresponding to a target region according to an embodiment of this application.
[0090] Figure 14 is a schematic diagram of the architecture of a computing device 1500 according to an embodiment of this application.
[0091] Figure 15 is a schematic diagram of the architecture of a computing device cluster according to an embodiment of this application.
[0092] Figure 16 is a schematic diagram of the connection between computing devices 1500A and 1500B via a network according to an embodiment of this application. Detailed Description
[0093] The technical solutions in this application will be described below with reference to the accompanying drawings.
[0094] This application will present various aspects, embodiments, or features around a system including multiple devices, components, modules, etc. It should be understood and appreciated that each system may include other devices, components, modules, etc., and / or may not include all devices, components, modules, etc. discussed in conjunction with the accompanying drawings. In addition, combinations of these solutions can also be used.
[0095] In addition, in the embodiments of this application, the words "exemplary," "for example," etc. are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as an "example" in this application should not be construed as being better or more advantageous than other embodiments or designs. Specifically, the term "example" is used to present concepts in a concrete manner.
[0096] In the embodiments of this application, "corresponding" and "corresponding" can sometimes be used interchangeably, and it should be noted that their intended meanings are consistent without emphasizing the difference.
[0097] The business scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will understand that with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of this application are equally applicable to similar technical problems.
[0098] References such as "an embodiment" or "some embodiments" described in this specification mean that one or more embodiments of this application include specific features, structures, or characteristics described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0099] In this application, "at least one" means one or more, and "more than one" means two or more. "And / or,"The description of the relationship between related objects indicates that there can be three kinds of relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the related objects before and after are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items in the specification 9 / 22 page 13 CN 120852972 A, including any combination of single items or plural items. For example, at least one of a, b, or c can mean: a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple.
[0100] At present, under the premise of no or only a small amount of prior information, it is a common need to identify the spatial location (also known as the geographical location) of the scene in the image. Several possible application scenarios are introduced below.
[0101] 1. Urban governance scenario:
[0102] For example, based on photos taken by patrol officers / grid workers' mobile phones, confirm the detailed information of the patrol object in the photo, such as the geographical location of the patrol object.
[0103] 2. Disaster prevention and mitigation scenario:
[0104] For example, quickly locate suspected fire points captured by surveillance cameras, find the geographical location of the suspected fire point in the city, and initiate rapid rescue.
[0105] 3. Internet application scenario:
[0106] For example, for photos of interest in WeChat Moments, query the shooting location, scene, and other information of the photo scene.
[0107] 4. Criminal investigation application scenario:
[0108] For example, through a photo or a surveillance video, identify the location of the crime scene, thereby improving the chain of evidence.
[0109] One related technology for image information recognition is network search technology, which, by uploading images to the network, identifies landmarks, objects, people, clothing, etc. in the captured images, thereby finding similar images or similar products in the images. This technology cannot determine the geographical location of the scene in the image.
[0110] Another related image information recognition technology is augmented reality (AR) visual positioning technology. This AR visual positioning technology obtains the initial location information of the image through hardware sensors (e.g., global positioning system (GPS), IMU (inertial measurement unit, IMU), etc.), narrows the search range based on the initial location information, obtains a small search range from the existing map of the target region where the image is located (e.g., the target city where the image is located), and identifies the scene in the image within this small search range.The technology identifies the geographical location of a scene in a captured image. When identifying the geographical location of a scene in a captured image, this technology relies not only on existing maps of the target region but also on the initial location information of the image. However, most images encountered in daily life or work lack initial location information, thus limiting the technology's ability to identify the geographical location of a scene in a captured image.
[0111] In view of this, this application provides an image information recognition method that can identify the geographical location of a scene in an image based on existing maps of the target region without any prior information.
[0112] For ease of description, a system architecture of this application will be described in detail below with reference to Figure 1.
[0113] As an example, Figure 1 is a schematic block diagram of a system architecture of this application. As shown in Figure 1, the system may include two parts: the first part is used to construct a feature index library corresponding to the target region (also called the region to be searched), and the second part is used to obtain the geographical location of the scene in the user-input image (an image captured in the target region, also called the image to be identified) from the feature index library. The functions of the modules involved in the above two parts are described in detail below.
[0114] 1. First part: Used to construct the feature index library corresponding to the target region
[0115] As an example, the first part includes a map acquisition module, a map construction module, and a feature index library. Among them, the map acquisition module is used to collect map information of the target region through at least one method, the map construction module is used to construct a map of the target region based on the map information collected through at least one method, and the feature index library is used to store the feature vectors of each viewpoint in the map of the target region. The specific implementation process of this part will be described in detail below with reference to Figure 2, which will not be detailed here.
[0116] 2. Second part: Used to obtain the geographical location of the scene in the image to be identified by the user input from the feature index library.
[0117] As an example, the second part includes a visualization interface, an image feature extraction module, an image location calculation / recommendation module, and an application platform. The system includes a visual interface for users to input images to be identified and to output the geographical location information of the scene within those images. An image feature extraction module extracts feature vectors from the user-input image. An image location calculation / recommendation module compares the feature vectors of the image with those stored in a feature index to obtain and display the possible geographical locations of the scene within the image. Application platforms include, but are not limited to: government integrated query platforms, automatic disaster identification platforms, and emergency response command systems.Platforms, public security clue acquisition platforms, consumer application platforms, internet application platforms, etc. The specific implementation process of this part will be described in detail below with reference to Figure 5, which will not be detailed here.
[0118] Figure 2 is a schematic flowchart of a method for constructing a feature index library corresponding to a target region provided by an embodiment of this application. As shown in Figure 2, the method includes steps 210-250, which will be described in detail below.
[0119] Step 210: Collect map information of the target region.
[0120] As an example, this application embodiment can collect map information of the target region. There are many ways to collect it, and this application does not make specific limitations on this. For example, the map information of the target region can be obtained by satellite images collected by satellites. For example, the map information of the target region can also be obtained by pictures taken by aircraft. For example, the map information of the target region can also be obtained by panoramic street images collected by sensors installed on vehicles.
[0121] Step 220: Construct a three-dimensional model of the target region based on the map information of the target region.
[0122] As an example, let's take obtaining map information of a target area through satellite images acquired by satellite as an example. As an example, a 3D white model of the target area can be obtained through satellite images acquired by satellite. For example, data related to buildings, ground, roads, water systems, etc., of the target area can be extracted from satellite images. Image processing techniques and algorithms can be used to identify and extract feature points, edges, and textures from the satellite images. And by using professional 3D modeling software, the extracted feature points and edges can be converted into geometric shapes and vertices in three-dimensional space, thereby constructing a 3D white model of the target area. During the modeling process, factors such as lighting, materials, and textures can also be considered to make the model more realistic and lifelike.
[0123] For example, Figure 3 is a schematic diagram of a 3D white model of the target area constructed based on the above method.
[0124] It should be understood that the above-mentioned 3D white model of the target area is an uncolored model widely used in computer graphics. It usually uses geometric shapes, vertices, wireframes, etc. to represent objects in the real world (e.g., topographical information in a region). It is a highly abstract and simple model. Through preprocessing, the 3D white model can be semantically divided into different categories such as buildings, ground, roads, and water systems, and the outlines of features in different categories can be further extracted.
[0125] It should also be understood that the 3D white model of the target area can also be subjected to various operations (such as rotation, scaling, and movement), rendering, etc. on a computer.
[0126] Step 230: Obtain usable viewpoints from the 3D model of the target area.
[0127] As an example, all viewpoints can be obtained from the 3D model of the target area, and usable viewpoints can be filtered from the all viewpoints. The specific implementation process of obtaining all viewpoints and usable viewpoints is described in detail below.
[0128] The above-mentioned available viewpoints can also be understood as viewpoints of passable areas, that is, viewpoints of passable areas are selected from the full set of viewpoints, and the viewpoints of passable areas are used as available viewpoints. Specification 11 / 22 pages 15 CN 120852972 A
[0129] It should be noted that, in the embodiments of this application, the viewpoints corresponding to a certain position in the three-dimensional model of the target region can be described by six degrees of freedom (6DOF), wherein 6DOF includes the degrees of freedom of movement along the three rectangular coordinate axes x, y, and z and the degrees of freedom of rotation around the three coordinate axes x, y, and z.
[0130] The implementation method of obtaining the full set of viewpoints from the three-dimensional model of the target region is described in detail below.
[0131] In one possible implementation, the spatial range of the three-dimensional model of the target region can be divided into multiple grids according to a fixed size, and multiple viewpoints can be obtained based on the multiple grids according to the sampling rules. These multiple viewpoints are also the full set of viewpoints mentioned above.
[0132] The following details the implementation method for selecting usable viewpoints from the full set of viewpoints.
[0133] In one possible implementation, for each viewpoint in the full set of viewpoints, the three-dimensional coordinates x, y, z of the spatial position of each viewpoint can be obtained by vertical ray tracing, as well as the semantic type of the scene within the field of view of each viewpoint. The semantic type may include, but is not limited to: road, ground, building, water system, etc.
[0134] Step 240: Generate a multimodal feature index for each usable viewpoint and store the multimodal feature index of each usable viewpoint in the feature index library.
[0135] As an example, at least one feature can be extracted for each usable viewpoint. As shown in Figure 4, the at least one feature may include, but is not limited to: depth features, semantic features, and contour features of the scene within the field of view of the viewpoint.
[0136] In one possible implementation, for each usable viewpoint, according to the principle of panoramic spherical projection, the depth, semantic, contour, and other features of the scene within the field of view of each usable viewpoint are extracted within a certain range in the horizontal and vertical directions.
[0137] In this embodiment of the application, after obtaining at least one feature of each available viewpoint, at least one feature vector of each available viewpoint can be extracted, such as scene depth feature vector, semantic feature vector, contour feature vector, and at least one feature vector of each available viewpoint can be fused to obtain a multimodal feature index, which is then stored in a feature index library.
[0138] In one possible implementation, as shown in Figure 5, a neural network can be used to extract at least one feature vector of each available viewpoint to obtain a multimodal feature vector of each viewpoint. An adaptive spatial density distribution multimodal feature index is constructed for the multimodal feature vector of each viewpoint and stored in a feature index library.
[0139] Step 250: Encode each available viewpoint and store the encoded value of each available viewpoint in the feature index library.
[0140] As an example, in this embodiment of the application, hash encoding can be performed based on the 6DOF (e.g., x / y value) and multimodal feature index of each available viewpoint to obtain the encoded value of each available viewpoint, and the encoded value of each available viewpoint can be stored in the feature index library.
[0141] It should be understood that the encoded value of each available viewpoint can reflect the multimodal features such as depth, semantics, and contour of the scene within the field of view of each available viewpoint.
[0142] Optionally, square blocks can be formed in units of a certain length, and the viewpoints in each block can be divided into a feature extraction task, and multiple feature extraction tasks can form a task list.
[0143] In the above technical solution, available viewpoints are adaptively obtained in the feasible ground area and low-altitude area on the map of the target region, and panoramic visual rendering is performed on each available viewpoint to extract features such as depth, semantics, and texture. A unified feature expression of features such as depth, semantics, and texture is established to form a multimodal feature vector and store it in the feature index library. In this way, the map features of the target region can be effectively extracted, eliminating the dependence on the initial location information, so that the geographical location of the image to be identified in the target region can be determined in the absence of other prior knowledge.
[0144] Figure 6 is a schematic flowchart of an image information recognition method provided by an embodiment of this application. As shown in Figure 6, page 12 / 22 of the specification, 16 CN 120852972 A, the method includes steps 610-640, which are described in detail below.
[0145] Step 610: The user imports the image to be identified.
[0146] For example, the user can input or import the image to be identified, which can also be called the image to be retrieved.
[0147] Optionally, the user can also input information about the region where the image to be identified is located. This region can be referred to as the region to be searched or the target region.
[0148] The information about the region where the image to be identified is located can include, but is not limited to: the region where the image to be identified was taken, and the target search range of the target region. The scene in the image to be identified is located within the target search range of the target region.
[0149] The target search range is used to indicate a search area within the target region. This allows the geographical location of the image to be identified to be determined from this search range, avoiding direct searching of the entire target region to determine the geographical location of the image to be identified, thus improving search efficiency.
[0150] For example, the file for the search range can be in geojson format, supporting multiple polygons.
[0151] Example 1: Taking a command and inspection scenario as an example, the inspector / grid member takes a real-scene photo of the image to be identified as shown in Figure 7.For images, the inspector / grid administrator needs to confirm information about the objects to be inspected in the image (e.g., the name or ID of the building in the upper left corner of the image). This allows for the construction of an urban patrol and governance application based on spatial search and positioning services, which is then linked with a city information modeling (CIM) platform. Through digital empowerment, this improves the efficiency and precision of business processing. The inspector / grid administrator can input the image to be identified, as shown in Figure 7. If the inspector / grid administrator knows that the image was taken in Xi'an, the target region entered by the user is Xi'an.
[0152] Example 2, taking the disaster monitoring scenario as an example, the image to be identified shown in Figure 8 is a frame from the urban surveillance video of Changsha City. However, due to inadequate urban data governance, the camera installation location information is lost, has errors, or is unreliable, making it impossible to identify the suspected fire point in the image to be identified, thus failing to meet the needs of disaster prevention and mitigation. Therefore, it is necessary to determine the information of the suspected fire point in the image to be identified (e.g., the name or ID of the building where the suspected fire point is located, or which unit or floor of the building the suspected fire point is located in, etc.) in order to carry out disaster prevention and mitigation work. After obtaining the image to be identified shown in Figure 8, the urban management staff can input the image to be identified shown in Figure 8. If the urban management staff knows that the image to be identified was taken in Changsha City, the target region input by the user is Changsha City.
[0153] Step 620: Extract multimodal feature vectors from the image to be identified.
[0154] For example, at least one feature can be extracted from the image to be identified input by the user. The at least two features of the image to be identified may include, but are not limited to, scene depth features, semantic features, and contour features of the image to be identified.
[0155] For example, taking the image to be identified as shown in Figure 7, the scene depth features, semantic features, and contour features extracted from the image to be identified are shown in Figure 9.
[0156] In this embodiment of the application, after obtaining at least two features of the image to be identified, a multimodal feature vector of the image to be identified can also be extracted, such as the scene depth feature vector, semantic feature vector, and contour feature vector of the image to be identified.
[0157] In one possible implementation, the image to be identified can be extracted according to the principle of panoramic spherical projection, taking the horizontal and vertical directions within a certain range, and the scene depth, semantic, and contour features of the image to be identified can be extracted. A neural network is then used to extract at least two feature vectors for each available viewpoint to obtain the multimodal feature vector of the image to be identified.
[0158] Step 630: Based on the multimodal feature vector of the image to be identified, search for target viewpoints in the target region that match the image to be identified from the feature index library. (Specification 13 / 22 pages 17 CN 120852972 A)
[0159] In this embodiment of the application, after obtaining the multimodal feature vector of the image to be identified, the target viewpoint in the target region that matches the image to be identified can be searched from the feature index library corresponding to the target region.
[0160] There are multiple ways to search for the viewpoint in the target region that matches the image to be identified from the feature index library. This application does not limit this in any specific way. Several possible implementation methods are described below.
[0161] In one possible implementation method, the multimodal feature vector of the image to be identified is compared with the multimodal feature index of each viewpoint stored in the feature index library to find the target viewpoint that matches the image to be identified.
[0162] In another possible implementation method, multiple recommended target viewpoints that match the multimodal feature vector of the image to be identified are determined according to the multimodal feature vector of the image to be identified and the encoding value of each viewpoint stored in the feature index library. Then, the multimodal feature vector of the image to be identified is compared with the multimodal feature index of each recommended target viewpoint stored in the feature index library to find the target viewpoint that matches the image to be identified from the multiple recommended target viewpoints.
[0163] In another possible implementation, multiple recommended target viewpoints can be displayed to the user, who can then determine a set of viewpoints from the multiple recommended target viewpoints. This set of viewpoints includes multiple viewpoints selected by the user from the multiple recommended target viewpoints. Then, based on the multimodal feature vector of the image to be identified and the multimodal feature vectors of the multiple recommended target viewpoints, at least one target viewpoint that matches the image to be identified is selected from the multiple recommended target viewpoints. This can further improve the efficiency of image retrieval.
[0164] Optionally, in some embodiments, if the target viewpoints include multiple target viewpoints, the matching degree between each target viewpoint and the image to be identified can also be obtained. This matching degree is used to indicate the similarity between the scene within the field of view of each target viewpoint and the scene of the image to be identified.
[0165] Optionally, in some embodiments, if the image to be identified includes buildings, the name or ID of the buildings included in the image to be identified can be determined based on the data of the three-dimensional model corresponding to the target region and the 6DOF of the target viewpoints.
[0166] Step 640: Display the target viewpoints to the user.
[0167] In this embodiment of the application, after obtaining the viewpoint that matches the image to be identified, the target viewpoint can be displayed to the user. For example, the target viewpoint can be output to the user through a visualization interface.
[0168] It should be understood that since the viewpoint corresponding to a certain position in the three-dimensional model of the target region is represented by 6DOF, the 6DOF of the target viewpoint that matches the image to be identified, along the x, y, and z rectangular coordinate axes, represents the viewpoint.The dynamic degrees of freedom are used to indicate the shooting point of the image to be identified, and the rotational degrees of freedom about the three coordinate axes x, y, and z are used to indicate the shooting direction of the image to be identified. Therefore, the user can determine the geographical location of the scene in the image to be identified in the target region by using the target viewpoint.
[0169] Optionally, in some embodiments, if the target viewpoint includes multiple target viewpoints, the matching degree between each target viewpoint and the image to be identified can also be shown to the user when showing the user the at least one target viewpoint. For example, the matching degree between each target viewpoint and the image to be identified can be output to the user through a visualization interface. In this way, the user can determine the geographical location of the scene in the image to be identified in the target region based on the matching degree between each target viewpoint and the image to be identified.
[0170] In some embodiments, if the image to be identified includes buildings, the name or ID of the buildings included in the image to be identified can also be shown to the user when showing the user the at least one target viewpoint. For example, the name or ID of the buildings included in the image to be identified can be output to the user through a visualization interface.
[0171] Example 1: Taking the image to be identified by the user as shown in Figure 7 as an example, the output of the visualization interface in this embodiment of the application is shown in Figure 10. For example, the building in the upper left corner of the image to be identified can be output as the building with ID 2926-038743 located on ×× Road, ×× District, Xi'an City.
[0172] Example 2: Taking the image to be identified by the user as shown in Figure 8 as an example, the output of the visualization interface in this embodiment of the application is shown in Figure 11. For example, the building in the image to be identified that is suspected to be the ignition point of a fire can be output as the building with ID 038496 located on ×× Road, ×× District, Changsha City.
[0173] In the above technical solution, by using the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region, at least one target viewpoint matching the image to be identified is determined, thereby identifying the geographical location of the scene in the image to be identified in the target region. In this way, without other prior information, the geographical location of the scene in the image to be identified can be identified based on the existing map of the target region where the image to be identified is located, thereby effectively improving the accuracy and efficiency of image retrieval and avoiding some limitations of existing methods for identifying the geographical location of images.
[0174] The method provided by the embodiments of this application has been described in detail above with reference to Figures 1 to 11. The embodiments of the device of this application will be described in detail below with reference to Figures 12 to 16. It should be understood that the description of the method embodiments corresponds to the description of the device embodiments. Therefore, the parts not described in detail can be referred to the above method embodiments.
[0175] Figure 12 is a schematic block diagram of an image information recognition device 1200 provided in an embodiment of this application. The device 1200 can be implemented by software, hardware, or a combination of both. The device 1200 provided in this embodiment can implement the method flow shown in Figure 6 of this embodiment. The device 1200 includes: a receiving module 1210, a processing module 1220, and a display module 1230. The receiving module 1210 is used to receive an image to be recognized input by a user; the processing module 1220 is used to obtain the multimodal feature vector of the image to be recognized, and based on the multimodal feature vector of the image to be recognized and the multimodal feature vector of each viewpoint in the target region, obtain at least one target viewpoint matching the image to be recognized; the display module 1230 is used to display the at least one target viewpoint matching the image to be recognized to the user. Wherein, the target region is the region where the scene in the image to be identified is located, and the multimodal feature vector of the image to be identified includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of the image to be identified. The target viewpoint is used to indicate the geographical location of the scene in the image to be identified in the target region.
[0176] Optionally, the processing module 1220 is further configured to: construct a model of the target region; obtain each viewpoint of the target region from the model of the target region; obtain and store the multimodal feature vector of each viewpoint, wherein the multimodal feature vector of each viewpoint includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of each viewpoint.
[0177] Optionally, the receiving module 1210 is further configured to receive information about the target region input by the user, the information about the target region being used to indicate the target retrieval range of the target region, and the scene in the image to be identified being located within the target retrieval range; the processing module 1220 is specifically configured to: obtain at least one target viewpoint matching the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target retrieval range.
[0178] Optionally, the processing module 1220 is further configured to: encode each viewpoint based on the multimodal feature vector of each viewpoint in the target region to obtain the encoded value of each viewpoint, the encoded value of each viewpoint being used to determine at least one target viewpoint matching the image to be identified.
[0179] Optionally, the processing module 1220 is specifically configured to: obtain multiple recommended target viewpoints matching the image to be identified based on the multimodal feature vector of the image to be identified and the encoding value of each viewpoint; and select at least one target viewpoint matching the image to be identified from the multiple recommended target viewpoints based on the multimodal feature vector of the image to be identified and the multimodal feature vectors of the multiple recommended target viewpoints. (Specification 15 / 22 pages 19 CN)120852972 A
[0180] Optionally, the display module 1230 is further configured to display the plurality of recommended target viewpoints to the user; the receiving module is further configured to receive the viewpoint set input by the user, the viewpoint set including the plurality of viewpoints selected by the user from the plurality of recommended target viewpoints; the processing module 1220 is specifically configured to: filter at least one target viewpoint matching the image to be identified from the viewpoint set according to the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the viewpoint set.
[0181] Optionally, the processing module 1220 is further configured to obtain information about buildings in the image to be identified, the information about buildings including at least one of the following: the name of the building, identification ID; the display module 1230 is further configured to display the information about buildings in the image to be identified to the user when displaying the at least one target viewpoint.
[0182] Optionally, the display module 1230 is further configured to, when displaying the at least one target viewpoint to the user, display the matching degree between each target viewpoint and the image to be identified, wherein the matching degree is used to indicate the similarity between the scene within the field of view of each target viewpoint and the scene of the image to be identified.
[0183] Optionally, the viewpoint is an available viewpoint adaptively obtained from the feasible ground area and low-altitude area on the model of the target region.
[0184] FIG13 is a schematic block diagram of an apparatus 1300 for constructing a corresponding feature index library of a target region provided in an embodiment of this application. The apparatus 1300 can be implemented by software, hardware, or a combination of both. The apparatus 1300 provided in this application embodiment can implement the method flow shown in FIG2 of this application embodiment. The apparatus 1300 includes: a processing module 1310 and a storage module 1320. The processing module 1310 is used to obtain each viewpoint in the target region and obtain the multimodal feature vector of each viewpoint in the target region. The storage module 1320 is used to store the multimodal feature vector of each viewpoint in the target region in a feature index library. The multimodal feature vector of each viewpoint includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of each viewpoint.
[0185] Optionally, the processing module 1310 is further used to fuse at least two feature vectors of each viewpoint to obtain a multimodal feature index of each viewpoint. The storage module 1320 is further used to store the multimodal feature index of each viewpoint in a feature index library.
[0186] Optionally, the processing module 1310 is specifically used to: construct a model of the target region; and obtain each viewpoint of the target region from the model of the target region.
[0187] Optionally, the device 1300 further includes: a receiving module and a display module, wherein the receiving module is used to receive user...The input is an image to be identified; the processing module 1310 is further configured to obtain the multimodal feature vector of the image to be identified, and based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region, obtain at least one target viewpoint that matches the image to be identified; the display module is configured to display to the user at least one target viewpoint that matches the image to be identified. The target region is the region where the scene in the image to be identified is located, and the multimodal feature vector of the image to be identified includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of the image to be identified. The target viewpoint is used to indicate the geographical location of the scene in the image to be identified within the target region.
[0188] Optionally, the receiving module is further configured to receive information about the target region input by the user, the information about the target region being used to indicate the target retrieval range of the target region, and the scene in the image to be identified being located within the target retrieval range; the processing module 1310 is specifically configured to: obtain at least one target viewpoint matching the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target retrieval range.
[0189] Optionally, the processing module 1310 is further configured to: encode each viewpoint based on the multimodal feature vector of each viewpoint in the target region to obtain the encoded value of each viewpoint, the encoded value of each viewpoint being used to determine at least one target viewpoint matching the image to be identified (page 16 / 22 of specification, 20 CN 120852972 A).
[0190] Optionally, the processing module 1310 is specifically configured to: obtain multiple recommended target viewpoints matching the image to be identified based on the multimodal feature vector of the image to be identified and the encoding value of each viewpoint; and filter at least one target viewpoint matching the image to be identified from the multiple recommended target viewpoints based on the multimodal feature vector of the image to be identified and the multimodal feature vectors of the multiple recommended target viewpoints.
[0191] Optionally, the display module is further configured to display the multiple recommended target viewpoints to the user; the receiving module is further configured to receive a set of viewpoints input by the user, the set of viewpoints including multiple viewpoints selected by the user from the multiple recommended target viewpoints; the processing module 1310 is specifically configured to: filter at least one target viewpoint matching the image to be identified from the set of viewpoints based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the set of viewpoints.
[0192] Optionally, the processing module 1310 is further configured to obtain information about a building in the image to be identified, the building information including at least one of the following: the building's name and identification ID; the display module is further configured to display the building to the user.When at least one target viewpoint is shown, information about buildings in the image to be identified is displayed to the user.
[0193] Optionally, the display module is further configured to, when showing the user the at least one target viewpoint, display to the user the matching degree between each target viewpoint and the image to be identified, the matching degree being used to indicate the similarity between the scene within the field of view of each target viewpoint and the scene of the image to be identified.
[0194] Optionally, the viewpoint is an available viewpoint adaptively obtained from the feasible ground area and low-altitude area on the model of the target region.
[0195] The device 1200 or device 1300 here may be embodied in the form of a functional module. The term "module" here may be implemented in software and / or hardware form, without specific limitation.
[0196] For example, a "module" may be a software program, hardware circuit or a combination of both that implements the above functions. For example, the implementation of the receiving module 1210 will be described below using device 1200 as an example. Similarly, the implementation of other modules, such as processing module 1220 and display module 1230, can refer to the implementation of receiving module 1210.
[0197] Receiving module 1210 is an example of a software functional unit. Receiving module 1210 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above-mentioned computing instance may be one or more. For example, receiving module 1210 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple geographically close data centers. Typically, a region may include multiple AZs.
[0198] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Typically, one VPC is set up within one region. Communication between two VPCs within the same region, and between VPCs in different regions, requires a communication gateway to be set up within each VPC to achieve interconnection between VPCs.
[0199] As an example of a hardware functional unit, the receiving module 1210 may include at least one...The receiving module 1210 can be a computing device, such as a server. Alternatively, it can be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD can be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0200] The multiple computing devices included in the receiving module 1210 can be distributed in the same region or in different regions. The multiple computing devices included in the receiving module 1210 can be distributed in the same Availability Zone (AZ) or in different AZs. Similarly, the multiple computing devices included in the receiving module 1210 can be distributed in the same Virtual Private Cloud (VPC) or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0201] Therefore, the modules of the various examples described in the embodiments of this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0202] It should be noted that: when the device provided in the above embodiments performs the above method, it is only an example of the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. For example, the receiving module 1210 can be used to perform any step in the above method, the processing module 1220 can be used to perform any step in the above method, and the display module 1230 can be used to perform any step in the above method. The steps implemented by the receiving module 1210, processing module 1220, and display module 1230 can be specified as needed. The receiving module 1210, processing module 1220, and display module 1230 respectively implement different steps in the above method to achieve all the functions of the above device.
[0203] In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process is detailed in the method embodiments above, which will not be repeated here.
[0204] The method provided in this application embodiment can be executed by a computing device, which can also be called a computer system. It includes a hardware layer, an operating system layer running on the hardware layer, and an application layer running on the operating system layer. The hardware layer includes hardware such as processing units, memory and memory control units, and the functions and structure of the hardware will be described in detail below. The operating system is any one or more computer operating systems that implement business processing through processes, such as Linux operating system, Unix operating system, Android operating system, iOS operating system or Windows operating system, etc. The application layer includes applications such as browsers, address books, word processing software, instant messaging software, etc. Optionally, the computer system is a handheld device such as a smartphone, or a terminal device such as a personal computer. This application does not particularly limit it, as long as it can be implemented by the method provided in this application embodiment. The execution subject of the method provided in this application embodiment can be a computing device, or a functional module in the computing device that can call and execute programs.
[0205] The computing device provided in this application embodiment will be described in detail below with reference to FIG14.
[0206] FIG14 is a schematic diagram of the architecture of a computing device 1500 provided in an embodiment of this application. The computing device 1500 may be a server, a computer, or other device with computing capabilities. The computing device 1500 shown in FIG14 includes: at least one processor 1510 and a memory 1520.
[0207] It should be understood that this application does not limit the number of processors and memories in the computing device 1500.
[0208] The processor 1510 executes the instructions in the memory 1520, causing the computing device 1500 to implement the method provided in this application. Alternatively, the processor 1510 executes the instructions in the memory 1520, causing the computing device 1500 to implement the various functional modules provided in this application, thereby implementing the method provided in this application.
[0209] Optionally, the computing device 1500 further includes a communication interface 1530. The communication interface 1530 uses a transceiver module, such as but not limited to a network interface card or transceiver, to enable communication between the computing device 1500 and other devices or communication networks.
[0210] Optionally, the computing device 1500 further includes a system bus 1540, wherein the processor 1510, memory 1520, and communication interface 1530 are respectively connected to the system bus 1540. The processor 1510 can access the memory through the system bus 1540.1520, for example, processor 1510 can read and write data or execute code in memory 1520 via system bus 1540. System bus 1540 is a peripheral component interconnect express (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The system bus 1540 is divided into address bus, data bus, control bus, etc. For ease of representation, only one thick line is used in Figure 14, but it does not mean that there is only one bus or one type of bus.
[0211] In one possible implementation, the main function of processor 1510 is to interpret the instructions (or code) of computer programs and process data in computer software. The instructions of the computer program and the data in the computer software can be stored in memory 1520 or cache 1516.
[0212] Optionally, processor 1510 may be an integrated circuit chip with signal processing capabilities. By way of example and not limitation, processor 1510 is a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. Among these, a general-purpose processor is a microprocessor, etc. For example, processor 1510 is a central processing unit (CPU).
[0213] Optionally, each processor 1510 includes at least one processing unit 1512 and a memory control unit 1514.
[0214] Optionally, the processing unit 1512, also called a core, is the most important component of the processor. The processing unit 1512 is manufactured from single-crystal silicon using a specific manufacturing process. All calculations, command reception, command storage, and data processing of the processor are performed by the core. The processing units independently execute program instructions, utilizing parallel computing capabilities to accelerate program execution. Each processing unit has a fixed logical structure. For example, a processing unit includes logical units such as a first-level cache, a second-level cache, an execution unit, an instruction-level unit, and a bus interface.
[0215] In one implementation example, the memory control unit 1514 is used to control the data interaction between the memory 1520 and the processing unit 1512. Specifically, the memory control unit 1514 receives a memory access request from the processing unit 1512 and, based on the memory access request...The request controls access to memory. As an example and not a limitation, the memory control unit is a device such as a memory management unit (MMU).
[0216] In one implementation example, each memory control unit 1514 addresses the memory 1520 via a system bus. An arbitrator (not shown in FIG. 14) is configured in the system bus to handle and coordinate contention for access by multiple processing units 1512.
[0217] In one implementation example, the processing unit 1512 and the memory control unit 1514 are connected via internal chip wiring, such as address lines, to enable communication between the processing unit 1512 and the memory control unit 1514.
[0218] Optionally, each processor 1510 also includes a cache 1516, where the cache is a buffer for data exchange (called a cache). When a processing unit 1512 needs to read data, it first searches for the required data in the cache. If found, it executes directly; otherwise, it searches for it in memory. Since the cache runs much faster than the memory, the cache helps the processing unit 1512 run faster.
[0219] The memory 1520 can provide running space for processes in the computing device 1500. For example, the memory 1520 stores the computer program (specifically, the program code) used to generate the process. After the computer program is run by the processor to generate the process, the processor allocates corresponding storage space for the process in the memory 1520. Further, the above-mentioned storage space specification 19 / 22 pages 23 CN 120852972 A further includes text segments, initialization data segments, bit initialization data segments, stack segments, heap segments, etc. The memory 1520 stores the data generated during the process execution in the storage space corresponding to the above-mentioned process, such as intermediate data or process data, etc.
[0220] Optionally, the memory is also called RAM, and its function is to temporarily store the computational data in the processor 1510, as well as the data exchanged with external storage such as hard disk. As long as the computer is running, the processor 1510 will load the data to be processed into memory for processing, and after the processing is completed, the processing unit 1512 will transmit the result.
[0221] By way of example and not limitation, the memory 1520 is volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. VolatileThe memory is random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory 1520 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0222] The structure of the computing device 1500 listed above is merely illustrative and is not limited thereto. The computing device 1500 of the embodiments of this application includes various hardware in prior art computer systems. For example, the computing device 1500 also includes other memories besides memory 1520, such as disk storage, etc. Those skilled in the art should understand that the computing device 1500 may also include other devices necessary for normal operation. Furthermore, depending on specific needs, those skilled in the art should understand that the computing device 1500 may also include hardware devices for implementing other additional functions. In addition, those skilled in the art should understand that the computing device 1500 may only include the devices necessary for implementing the embodiments of this application, and may not necessarily include all the devices shown in FIG. 14.
[0223] This application embodiment also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device may be a server. In some embodiments, the computing device may also be a desktop computer, a laptop computer, or a smartphone, etc., a terminal device.
[0224] As shown in FIG. 15, the computing device cluster includes at least one computing device 1500. The memory 1520 in one or more computing devices 1500 in the computing device cluster may store the same instructions for performing the above-described method.
[0225] In some possible implementations, the memory 1520 in one or more computing devices 1500 in the computing device cluster may also store a portion of the instructions for performing the above-described method. In other words, a combination of one or more computing devices 1500 can jointly execute the instructions of the above method.
[0226] It should be noted that the memory 1520 in different computing devices 1500 in the computing device cluster can store different instructions, which are used to execute some functions of the above-mentioned device. That is, the instructions stored in the memory 1520 in different computing devices 1500 can realize the functions of one or more modules in the above-mentioned device.
[0227] In some possible implementations, one or more computing devices in the computing device cluster can be connected through a network. The network can be a wide area network or a local area network, etc. Figure 16 shows a possible implementation. As shown in Figure 16, two computing devices 1500A and 1500B are connected through a network. Specifically, they are connected to the network through the communication interface in each computing device. Specification 20 / 22 pages 24 CN 120852972 A
[0228] It should be understood that the function of computing device 1500A shown in Figure 16 can also be performed by multiple computing devices 1500. Similarly, the function of computing device 1500B can also be performed by multiple computing devices 1500.
[0229] In this embodiment, a computer program product containing instructions is also provided. The computer program product may be a software or program product containing instructions that can run on a computing device or be stored in any available medium. When it runs on a computing device, it causes the computing device to perform the methods provided above, or causes the computing device to implement the functions of the apparatus provided above.
[0230] In this embodiment, a computer-readable storage medium is also provided. The computer-readable storage medium may be any available medium that the computing device can store, or a data storage device such as a data center containing one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that, when executed on a computing device, cause the computing device to perform the methods provided above.
[0231] It should be understood that in the various embodiments of this application, the order of the above processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0232] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0233] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0234] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, and the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.
[0235] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0236] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0237] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of this application. The aforementioned storage medium includes: USB flash drive, mobile hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, and other media that can store program code.
[0238] The above description is merely a specific embodiment of this application, but the protection scope of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the protection scope of this application. Therefore, the protection scope of this application should be determined by the protection scope of the claims. Specification 22 / 22 pages 26 CN120852972 A Figure 1 Instruction Manual Appendix 1 / 9 Page 27 CN 120852972 A Figure 2 Figure 3 Instruction Manual Appendix 2 / 9 Page 28 CN 120852972 A Figure 4 Figure 5 Instruction Manual Appendix 3 / 9 Page 29 CN 120852972 A Figure 6 Figure 7 Instruction Manual Appendix 4 / 9 Page 30 CN 120852972 A Figure 8 Figure 9 Instruction Manual Appendix 5 / 9 Page 31 CN 120852972 A Figure 10 Figure 11 Figure 12 Instruction Manual Appendix 6 / 9 Page 32 CN 120852972 A Figure 13 Figure 14 Instruction Manual Appendix 7 / 9 Page 33 CN 120852972 A Figure 15 Instruction Manual Appendix 8 / 9 Page 34 CN 120852972 A Figure 16 Instruction Manual Appendix 9 / 9 Page 35 CN 120852972 A METHOD AND APPARATUS FOR IMAGE INFORMATION RECOGNITION, AND COMPUTING DEVICE Abstract The present application provides a method for image information recognition. The method includes: receiving a to-be-recognized image input by a user; obtaining a multimodal feature vector of the to-be-recognized image; to-be-recognized image, and displaying the at least one target viewpoint to the user. The target region is a regionwhere a scene of the to- be-recognized image is located. The multimodal feature vector of the to-be-recognized image includes at least two of the following feature vectors: a scene depth feature vector, a semantic feature vector, and a contour feature vector of the to-be-recognized image. The target viewpoint is used to indicate a geographical location of the scene of the to-be-recognized image in the target region. With the method, the geographical location of the scene of the to-be-recognized image can be recognized based on an existing map of the target region where the to-be-recognized image is located.
Claims
1. A method for image information recognition, characterized in that, The method includes: Receives the image to be recognized from the user input; The multimodal feature vector of the image to be identified is obtained, and the multimodal feature vector of the image to be identified includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of the image to be identified; Based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region, at least one target viewpoint matching the image to be identified is obtained. The target region is the region where the scene in the image to be identified is located, and the target viewpoint is used to indicate the geographical location of the scene in the image to be identified in the target region. The user is shown at least one target viewpoint.
2. The method according to claim 1, characterized in that, The method further includes: Construct a model of the target region; Obtain each viewpoint of the target region from the model of the target region, and each viewpoint is used to indicate each location in the target region; Obtain the multimodal feature vector of each viewpoint, wherein the multimodal feature vector of each viewpoint includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of each viewpoint.
3. The method according to claim 1 or 2, characterized in that, The method further includes: The system receives information about the target region input by the user, which is used to indicate the target retrieval range of the target region. The scene in the image to be identified is located within the target retrieval range of the target region. The step of obtaining at least one target viewpoint matching the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region includes: Based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target retrieval range, at least one target viewpoint that matches the image to be identified is obtained.
4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Based on the multimodal feature vector of each viewpoint in the target region, the encoding value of each viewpoint is obtained, and the encoding value of each viewpoint is used to determine at least one target viewpoint that matches the image to be identified.
5. The method according to claim 4, characterized in that, The step of obtaining at least one target viewpoint matching the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region includes: Based on the multimodal feature vector of the image to be identified and the encoding value of each viewpoint, multiple recommended target viewpoints matching the image to be identified are obtained; Based on the multimodal feature vector of the image to be identified and the multimodal feature vectors of the plurality of recommended target viewpoints, at least one target viewpoint that matches the image to be identified is selected from the plurality of recommended target viewpoints.
6. The method according to claim 5, characterized in that, The method further includes: Show the user the multiple recommended target viewpoints; Receive the set of viewpoints input by the user, the set of viewpoints including multiple viewpoints selected by the user from the multiple recommended target viewpoints; The step of selecting at least one target viewpoint matching the image to be identified from the multiple recommended target viewpoints based on the multimodal feature vector of the image to be identified and the multimodal feature vectors of the multiple recommended target viewpoints includes: Based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the viewpoint set, at least one target viewpoint that matches the image to be identified is selected from the viewpoint set.
7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: Obtain information about buildings in the image to be identified, wherein the building information includes at least one of the following: the name of the building, and its identification ID; When displaying the at least one target viewpoint to the user, information about the buildings in the image to be identified is displayed to the user.
8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: When displaying the at least one target viewpoint to the user, the user is shown the degree of matching between each target viewpoint and the image to be identified, the degree of matching indicating the similarity between the scene within the field of view of each target viewpoint and the scene of the image to be identified.
9. An image information recognition device, characterized in that, The device includes: The receiving module is used to receive the image to be recognized input by the user; The processing module is used to obtain the multimodal feature vector of the image to be identified, wherein the multimodal feature vector of the image to be identified includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of the image to be identified; The processing module is further configured to obtain at least one target viewpoint that matches the image to be identified based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target region, wherein the target region is the region where the scene in the image to be identified is located, and the target viewpoint is used to indicate the geographical location of the scene in the image to be identified in the target region. The display module is used to display the at least one target viewpoint to the user.
10. The apparatus according to claim 9, characterized in that, The processing module is also used for: Construct a model of the target region; Obtain each viewpoint of the target region from the model of the target region, and each viewpoint is used to indicate each location in the target region; Obtain the multimodal feature vector of each viewpoint, wherein the multimodal feature vector of each viewpoint includes at least two of the following feature vectors: scene depth feature vector, semantic feature vector, and contour feature vector of each viewpoint.
11. The apparatus according to claim 9 or 10, characterized in that, The receiving module is further configured to receive information about the target region input by the user, the information about the target region being used to indicate the target retrieval range of the target region, and the scene in the image to be identified being located within the target retrieval range of the target region; The processing module is specifically used for: Based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the target retrieval range, at least one target viewpoint that matches the image to be identified is obtained.
12. The apparatus according to any one of claims 9 to 11, characterized in that, The processing module is further configured to obtain the encoding value of each viewpoint based on the multimodal feature vector of each viewpoint in the target region, and the encoding value of each viewpoint is used to determine at least one target viewpoint that matches the image to be identified.
13. The apparatus according to claim 12, characterized in that, The processing module is specifically used for: Based on the multimodal feature vector of the image to be identified and the encoding value of each viewpoint, multiple recommended target viewpoints matching the image to be identified are obtained; Based on the multimodal feature vector of the image to be identified and the multimodal feature vectors of the plurality of recommended target viewpoints, at least one target viewpoint that matches the image to be identified is selected from the plurality of recommended target viewpoints.
14. The apparatus according to claim 13, characterized in that, The display module is also used to display the multiple recommended target viewpoints to the user; The receiving module is further configured to receive the set of viewpoints input by the user, wherein the set of viewpoints includes multiple viewpoints selected by the user from the multiple recommended target viewpoints; The processing module is specifically used for: Based on the multimodal feature vector of the image to be identified and the multimodal feature vector of each viewpoint in the viewpoint set, at least one target viewpoint that matches the image to be identified is selected from the viewpoint set.
15. The apparatus according to any one of claims 9 to 14, characterized in that, The processing module is further configured to obtain information about buildings in the image to be identified, wherein the building information includes at least one of the following: the name of the building and its identifier ID; The display module is also used to display information about buildings in the image to be identified to the user when displaying the at least one target viewpoint to the user.
16. The apparatus according to any one of claims 9 to 15, characterized in that, The display module is further configured to, when displaying the at least one target viewpoint to the user, display the degree of matching between each target viewpoint and the image to be identified, wherein the degree of matching is used to indicate the similarity between the scene within the field of view of each target viewpoint and the scene of the image to be identified.
17. A computing device cluster, characterized in that, The system includes at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1 to 8.
18. A computer program product containing instructions, characterized in that, When the instruction is executed by the computing device cluster, the computing device cluster causes the computing device cluster to perform the method as described in any one of claims 1 to 8.
19. A computer-readable storage medium, characterized in that, It includes computer program instructions, which, when executed by a cluster of computing devices, perform the method as described in any one of claims 1 to 8.