Image navigation positioning method, device, equipment and medium
By matching image and text features in a location knowledge graph, the accuracy problem of visual positioning technology in areas with scarce data is solved, achieving higher precision and universality in image navigation and positioning.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-27
- Publication Date
- 2026-04-10
AI Technical Summary
Existing visual localization technologies suffer from decreased accuracy in areas with indistinct features or scarce data, leading to insufficient reliability of the models in global applications. Furthermore, the imbalance of training datasets may result in bias.
By identifying the image features of the target image, data matching is performed in a pre-stored location knowledge graph. The location knowledge graph, constructed by combining image and text features, determines the attribute information of the matching entity to determine the geographic information, and the location result is calibrated using real-time shooting location.
It improves the universality and accuracy of image-based navigation and positioning, reduces positioning errors in data-scarce areas, and enhances the reliability of the model for global application.
Smart Images

Figure CN121829495A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to an image navigation and positioning method, apparatus, device, and medium. Background Technology
[0002] Visual navigation is a crucial component of the PNT (Personal Navigation Technology) system. There are many technical approaches to achieving visual navigation, among which SLAM (Simultaneous Localization and Mapping) technology is widely used. SLAM technology can be divided into two main categories: Filter-based SLAM: This type of method uses Bayesian filtering or its derivative algorithms, such as Kalman filtering and particle filtering, to predict and update the robot's state using the state and control inputs from the previous time step, as well as sensor observations.
[0003] Graph-based SLAM: This type of method constructs a graph structure for the problem and uses optimization techniques such as least squares to optimize the robot's trajectory and map feature points, thereby improving the accuracy of localization and map building.
[0004] SLAM technology has a wide range of applications. For example, in autonomous vehicles, SLAM helps vehicles achieve precise localization and environmental perception; in the field of home robots, such as robotic vacuum cleaners, SLAM technology is used to achieve effective path planning and environmental mapping. SLAM is an important branch of visual navigation technology and is crucial for spatial perception and navigation in robots and automated systems.
[0005] Visual positioning technology has made significant progress in recent years, but research on positioning based on information mining from images themselves is still relatively limited, and existing methods also have certain limitations. These limitations not only affect the scope of application of the technology but also restrict its effectiveness in real-world scenarios.
[0006] Traditional approaches divide the Earth's surface into geographical units of varying scales and rely on geographic labels attached to images during training. This approach performs poorly in areas with indistinct features or scarce data.
[0007] The performance of traditional methods largely depends on their training dataset. If there are significantly more images in some regions than others, the model may become biased towards those regions. This bias can lead to decreased accuracy in localization in data-scarce areas, thus affecting the reliability of its global application. Summary of the Invention
[0008] To address the problems existing in the prior art, the present invention provides an image navigation and positioning method, apparatus, device, and medium.
[0009] This invention provides an image navigation and positioning method, comprising: Identify image features of the target image; The image features are matched against a pre-stored location knowledge graph to determine the matching entities in the location knowledge graph. The geographic information of the target image is determined based on the attribute information of the matching entity; The location knowledge graph is constructed from image features derived from analyzing image samples and text features derived from analyzing text samples.
[0010] According to an image navigation and positioning method provided by the present invention, the method further includes a step of acquiring the positioning knowledge graph, comprising: Obtain the image features of image samples and the text features of text samples; each image sample corresponds to at least one text sample; The image features of the image samples and the text features of the text samples are divided into entity-relationship structures to construct the localization knowledge graph.
[0011] According to an image navigation and positioning method provided by the present invention, data matching is performed on the image features in a pre-stored positioning knowledge graph to determine the matching entities in the positioning knowledge graph, including: The image features are matched against the attribute information of each entity in the localization knowledge graph to obtain entities whose matching degree with the image features is greater than a preset value. If there is only one entity, the entity is determined as the matching entity. If there are two or more entities, the entity with the larger matching degree is selected as the matching entity.
[0012] According to an image navigation and positioning method provided by the present invention, the method further includes: acquiring the real-time shooting location corresponding to the target image; accordingly, determining the positioning distance based on the geographic information and the real-time shooting location, and determining the geographic information of the target image as valid information when the positioning distance is within a preset numerical range.
[0013] According to an image navigation and positioning method provided by the present invention, the method further includes: when the positioning distance is outside a preset numerical range, determining the positioning distance again based on the geographic information determined by the next matched entity and the real-time shooting location.
[0014] According to an image navigation and positioning method provided by the present invention, the method further includes: acquiring text data for estimating a target image; accordingly, determining location features and object features on the image based on the text data; and performing data matching in a pre-stored positioning knowledge graph based on the location features and the object features to determine matching entities in the positioning knowledge graph.
[0015] According to an image navigation and positioning method provided by the present invention, the step of identifying the image features of the target image includes: segmenting the target image into image blocks, extracting features from each image block, and obtaining the image features of the target image.
[0016] The present invention also provides an image navigation and positioning device, comprising: The recognition module is used to identify the image features of the target image; The analysis module is used to perform data matching of the image features in a pre-stored location knowledge graph to determine the matching entities in the location knowledge graph; The positioning module is used to determine the geographic information of the target image based on the attribute information of the matching entity; wherein, the positioning knowledge graph is constructed from image features obtained by analyzing image samples and text features obtained by analyzing text samples.
[0017] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement any of the image navigation and positioning methods described above.
[0018] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the image navigation and positioning methods described above.
[0019] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements any of the image navigation and positioning methods described above.
[0020] This invention provides an image navigation and positioning method, apparatus, device, and medium. By identifying the image features of a target image, the image features are matched in a positioning knowledge graph to determine matching entities in the positioning knowledge graph. Based on the attribute information of the matching entities, the geographic information of the target image is determined. This achieves the combination and matching of image information and geographic information in the knowledge graph, providing a more universal and high-precision image navigation and positioning method and improving the accuracy of image positioning. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0022] Figure 1This is a flowchart illustrating the image navigation and positioning method provided by the present invention.
[0023] Figure 2 This is a schematic diagram of the image navigation and positioning device provided by the present invention.
[0024] Figure 3 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0026] The following is combined Figures 1-3 The present invention describes the image navigation and positioning method, apparatus, device, and medium.
[0027] Figure 1 This diagram illustrates a flowchart of an image navigation and positioning method provided by the present invention. (See attached diagram.) Figure 1 The method includes the following steps: Step 11: Identify the image features of the target image.
[0028] Step 12: Perform data matching on the image features in the pre-stored localization knowledge graph to determine the matching entities in the localization knowledge graph.
[0029] Step 13: Determine the geographic information of the target image based on the attribute information of the matching entity; the location knowledge graph is constructed by analyzing the image features obtained from the analysis of image samples and the text features obtained from the analysis of text samples.
[0030] Regarding steps 11-13, it should be noted that the method of this invention is used to investigate the source of image content in order to determine the location of the image content. For example, if an image shows a temple in Wutai Mountain, this invention can determine that the temple in the image is "Wutai Mountain Temple in Xinzhou City, Shanxi Province". As another example, if scenery on both sides of a highway is recorded, this invention can determine that the bridge in the recorded video is "Wuhan Yangtze River Bridge".
[0031] In this invention, image processing begins with extracting the image features of the target image. Images contain regions that can be used to identify image content; these regions serve as the objects of image recognition. For example, if an image contains Mount Wutai, then the temples of Mount Wutai can be the objects of recognition. In other words, the target image is segmented into image blocks, features are extracted from each image block, and the features from all image blocks are combined to form the image features of the target image.
[0032] In this invention, image features are matched against a pre-stored location knowledge graph to determine matching entities within the graph. A knowledge graph, as a structured semantic knowledge base, represents entities and their relationships graphically, providing a powerful tool for rapidly describing concepts and their associations in the physical world. A knowledge graph includes nodes (entities) and edges (relationships), where each node represents a real-world entity. Edges capture the relationships of interest between two nodes. Therefore, after constructing the knowledge graph, relationships are established between each entity, and each entity possesses its own attribute information, including image features and text features. Image features are object features identified from images. Text features are the textual representations of these object features. These textual representations contain geographic information.
[0033] In this invention, the localization knowledge graph is constructed based on image features derived from the analysis of image samples and text features derived from the analysis of text samples.
[0034] The image navigation and positioning method provided by this invention identifies the image features of the target image, performs data matching of the image features in a positioning knowledge graph, determines the matching entities in the positioning knowledge graph, and determines the geographic information of the target image based on the attribute information of the matching entities. This achieves the combination and matching of image information and geographic information in the knowledge graph, providing a more universal and high-precision image navigation and positioning method and improving the accuracy of image positioning.
[0035] In a further step of the above method, the acquisition of the location knowledge graph is explained in detail as follows: Obtain the image features of the image samples and the text features of the text samples; each image sample corresponds to at least one text sample.
[0036] The image features of image samples and the text features of text samples are divided into entity-relationship structures to construct a localization knowledge graph.
[0037] It's important to note that each image in the dataset undergoes preprocessing. Preprocessing operations include image resizing, cropping, and noise removal. Preprocessing reduces the complexity of subsequent processing and improves data quality. Each image is divided into multiple image patches, features are extracted from these patches, and the features from each patch are combined to form the image features of the target image.
[0038] Each image sample corresponds to at least one text sample. These text samples can correctly describe the content of each image, at least the geographical information. Feature extraction is performed on the text samples to extract their textual features.
[0039] The image features of related image samples and the text features of text samples are divided into entity-relationship structures to construct a localization knowledge graph.
[0040] Furthermore, the main focus is on explaining the process of matching image features against a pre-stored localization knowledge graph to determine matching entities within the graph, as detailed below: The image features are matched against the attribute information of each entity in the localization knowledge graph. Entities whose matching degree with the image features in the attribute information is greater than a preset value are obtained. If there is only one entity, it is determined as the matching entity. If there are two or more entities, the entity with the higher matching degree is selected as the matching entity.
[0041] It should be noted that in this invention, the extracted image features will inevitably differ from those extracted from different images. Therefore, when matching image features with the attribute information of various entities in the localization knowledge graph, the matching degree needs to be considered. However, there are many entities in the localization knowledge graph, so entities with a matching degree of a certain threshold need to be filtered out. For example, entities with a matching degree greater than 98% can be filtered out.
[0042] If only one entity is selected, that entity is chosen as the matching entity; if two or more entities are selected, the entity with the higher matching degree is chosen as the matching entity.
[0043] In this invention, since people record images they take at a location in real time, the purpose may be to identify the object being photographed. Therefore, the real-time shooting location corresponding to the target image is obtained. Because it is real-time shooting, if the location information identified from the image is correct, the real-time shooting location will be close to the object in the image content. Therefore, based on the geographical information determined from the image content and the real-time shooting location, a positioning distance is determined. When the positioning distance is within a preset value range, the geographical information of the target image is determined to be valid information.
[0044] Furthermore, if the positioning distance is outside the preset value range, it may indicate that the determined geographic information of the image is not accurate enough. Therefore, if a large number of entities have been matched before, the positioning distance can be determined again based on the geographic information determined by the next matched entity and the real-time shooting location. Then, the positioning distance is compared with the preset value range, and the validity of the geographic information is determined based on the comparison result.
[0045] In a further step of the above method, it is sometimes necessary to locate the image content and simultaneously assess its specific location. In this case, it is necessary to compare the assessed location with the determined location to comprehensively judge the localization result of the image content.
[0046] At this point, text data is acquired to predict the target image. Then, location features and object features on the image are determined based on the text data. Here, object features are actually descriptive content about objects within the image.
[0047] Location features and object features on images are determined based on text data; data matching is performed in a pre-stored location knowledge graph based on location features and object features to determine matching entities in the location knowledge graph.
[0048] The image navigation and positioning device provided by the present invention is described below. The image navigation and positioning device described below can be referred to in correspondence with the image navigation and positioning method described above.
[0049] Figure 2 A schematic flowchart of an image navigation and positioning device provided by the present invention is shown below. Figure 2 The device includes an identification module 21, an analysis module 22, and a positioning module 23, wherein: The recognition module is used to identify the image features of the target image; The analysis module is used to perform data matching of image features in a pre-stored localization knowledge graph to determine the matching entities in the localization knowledge graph; The localization module is used to determine the geographic information of the target image based on the attribute information of the matching entity; the localization knowledge graph is constructed from image features obtained by analyzing image samples and text features obtained by analyzing text samples.
[0050] Since the apparatus of this embodiment is based on the same principle as the method of the above embodiment, more detailed explanations will not be repeated here.
[0051] It should be noted that, in the embodiments of the present invention, the relevant functional modules can be implemented by a hardware processor.
[0052] The image navigation and positioning device provided by this invention identifies the image features of a target image, performs data matching of the image features in a positioning knowledge graph, determines the matching entities in the positioning knowledge graph, and determines the geographic information of the target image based on the attribute information of the matching entities. This realizes the combination and matching of image information and geographic information in the knowledge graph, providing a more universal and high-precision image navigation and positioning method, and improving the accuracy of image positioning.
[0053] Figure 3 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 3 As shown, the electronic device may include a processor 31, a communication interface 32, a memory 33, and a communication bus 34. The processor 31, communication interface 32, and memory 33 communicate with each other via the communication bus 34. The processor 31 can call logical instructions stored in the memory 33 to execute an image navigation and positioning method. This method includes: identifying image features of a target image; performing data matching of the image features in a pre-stored positioning knowledge graph to determine matching entities in the positioning knowledge graph; and determining the geographic information of the target image based on the attribute information of the matching entities. The positioning knowledge graph is constructed from image features obtained by analyzing image samples and text features obtained by analyzing text samples.
[0054] Furthermore, the logical instructions in the aforementioned memory 33 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0055] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the image navigation and positioning method provided by the above methods. The method includes: identifying image features of a target image; performing data matching of the image features in a pre-stored positioning knowledge graph to determine matching entities in the positioning knowledge graph; determining the geographic information of the target image based on the attribute information of the matching entities; the positioning knowledge graph is constructed from image features obtained by analyzing image samples and text features obtained by analyzing text samples.
[0056] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the image navigation and positioning method provided by the above methods. The method includes: identifying image features of a target image; performing data matching of the image features in a pre-stored positioning knowledge graph to determine matching entities in the positioning knowledge graph; determining the geographic information of the target image based on the attribute information of the matching entities; the positioning knowledge graph is constructed from image features obtained by analyzing image samples and text features obtained by analyzing text samples.
[0057] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0058] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0059] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An image navigation positioning method, characterized by, The method comprises: identifying image features of a target image; performing data matching on the image features in a pre-stored positioning knowledge graph to determine a matching entity in the positioning knowledge graph; determining geographical information of the target image based on attribute information of the matching entity; wherein the positioning knowledge graph is constructed by image features obtained by analyzing image samples and text features obtained by analyzing text samples.
2. The image navigation and positioning method of claim 1, wherein, The method further comprises a step of acquiring the positioning knowledge graph, comprising: acquiring image features of image samples and text features of text samples; each image sample corresponds to at least one text sample; performing entity-relation structure division on the image features of the image samples and the text features of the text samples to construct the positioning knowledge graph.
3. Image navigation positioning method according to claim 1 or 2, characterized in that, The data matching of the image features in the pre-stored positioning knowledge graph to determine the matching entity in the positioning knowledge graph comprises: performing data matching on the image features in attribute information of each entity in the positioning knowledge graph to obtain an entity whose matching degree in the attribute information to the image features is greater than a preset value; if there is only one entity, the entity is determined as the matching entity; if there are two or more entities, the entity with a greater matching degree is selected as the matching entity.
4. The image navigation and positioning method of claim 3, wherein, The method further comprises: acquiring a real-time shooting position corresponding to the target image; accordingly, based on the geographical information and the real-time shooting position, a positioning distance is determined, and when the positioning distance is within a preset numerical range, the geographical information of the target image is determined as valid information.
5. The image navigation and positioning method of claim 4, wherein, The method further comprises: when the positioning distance is outside the preset numerical range, the positioning distance is determined again based on geographical information determined by a next entity matched and the real-time shooting position.
6. The image navigation and localization method of claim 3, wherein, The method further comprises: acquiring text data for estimating the target image; accordingly, based on the text data, position features and object features on the image are determined; based on the position features and the object features, data matching is performed in a pre-stored positioning knowledge graph to determine a matching entity in the positioning knowledge graph.
7. The image navigation and localization method of claim 1, wherein, The method further comprises: acquiring text data for estimating the target image; accordingly, based on the text data, position features and object features on the image are determined; based on the position features and the object features, data matching is performed in a pre-stored positioning knowledge graph to determine a matching entity in the positioning knowledge graph.
8. An image navigation positioning apparatus characterized by comprising: The method further comprises: an identification module configured to identify image features of a target image; an analysis module configured to perform data matching on the image features in a pre-stored positioning knowledge graph to determine a matching entity in the positioning knowledge graph; a positioning module configured to determine geographical information of the target image based on attribute information of the matching entity; wherein the positioning knowledge graph is constructed by image features obtained by analyzing image samples and text features obtained by analyzing text samples.
9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the program to implement the image navigation positioning method according to any one of claims 1-7. 10.A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the image navigation positioning method according to any one of claims 1-7.