Geolocation method and apparatus based on multi-modal large model, device and medium

By generating text embedding vectors through a large multimodal model and matching them with a preset vector library, the problem of inaccurate positioning of SLAM technology in complex environments is solved, and stable and high-precision absolute geographic positioning is achieved.

CN120256489BActive Publication Date: 2025-10-10北京观微科技有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510751681.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-10-10
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

Existing SLAM technology cannot perform stable and reliable positioning in complex environments, especially in large-scale, dynamic, low-texture and extreme lighting scenarios, where there are problems of inaccurate positioning and susceptibility to dynamic environmental interference.

Method used

A geolocation method based on a multimodal large model is adopted. Through the connection between images and texts, a multimodal large model is used to generate descriptive text, which is converted into a text embedding vector and matched with a preset vector library to determine the positioning information of the target geographic area.

Benefits of technology

It achieves stable high-precision positioning in complex environments and can obtain absolute geographic location information, eliminating dependence on local feature points in the image and improving the reliability and accuracy of positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256489B_ABST
    Figure CN120256489B_ABST
Patent Text Reader

Abstract

The application provides a geographic positioning method and device based on a multi-modal large model, equipment and a medium, relating to the technical field of geographic positioning. The method comprises: inputting a target picture corresponding to a target geographic area into a multi-modal large model to obtain a target description text corresponding to the target picture output by the multi-modal large model; the multi-modal large model is used to generate the target description text based on the target picture; inputting the target description text into a text embedding model to obtain a target text embedding vector corresponding to the target description text output by the text embedding model; and matching the target text embedding vector with each text embedding vector in a preset vector library to determine target positioning information corresponding to the target geographic area. The application does not depend on local feature points of an image, so it is not easily affected by a dynamic environment, can stably position in a complex environment, and has high positioning accuracy. Compared with a simultaneous localization and mapping technology, the application can also obtain absolute position information through a matching process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of geolocation technology, and in particular to a geolocation method, device, equipment and medium based on a multimodal large model. Background Art

[0002] Geolocation technology is widely used in navigation, autonomous driving, augmented reality, robot positioning, geographic information systems and other fields.

[0003] Existing geolocation mainly uses the following two methods: (1) Simultaneous Localization and Mapping (SLAM), which allows robots and other devices to build environmental maps in real time through sensor data and determine their own positions; (2) Image matching positioning method based on panoramic database, which is to take photos and match features with street view images in the database to determine the shooting location.

[0004] However, SLAM technology still has many limitations in complex scenarios such as large-scale, dynamic, low-texture, and extreme lighting conditions. First, it relies on local feature points in the image, making it less adaptable to dynamic environments and susceptible to interference from moving objects such as pedestrians and vehicles, leading to tracking drift or failure. Second, in scenes with low texture or a large number of repetitive structures (such as white walls, glass curtain walls, and corridors), reliable tracking and positioning are difficult due to the lack of stable feature points. Furthermore, SLAM technology derives relative position information based on relative motion and cannot directly obtain absolute geographic location (such as longitude and latitude). Summary of the Invention

[0005] The present invention provides a geolocation method, device, equipment and medium based on a multimodal large model, which is used to solve the defect that the SLAM technology in the prior art cannot perform stable and reliable positioning in complex environments. The technical solution of the present invention connects images, texts, and vectors, and matches a preset vector library with the target text embedding vector, so that stable positioning can still be performed in complex environments, and the positioning accuracy is relatively high.

[0006] The present invention provides a geolocation method based on a multimodal large model, comprising the following steps.

[0007] Inputting a target image corresponding to a target geographical area into a multimodal large model, and obtaining a target description text corresponding to the target image output by the multimodal large model; the multimodal large model is used to generate the target description text based on the target image;

[0008] Inputting the target description text into a text embedding model to obtain a target text embedding vector corresponding to the target description text output by the text embedding model;

[0009] match the target text embedding vector with each text embedding vector in the preset vector library, and determine the target positioning information corresponding to the target geographic region.

[0010] According to the geographical positioning method based on the multi-modal large model provided by the application, the target text embedding vector is matched with each text embedding vector in the preset vector library, and the target positioning information corresponding to the target geographic region is determined, which comprises:

[0011] For each text embedding vector in the preset vector library, the similarity value of the target text embedding vector and the text embedding vector is calculated;

[0012] The first preset number of text embedding vectors sorted in descending order of similarity value are determined as matching vectors, and the target positioning information is determined according to each matching vector.

[0013] According to the geographical positioning method based on the multi-modal large model provided by the application, the target positioning information is determined according to each matching vector, which comprises:

[0014] For each matching vector, the relative similarity value corresponding to the matching vector is determined based on the similarity value of the matching vector, the highest matching vector with the highest similarity value and the lowest matching vector with the lowest similarity value among all the matching vectors;

[0015] The sum of the relative similarity values of all the matching vectors and the relative similarity value corresponding to the matching vector are used to determine the confidence value corresponding to the matching vector;

[0016] The target positioning information is determined based on the confidence value corresponding to each matching vector.

[0017] According to the geographical positioning method based on the multi-modal large model provided by the application, the target positioning information is determined based on the confidence value corresponding to each matching vector, which comprises:

[0018] The confidence threshold is determined based on the highest confidence value among the confidence values and a preset constant, the matching vector with a confidence value greater than or equal to the confidence threshold is determined as the first positioning vector, a plurality of first panoramic map data with the same identifier as each first positioning vector is determined in the panoramic map database, and the target positioning information corresponding to the target geographic region is determined based on each first panoramic map data; or,

[0019] determining a matching vector corresponding to a highest confidence value among the confidence values ​​as a second positioning vector, determining second panoramic image data having the same identifier as the second positioning vector in the panoramic image database, and determining target positioning information corresponding to the target geographic area based on the second panoramic image data;

[0020] The identifier of the panoramic image data in the panoramic image database corresponds to the identifier of the text embedding vector in the preset vector library.

[0021] According to a geolocation method based on a multimodal large model provided by the present invention, the panoramic image database is constructed by the following steps:

[0022] Acquire multiple panoramic images and panoramic image data corresponding to each of the panoramic images; the panoramic image data includes an identifier, positioning information, and shooting time corresponding to the panoramic image;

[0023] The panoramic image database is constructed based on the identifiers, positioning information and shooting times corresponding to all the panoramic images.

[0024] According to a geolocation method based on a multimodal large model provided by the present invention, the panoramic image includes a plurality of pixels; and the preset vector library is constructed by the following steps:

[0025] For each of the panoramic images, determining the spherical coordinates corresponding to each pixel based on the original coordinates of each pixel in the panoramic image; determining the pixel color value corresponding to each pixel based on the original coordinates of each pixel in the panoramic image and the pixel color value of the pixel corresponding to each pixel;

[0026] Determining the three-dimensional Cartesian coordinates corresponding to each pixel based on the spherical coordinates corresponding to each pixel;

[0027] Projecting the panoramic image onto a plurality of cubic faces, and determining the cubic face coordinates of each pixel on each cubic face based on the three-dimensional Cartesian coordinates corresponding to each pixel;

[0028] For each of the cube faces, determining, based on the pixel size of the cube face and the cube face coordinates of each of the pixels on the cube face, a texture pixel coordinate of each of the pixels on the cube face; and determining, based on the texture pixel coordinates and pixel color values ​​of each of the pixels on the cube face, a local scene image corresponding to the cube face;

[0029] For each of the panoramic images, multiple local scene images corresponding to the panoramic image are input into the multimodal large model to obtain description texts corresponding to each of the local scene images output by the multimodal large model;

[0030] Inputting each of the description texts corresponding to the panoramic image into the text embedding model to obtain a text embedding vector corresponding to each of the description texts output by the text embedding model;

[0031] The preset vector library is constructed based on the text embedding vectors corresponding to all the panoramic images.

[0032] According to a multimodal large model-based geolocation method provided by the present invention, inputting the target description text into a text embedding model to obtain a target text embedding vector corresponding to the target description text output by the text embedding model includes:

[0033] Inputting the target description text into a large language model to obtain a standardized text output by the large language model; the large language model is used to perform standardization processing on the target description text; the standardization processing includes vocabulary unification processing, sentence unification processing, removal of irrelevant characters, data format unification processing and stop word filtering processing;

[0034] The standardized text is input into the text embedding model to obtain the target text embedding vector output by the text embedding model; the text embedding model is used to generate the target text embedding vector after executing a long text splitting strategy, a word segmentation strategy and an embedding strategy on the standardized text.

[0035] The present invention also provides a geo-positioning device based on a multimodal large model, comprising the following modules:

[0036] A text module is configured to input a target image corresponding to a target geographic area into a multimodal large model, and obtain a target description text corresponding to the target image output by the multimodal large model; the multimodal large model is configured to generate the target description text based on the target image;

[0037] A vector module is used to input the target description text into a text embedding model to obtain a target text embedding vector corresponding to the target description text output by the text embedding model;

[0038] The positioning module is used to match the target text embedding vector with each text embedding vector in a preset vector library to determine the target positioning information corresponding to the target geographical area.

[0039] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the geolocation method based on the multimodal large model as described above is implemented.

[0040] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the multimodal large model-based geolocation methods described above.

[0041] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the above-described multimodal large model-based geolocation methods.

[0042] The geolocation method, device, equipment and medium based on the multimodal large model provided by the present invention are as follows: a target image corresponding to the target geographical area is input into the multimodal large model, and a target description text corresponding to the target image output by the multimodal large model is obtained; the multimodal large model is used to generate a target description text based on the target image; the target description text is input into a text embedding model, and a target text embedding vector corresponding to the target description text output by the text embedding model is obtained; the target text embedding vector is matched with each text embedding vector in a preset vector library to determine the target positioning information corresponding to the target geographical area. The technical solution of the present invention does not rely on the local feature points of the image through the connection between images, texts and vectors, and the matching of the preset vector library with the target text embedding vector. Therefore, it will not be easily affected by the dynamic environment. It can still be stably positioned in a complex environment, and the positioning accuracy is high. In addition, compared with SLAM technology, the present invention can also obtain absolute position information through the matching process. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0044] Figure 1 It is a flowchart of the geolocation method based on the multimodal large model provided by the present invention.

[0045] Figure 2 It is a structural schematic diagram of the geo-positioning device based on the multimodal large model provided by the present invention.

[0046] Figure 3 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0047] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0048] In view of the above problems in the prior art, the present invention provides a geolocation method based on a multimodal large model. Figure 1 This is a flow chart of the geolocation method based on the multimodal large model provided by the present invention, such as Figure 1 As shown, the method includes the following steps 110 to 130.

[0049] Step 110: Input a target image corresponding to a target geographical area into a multimodal large model to obtain a target description text corresponding to the target image output by the multimodal large model; the multimodal large model is used to generate the target description text based on the target image.

[0050] Specifically, the target geographical area can be any geographical area, for example, it can be Square A in City A. The target picture is a picture taken of the target geographical area, for example, it can be a picture taken with a mobile phone, a camera or other device with a camera function, and the target picture should be horizontal. For example, when it is necessary to locate a certain square, a picture of the square can be taken as the target picture. The multimodal large model can be, for example, GLM-4V, or other multimodal large models that can generate description text based on pictures. The embodiment of the present invention does not impose specific restrictions here. After the target picture is input into the multimodal large model, the multimodal large model can generate a target description text based on the target picture and a preset prompt word template, and then obtain the target description text corresponding to the target picture output by the multimodal large model, and the target description text is used to describe the characteristics of the target picture.

[0051] It's important to note that large multimodal models present several challenges: the random nature of language generation can lead to varying descriptions; differences in focus can lead to omissions of key information; vocabulary diversity can lead to different representations of the same thing; and environmental factors (such as time of day, weather, and lighting) can lead to variations in descriptions. Consequently, the descriptions generated by large multimodal models for the same image may vary. Therefore, preset prompt word templates can be pre-set based on needs. Furthermore, during application, the preset prompt word templates and the target image can be simultaneously input into the large multimodal model.

[0052] For example, the preset prompt word template may be as follows:

[0053] (1) Buildings:

[0054] Appearance description: {the shape, material, and color of the building, such as "a three-story red brick building"}.

[0055] Building use: {the function of a building, such as "the ground floor is a shop and the upper floors are residential"}.

[0056] Building location: {the building's relative position within the image, e.g., "located in the center of the image, adjacent to a main road on the left"}.

[0057] (2) Signboard:

[0058] Signboard content: {the complete text on the signboard, such as "XX Road 100 meters"}.

[0059] Position: {The relative position of the sign in the image, e.g., "located in the lower right corner, near the sidewalk"}.

[0060] (3) Roads and greening:

[0061] Road description: {road type, width, material, and markings, such as "asphalt two-way four-lane road with a yellow center line"}.

[0062] Green belt: {the vegetation type of the green belt, such as "there is a row of trees next to the sidewalk, and the tree species is sycamore"}.

[0063] (4) Other fixed facilities:

[0064] Infrastructure: {such as bus stops, benches, mailboxes, such as "bus stop, the sign says 'XX Station'"}.

[0065] Please describe the picture strictly according to the above format to ensure the content is complete and avoid missing key information. If any content is missing, please fill in "None".

[0066] The target text description A generated by the multimodal large model based on the target image and the preset prompt word template in the above example can be as follows:

[0067] (1) Buildings:

[0068] Appearance Description: The buildings are primarily made of glass and concrete. The building on the left is cylindrical, while the building on the right is rectangular. The glass windows reflect the yellow light, giving it a modern and technological feel.

[0069] Building Use: Based on the lights on in the windows, it is speculated that these buildings may be offices or commercial spaces.

[0070] Building Location: The building is located in an open square surrounded by other buildings and looks like it is in a business district of a city.

[0071] (2) Signboard:

[0072] Signage content: There is no obvious signage in the image.

[0073] Location: There is no sign in the picture.

[0074] (3) Roads and greening:

[0075] Road description: The road in the square is paved with gray bricks, of moderate width, and has no obvious traffic markings.

[0076] Green belt: There is no obvious green belt in the square.

[0077] (4) Other fixed facilities:

[0078] Other infrastructure: There are a few streetlights and trash cans in the square. On the right side of the picture, you can see a parked white truck.

[0079] Step 120: Input the target description text into a text embedding model to obtain a target text embedding vector corresponding to the target description text output by the text embedding model.

[0080] Specifically, after obtaining the target description text, the target description text can be input into a text embedding model, which can then process the target description text and output a target text embedding vector corresponding to the target description text. For example, the text embedding model can be the bge-base-1.5v-zh embedding model, which is a Chinese text embedding model developed by the Beijing Artificial Intelligence Research Institute. The text embedding model can also be other embedding models, which are not specifically limited in this embodiment of the present invention.

[0081] In one embodiment, inputting the target description text into a text embedding model to obtain a target text embedding vector corresponding to the target description text output by the text embedding model includes:

[0082] Inputting the target description text into a large language model to obtain a standardized text output by the large language model; the large language model is used to perform standardization processing on the target description text; the standardization processing includes vocabulary unification processing, sentence unification processing, removal of irrelevant characters, data format unification processing and stop word filtering processing;

[0083] The standardized text is input into the text embedding model to obtain the target text embedding vector output by the text embedding model; the text embedding model is used to generate the target text embedding vector after executing a long text splitting strategy, a word segmentation strategy and an embedding strategy on the standardized text.

[0084] Specifically, to improve the consistency of the target description text, the target description text can also be standardized. For example, the target description text can be input into a large language model, and the large language model can perform standardization on the target description text to obtain standardized text output by the large language model. This standardization process may include:

[0085] (1) Lexical unification: replace synonyms, simplify redundant descriptions, and make expressions more consistent. For example, "high-rise building" → "building", "office building" → "office building".

[0086] (2) Sentence structure unification: Unify the description structure so that different texts maintain the same grammatical format. For example, "This is a building with glass curtain walls" → "Building: glass curtain wall structure".

[0087] (3) Remove irrelevant characters: Delete unnecessary punctuation marks, spaces, etc. to ensure the text is clean and neat. For example, "Building: glass curtain wall." → "Building: glass curtain wall".

[0088] (4) Data format unification: Unify the measurement units to ensure data consistency. For example, "5m" → "5 meters", "10 meters" → "10 meters".

[0089] (5) Stop word filtering: remove words that affect semantic matching, such as "maybe", "seem" and other ambiguous words. For example, "This may be an office building" → "Building: office building".

[0090] For example, the standardized text A obtained by normalizing the target text description A in the above example may be as follows:

[0091] (1) Buildings:

[0092] Appearance Description: A glass and concrete structure, with the left building being cylindrical and the right building being rectangular. The glass windows reflect yellow light, creating a modern and technological feel.

[0093] Building Use: There are lights on in the windows, presumably an office building or commercial space.

[0094] Building location: Located in an open square, surrounded by other buildings, possibly belonging to the city's commercial area.

[0095] (2) Signboard:

[0096] Signage content: No obvious signage.

[0097] Location: No sign.

[0098] (3) Roads and greening:

[0099] Road Description: The square road is paved with gray bricks, with a moderate width and no obvious traffic markings.

[0100] Green belt: No obvious green belt.

[0101] (4) Other fixed facilities:

[0102] Infrastructure: The square is equipped with streetlights and trash cans. There is a parked white truck on the right side of the picture.

[0103] The specific adjustments from target description text A to standardized text A in the above example are as follows:

[0104] Vocabulary was unified: for example, "the building is mainly composed of glass and concrete" → "glass and concrete structure"; sentence structure was unified: redundant modifiers were removed, for example, "gives people a sense of modernity and technology" → "the overall style is modern and technological"; irrelevant characters were removed: irrelevant words such as "looks like" and "can be seen" were deleted; data format was unified: no values ​​needed to be adjusted; stop word filtering was performed: ambiguous words such as "maybe" and "seem" were removed.

[0105] Furthermore, after obtaining the standardized text, the standardized text can be input into a text embedding model. The text embedding model can perform long text splitting strategy, word segmentation strategy and embedding strategy on the standardized text, thereby generating a target text embedding vector.

[0106] Among them, in the long text splitting strategy, it can be split according to semantic units, split according to fixed length, or split in combination with themes. For example, in the embodiment of the present invention, it can be split according to fixed length, and every 800 words can be split into a text block. The last 200 words of the previous text block and the first 200 words of the next text block are the overlapping information of the previous and next contexts. Usually, the text description of a picture is about 350 words, and basically does not exceed 400 words. Therefore, splitting every 800 words into a text block ensures that each text block can independently cover all the information of a picture.

[0107] In word segmentation strategies, neural networks can be used to learn contextual features of text and dynamically adjust the strategy. SentencePiece can be used as a sub-model in word segmentation strategies. SentencePiece is a word segmentation tool that breaks text into smaller sub-word units with a certain frequency and statistical significance in the dictionary. The advantages of SentencePiece include its suitability for long Chinese texts, its ability to handle complex syntactic structures, its support for out-of-vocabulary words, its ability to avoid information loss, and its ability to maintain semantic integrity, thus improving the quality of downstream tasks such as search and dialogue generation.

[0108] Embedding strategies can include direct embedding, segmented embedding and aggregation, key content extraction and embedding, and multi-granularity embedding fusion. In one embodiment of the present invention, for example, direct embedding can be performed by inputting text into the Transformer-based bge-base-1.5v-zh model. bge-base-1.5v-zh uses a multi-layer self-attention mechanism to globally model the text and capture long-range dependencies.

[0109] In the above embodiment, standardizing the target description text increases the accuracy of the text description. Furthermore, a text embedding model can convert the standardized text into a text embedding vector that retains the corresponding features, enabling subsequent matching based on the text embedding vector, thereby improving the accuracy of matching and positioning. Furthermore, by analyzing core visual information such as buildings, signs, and roads, the system can accurately infer the actual geographic location of the photo, achieving high-precision positioning in urban environments or unlabeled areas.

[0110] Step 130: Match the target text embedding vector with each text embedding vector in a preset vector library to determine target positioning information corresponding to the target geographic area.

[0111] Specifically, the target text embedding vector may be matched with each text embedding vector in a preset vector library, and the target positioning information corresponding to the target geographical area may be determined based on the text embedding vector that matches the target text embedding vector.

[0112] In one embodiment, matching the target text embedding vector with each text embedding vector in a preset vector library to determine the target positioning information corresponding to the target geographical area includes:

[0113] For each of the text embedding vectors in the preset vector library, calculating a similarity value between the target text embedding vector and the text embedding vector;

[0114] A first preset number of text embedding vectors sorted in descending order according to similarity values ​​are determined as matching vectors, and the target positioning information is determined according to each of the matching vectors.

[0115] Specifically, for each text embedding vector in the preset vector library, the cosine similarity can be used to calculate the similarity value between the target text embedding vector and the text embedding vector. The similarity value corresponding to the text embedding vector It can be calculated using the following formula:

[0116]

[0117] in, represents the target text embedding vector, Indicates the number The text embedding vector of .

[0118] Furthermore, all text embedding vectors can be sorted from largest to smallest according to their corresponding similarity values, and then a first preset number of text embedding vectors ranked at the top can be determined as matching vectors. The first preset number can be set as needed, for example, 5 to 10. It should be noted that each text embedding vector has a corresponding identifier, and each matching vector also naturally has a corresponding identifier. The positioning information corresponding to each matching vector can be determined based on the corresponding identifier, and then the target positioning information of the target geographic area can be determined based on the positioning information corresponding to each matching vector.

[0119] In the above embodiment, the target text embedding vector is matched with the text embedding vector in the preset vector library according to the similarity value, and the target positioning information can be determined according to the matched text embedding vector, thereby getting rid of the dependence of the global positioning system in the traditional positioning method, and reliable positioning can be performed even in signal-limited environments (indoors or dense urban areas).

[0120] In one embodiment, determining the target positioning information according to each matching vector includes:

[0121] For each of the matching vectors, determining a relative similarity value corresponding to the matching vector based on the similarity value of the matching vector, a highest matching vector with the highest similarity value among all the matching vectors, and a lowest matching vector with the lowest similarity value;

[0122] Determining a confidence value corresponding to the matching vector based on a sum of the relative similarity values ​​of all the matching vectors and the relative similarity values ​​corresponding to the matching vectors;

[0123] The target positioning information is determined based on the confidence value corresponding to each matching vector.

[0124] Specifically, for each matching vector, the relative similarity value of the matching vector can be determined based on the similarity value of the matching vector, the highest matching vector corresponding to the highest similarity value and the lowest matching vector corresponding to the lowest similarity value. The process of solving the relative similarity value is essentially the process of normalizing the similarity value. The relative similarity value can ensure that different matching vectors are comparable. The relative similarity value corresponding to the matching vector It can be expressed by the following formula:

[0125]

[0126] in, represents the lowest matching vector, represents the highest matching vector.

[0127] Furthermore, the relative similarity value can measure the matching degree of two vectors, but it cannot measure the reliability of the result. Therefore, the confidence value can be further calculated based on the relative similarity value. The confidence value corresponding to the matching vector It can be expressed by the following formula:

[0128]

[0129] in, Represents the sum of the relative similarity values ​​of all matching vectors, The range is (0, 1], it is easy to understand that The larger the representation number, the The larger the proportion of the matching vector in all matching vectors.

[0130] After determining the confidence values ​​corresponding to the respective matching vectors, target positioning information may be determined based on the confidence values ​​corresponding to the respective matching vectors.

[0131] In the above embodiment, calculating the relative similarity value can ensure the comparability between different matching vectors. The further determined confidence value can reduce the risk of misjudgment and improve matching accuracy in complex scenarios (such as image matching and long text retrieval).

[0132] In one embodiment, determining the target positioning information based on the confidence value corresponding to each matching vector includes:

[0133] Determining a confidence threshold based on the highest confidence value among the confidence values ​​and a preset constant; determining a matching vector having a confidence value greater than or equal to the confidence threshold as a first positioning vector; determining a plurality of first panoramic image data having the same identifier as each of the first positioning vectors in a panoramic image database, and determining the target positioning information corresponding to the target geographic area based on each of the first panoramic image data; or,

[0134] determining a matching vector corresponding to a highest confidence value among the confidence values ​​as a second positioning vector, determining second panoramic image data having the same identifier as the second positioning vector in the panoramic image database, and determining target positioning information corresponding to the target geographic area based on the second panoramic image data;

[0135] The identifier of the panoramic image data in the panoramic image database corresponds to the identifier of the text embedding vector in the preset vector library.

[0136] Specifically, the target positioning information corresponding to the target geographic region can be determined in two ways. In the first way: a confidence threshold can be determined based on the highest confidence value among the confidence values and a preset constant, and the confidence threshold The confidence threshold can be determined by the following formula:

[0137]

[0138] wherein, the highest confidence value among the confidence values, the preset constant, which can be set as needed, for example, can be 0.15.

[0139] Further, the matching vectors can be screened by the confidence threshold, and the matching vectors with confidence values less than the confidence threshold are removed, and the matching vectors with confidence values greater than or equal to the confidence threshold are determined as the first positioning vectors. Further, the identity of the first positioning vector can be obtained and matched with the identity of the panoramic map data in the panoramic map database to determine a plurality of first panoramic map data with the same identity as the first positioning vectors. It should be noted that one first positioning vector corresponds to only one first panoramic map data, and the target positioning information can be determined according to the first panoramic map data corresponding to the first positioning vector with the highest confidence value, or the target positioning information can be determined according to the first panoramic map data corresponding to any first positioning vector. It should be noted that if the confidence values of all matching vectors are less than the confidence threshold, the positioning fails.

[0140] In the second way, the matching vector corresponding to the highest confidence value among the confidence values can be determined as the second positioning vector, and the remaining matching vectors are all removed. The second panoramic map data with the same identity as the second positioning vector is determined in the panoramic map database, and the target positioning information corresponding to the target geographic region is determined based on the second panoramic map data.

[0141] wherein the identity of the panoramic map data in the panoramic map database corresponds to the identity of the text embedding vector in the preset vector library. The panoramic map data can include the identity, positioning information and shooting time corresponding to the panoramic map, so that when the positioning vector has the same identity as the panoramic map data, the positioning information corresponding to the panoramic map data can be determined as the target positioning information.

[0142] In the above embodiments, by efficiently matching vectors, the most similar vector is found from the preset vector library, and the positioning information is queried in the panoramic map database based on the identity of the matching result, so that the target positioning information corresponding to the target geographic region can be accurately determined.

[0143] In one embodiment, the panoramic map database is constructed by the following steps:

[0144] Acquire multiple panoramic images and panoramic image data corresponding to each of the panoramic images; the panoramic image data includes an identifier, positioning information, and shooting time corresponding to the panoramic image;

[0145] The panoramic image database is constructed based on the identifiers, positioning information and shooting times corresponding to all the panoramic images.

[0146] Specifically, multiple panoramas can be obtained from open-source panorama libraries online, such as the Google X Street View library. Panoramas use an equirectangular projection, resulting in the following characteristics: the horizontal image (X-axis) represents a 360-degree panoramic view, covering the yaw angle from -180° to 180°. The vertical image (Y-axis) represents a 180-degree viewing angle, covering the pitch angle from -90° (zenith) to 90° (ground). Horizontally, the angle corresponding to each pixel is uniform, but the vertical angle distribution is distorted, with areas near the zenith and ground appearing stretched.

[0147] The panoramic data corresponding to each panoramic image can also be obtained, and the panoramic data includes the identifier, positioning information and shooting time corresponding to the panoramic image. The positioning information may include longitude and latitude.

[0148] The constructed panoramic image database may be shown in Table 1 below, for example:

[0149] Table 1 Panorama database example table

[0150]

[0151] As shown in Table 1 above, the panorama marked as 123 has a longitude of 112.34 and a latitude of 45.67, and was taken at 12:00 on March 15, 2025.

[0152] In the above embodiment, a panoramic image database is constructed based on the identifier, positioning information and shooting time. The identifier enables the panoramic image to correspond to the text embedding vector, and then based on the correspondence, the target text embedding vector corresponding to the text embedding vector can be determined according to the positioning information in the panoramic image database, and then the target positioning information is determined, which improves the positioning flexibility and accuracy, while the shooting time can improve the user experience.

[0153] In one embodiment, the panoramic image includes a plurality of pixels; and the preset vector library is constructed by the following steps:

[0154] For each of the panoramic images, determining the spherical coordinates corresponding to each pixel based on the original coordinates of each pixel in the panoramic image; determining the pixel color value corresponding to each pixel based on the original coordinates of each pixel in the panoramic image and the pixel color value of the pixel corresponding to each pixel;

[0155] Determining the three-dimensional Cartesian coordinates corresponding to each pixel based on the spherical coordinates corresponding to each pixel;

[0156] Projecting the panoramic image onto a plurality of cubic faces, and determining the cubic face coordinates of each pixel on each cubic face based on the three-dimensional Cartesian coordinates corresponding to each pixel;

[0157] For each of the cube faces, determining, based on the pixel size of the cube face and the cube face coordinates of each of the pixels on the cube face, a texture pixel coordinate of each of the pixels on the cube face; and determining, based on the texture pixel coordinates and pixel color values ​​of each of the pixels on the cube face, a local scene image corresponding to the cube face;

[0158] For each of the panoramic images, multiple local scene images corresponding to the panoramic image are input into the multimodal large model to obtain description texts corresponding to each of the local scene images output by the multimodal large model;

[0159] Inputting each of the description texts corresponding to the panoramic image into the text embedding model to obtain a text embedding vector corresponding to each of the description texts output by the text embedding model;

[0160] The preset vector library is constructed based on the text embedding vectors corresponding to all the panoramic images.

[0161] Specifically, if a semantic description is directly performed on a panoramic image, the following problems will arise: ① Projection deformation causes the shape of objects to be distorted, and the shape and proportion of the objects are distorted, affecting the understanding of the multimodal large model. For example, text (such as street signs) may become unrecognizable in the stretched area. ② The unbalanced perspective reduces the multimodal large model's ability to perceive vertical information. The semantic description may focus too much on horizontal information and ignore important features in the vertical direction. Vertical information may be misinterpreted. For example, the stretched part of the ground may be mistaken for a large area of ​​flat land. ③ The loss of details may cause important positioning information (such as signs and logos) to be ignored. For example, if there is a store sign in the lower right corner of a panoramic image, but it is small and blurry due to projection, the model may ignore it.

[0162] Therefore, it is necessary to convert the equirectangular projection panorama to the cubic projection. First, the spherical coordinates corresponding to each pixel can be determined based on the original coordinates of each pixel in the panorama. The spherical coordinates include the azimuth coordinates and the pitch coordinates. The spherical coordinates corresponding to the pixel are It can be solved by the following formula:

[0163]

[0164]

[0165] in, Indicates the azimuth coordinate, ranging from , represents pi, Indicates the pitch angle coordinate, ranging from , represents the original coordinates of the pixel, Indicates the width of the panorama. Indicates the height of the panorama.

[0166] Then, the three-dimensional Cartesian coordinates corresponding to each pixel can be determined based on the spherical coordinates corresponding to each pixel. The three-dimensional Cartesian coordinates It can be solved by the following formula:

[0167]

[0168]

[0169]

[0170] Among them, it is easy to understand that Control the horizontal direction and determine the X-axis and Z-axis coordinates. Controls the vertical direction and determines the Y-axis coordinate.

[0171] After obtaining the three-dimensional Cartesian coordinates corresponding to each pixel, the panorama can be projected onto multiple cube faces. It is easy to understand that a cube has six faces, namely front, back, left, right, top, and bottom. Each face is a square. The center of the cube can be set as the origin, and the side length can be set to 2, for example. Then the range of pixel coordinates on the cube face is [-1,1]. In order to facilitate projection, it is also necessary to define the center direction vector of each cube face. The center direction vectors of the front (+Z), back (-Z), left (-X), right (+X), top (+Y), and bottom (-Y) are (0,0,1), (0,0,-1), (-1,0,0), (1,0,0), (0,1,0), and (0,-1,0) respectively. When projecting the panorama onto the cube face, it can be projected according to its maximum component (maxComponen). Let maxComponent=max(|X|,|Y|,|Z|). If maxComponent=Z, then project to the front. If maxComponent= , then project it to the back. If maxComponent= , then project to the left. If maxComponent=X, then project to the right. If maxComponent=Y, then project to the top. If maxComponent= , then projected below.

[0172] After projecting the panorama, the cube face coordinates of each pixel on each cube face It can be determined by the following formula:

[0173] Front (+Z): ,

[0174] After (-Z): ,

[0175] Left (-X): ,

[0176] Right (+X): ,

[0177] Up (+Y): ,

[0178] Down (-Y): ,

[0179] Among them, it is easy to understand that The range is [-1,1].

[0180] For each cube face, the texture pixel coordinates of each pixel on the cube face can be determined based on the pixel size of the cube face and the cube face coordinates of each pixel on the cube face. It can be solved by the following formula:

[0181]

[0182]

[0183] in, Indicates the pixel size of the cube face, that is, the cube face includes pixels.

[0184] In addition, the four pixel points corresponding to each pixel in the panoramic image can determine the pixel color value corresponding to each pixel based on the original coordinates of each pixel and the pixel color value of the pixel point corresponding to each pixel. The corresponding pixels are the upper left pixel , upper right pixel , lower left pixel , lower right pixel ,in for The value rounded down, for The value rounded down, the pixel color values ​​corresponding to the four pixels are , , , .

[0185] Furthermore, the pixel color value can be solved by bilinear interpolation , bilinear interpolation method to solve pixel color value This can reduce distortion and achieve a smoother visual effect. In bilinear interpolation, we can first Interpolate in the direction and get , ,in, express The weight of the direction. Interpolate in the direction and get ,in, express The weight of the direction.

[0186] Finally, based on the texture pixel coordinates of each pixel on the cube face and pixel color values The local scene image corresponding to the cube face can be determined.

[0187] Furthermore, for each panoramic image, multiple local scene images corresponding to the panoramic image can be input into the multimodal large model, and the multimodal large model outputs descriptive text corresponding to each local scene image. It should be noted that only the local scene images corresponding to the front, back, left, and right sides can be input into the multimodal large model. The upper view is usually the sky, without obvious landmarks, buildings, or street signs, making it difficult to provide a meaningful semantic description. The sky's color, cloud cover, and other information vary greatly and are not very helpful for positioning. The lower view is mainly the ground, sidewalk, and road, usually only having texture (asphalt, concrete) or shadows, lacking unique positioning features, and may contain shadows of pedestrians or vehicles, which is not very useful for position matching. Eliminating the local scene images corresponding to the upper and lower sides can reduce storage usage and computational complexity, improving positioning efficiency.

[0188] Furthermore, the description texts corresponding to the panoramic image (after excluding the upper and lower parts, that is, one panoramic image corresponds to four partial image images, and one partial image corresponds to one description text) can be input into the text embedding model to obtain the text embedding vectors corresponding to each description text output by the text embedding model. The process of determining the text embedding vector can refer to step 120, and this embodiment will not be repeated. Finally, a preset vector library can be constructed based on the text embedding vectors corresponding to all panoramic images. Exemplarily, the preset vector library can be shown in Table 2 below:

[0189] Table 2 Preset vector library example table

[0190]

[0191] As shown in Table 2 above, the panoramic image identified as UUID-123 corresponds to four text embedding vectors. After the target text embedding vector matches a certain text embedding vector, the panoramic image database can be queried based on the UUID, and the target positioning information corresponding to the target geographic area can be determined based on the query results.

[0192] It should be noted that there are image matching and positioning methods based on panoramic databases in the prior art, but this method mainly relies on the local features of the captured panorama. These local features are sensitive to changes in lighting, occlusion, and perspective, and are prone to ignoring high-level semantic information such as buildings, signs, and roads, resulting in low matching accuracy. In addition, image matching and positioning methods based on panoramic databases usually require searching for the most similar images in a large-scale database, which consumes a lot of storage and computing resources. The technical solution of the present invention converts equirectangular projection into cubic projection to reduce image deformation and improve semantic extraction accuracy. At the same time, only the four most informative perspectives of front, back, left, and right are screened for processing, and efficient vector indexing technology is used to accelerate retrieval, thereby maintaining fast and stable response capabilities in a large-scale data environment.

[0193] In the above embodiment, the identification and positioning information corresponding to the panoramic image is stored. Since the panoramic image uses an equidistant cylindrical projection, which is subject to deformation issues, it needs to be converted to a cubic projection to obtain images from six perspectives: front, back, left, right, top, and bottom. Of the six perspectives of the cubic projection, only the front, back, left, and right perspectives are selected for processing to reduce redundant information and improve analysis efficiency. Subsequently, a multimodal large model is used to generate descriptive text, and a unified preset prompt word template is used for standardization to enhance the consistency and accuracy of the matching. This enables the stable provision of effective location information even in complex geographical environments, improving the applicability of navigation and positioning systems.

[0194] The geolocation method based on the multimodal large model provided by the present invention comprises the following steps: inputting a target image corresponding to the target geographical area into the multimodal large model, obtaining a target description text corresponding to the target image output by the multimodal large model; the multimodal large model is used to generate a target description text based on the target image; inputting the target description text into a text embedding model, obtaining a target text embedding vector corresponding to the target description text output by the text embedding model; matching the target text embedding vector with each text embedding vector in a preset vector library, and determining the target positioning information corresponding to the target geographical area. The technical solution of the present invention does not rely on the local feature points of the image through the connection between images, texts, and vectors, and the matching of the preset vector library with the target text embedding vector, and is therefore not easily affected by the dynamic environment. It can still be stably positioned in a complex environment, and the positioning accuracy is high. In addition, compared with SLAM technology, the present invention can also obtain absolute position information through the matching process.

[0195] The following describes the geo-positioning device based on the multimodal large model provided by the present invention. The geo-positioning device based on the multimodal large model described below and the geo-positioning method based on the multimodal large model described above can refer to each other.

[0196] Figure 2 This is a schematic diagram of the structure of the geo-positioning device based on the multi-modal large model provided by the present invention, such as Figure 2 As shown, the multimodal large model-based geolocation device 200 includes the following modules:

[0197] The text module 210 is configured to input a target image corresponding to a target geographic area into a multimodal large model, and obtain a target description text corresponding to the target image output by the multimodal large model; the multimodal large model is configured to generate the target description text based on the target image;

[0198] A vector module 220 is configured to input the target description text into a text embedding model to obtain a target text embedding vector corresponding to the target description text output by the text embedding model;

[0199] The positioning module 230 is configured to match the target text embedding vector with each text embedding vector in a preset vector library to determine target positioning information corresponding to the target geographical area.

[0200] In one embodiment, the vector module 220 is specifically configured to:

[0201] Inputting the target description text into a large language model to obtain a standardized text output by the large language model; the large language model is used to perform standardization processing on the target description text; the standardization processing includes vocabulary unification processing, sentence unification processing, removal of irrelevant characters, data format unification processing and stop word filtering processing;

[0202] The standardized text is input into the text embedding model to obtain the target text embedding vector output by the text embedding model; the text embedding model is used to generate the target text embedding vector after executing a long text splitting strategy, a word segmentation strategy and an embedding strategy on the standardized text.

[0203] In one embodiment, the positioning module 230 is specifically configured to:

[0204] For each of the text embedding vectors in the preset vector library, calculating a similarity value between the target text embedding vector and the text embedding vector;

[0205] A first preset number of text embedding vectors sorted in descending order according to similarity values ​​are determined as matching vectors, and the target positioning information is determined according to each of the matching vectors.

[0206] In one embodiment, the positioning module 230 is further configured to:

[0207] For each of the matching vectors, determining a relative similarity value corresponding to the matching vector based on the similarity value of the matching vector, a highest matching vector with the highest similarity value among all the matching vectors, and a lowest matching vector with the lowest similarity value;

[0208] Determining a confidence value corresponding to the matching vector based on a sum of the relative similarity values ​​of all the matching vectors and the relative similarity values ​​corresponding to the matching vectors;

[0209] The target positioning information is determined based on the confidence value corresponding to each matching vector.

[0210] In one embodiment, the positioning module 230 is further configured to:

[0211] Determining a confidence threshold based on the highest confidence value among the confidence values ​​and a preset constant; determining a matching vector having a confidence value greater than or equal to the confidence threshold as a first positioning vector; determining a plurality of first panoramic image data having the same identifier as each of the first positioning vectors in a panoramic image database, and determining the target positioning information corresponding to the target geographic area based on each of the first panoramic image data; or,

[0212] determining a matching vector corresponding to a highest confidence value among the confidence values ​​as a second positioning vector, determining second panoramic image data having the same identifier as the second positioning vector in the panoramic image database, and determining target positioning information corresponding to the target geographic area based on the second panoramic image data;

[0213] The identifier of the panoramic image data in the panoramic image database corresponds to the identifier of the text embedding vector in the preset vector library.

[0214] In one embodiment, the geo-positioning device based on the multimodal large model further includes a first building module, which is specifically configured to:

[0215] Acquire multiple panoramic images and panoramic image data corresponding to each of the panoramic images; the panoramic image data includes an identifier, positioning information, and shooting time corresponding to the panoramic image;

[0216] The panoramic image database is constructed based on the identifiers, positioning information and shooting times corresponding to all the panoramic images.

[0217] In one embodiment, the panoramic image includes a plurality of pixels; the multimodal large model-based geo-positioning device further includes a second construction module, the second construction module being specifically configured to:

[0218] For each of the panoramic images, determining the spherical coordinates corresponding to each pixel based on the original coordinates of each pixel in the panoramic image; determining the pixel color value corresponding to each pixel based on the original coordinates of each pixel in the panoramic image and the pixel color value of the pixel corresponding to each pixel;

[0219] Determining the three-dimensional Cartesian coordinates corresponding to each pixel based on the spherical coordinates corresponding to each pixel;

[0220] Projecting the panoramic image onto a plurality of cubic faces, and determining the cubic face coordinates of each pixel on each cubic face based on the three-dimensional Cartesian coordinates corresponding to each pixel;

[0221] For each of the cube faces, determining, based on the pixel size of the cube face and the cube face coordinates of each of the pixels on the cube face, a texture pixel coordinate of each of the pixels on the cube face; and determining, based on the texture pixel coordinates and pixel color values ​​of each of the pixels on the cube face, a local scene image corresponding to the cube face;

[0222] For each of the panoramic images, multiple local scene images corresponding to the panoramic image are input into the multimodal large model to obtain description texts corresponding to each of the local scene images output by the multimodal large model;

[0223] Inputting each of the description texts corresponding to the panoramic image into the text embedding model to obtain a text embedding vector corresponding to each of the description texts output by the text embedding model;

[0224] The preset vector library is constructed based on the text embedding vectors corresponding to all the panoramic images.

[0225] The present invention provides a geographic positioning device based on a multimodal large model. The target image corresponding to the target geographic area is input into the multimodal large model to obtain a target description text corresponding to the target image output by the multimodal large model; the multimodal large model is used to generate a target description text based on the target image; the target description text is input into a text embedding model to obtain a target text embedding vector corresponding to the target description text output by the text embedding model; the target text embedding vector is matched with each text embedding vector in a preset vector library to determine the target positioning information corresponding to the target geographic area. The technical solution of the present invention does not rely on the local feature points of the image through the connection between images, texts, and vectors, and the matching of the preset vector library with the target text embedding vector. Therefore, it will not be easily affected by the dynamic environment. It can still be stably positioned in a complex environment, and the positioning accuracy is high. In addition, compared with SLAM technology, the present invention can also obtain absolute position information through the matching process.

[0226] Figure 3 An example of a physical structure diagram of an electronic device is shown below. Figure 3 As shown, the electronic device may include: a processor 310, a communications interface 320, a memory 330, and a communication bus 340. The processor 310, the communications interface 320, and the memory 330 communicate with each other via the communication bus 340. The processor 310 may call logic instructions in the memory 330 to execute a geolocation method based on a multimodal large model.

[0227] In addition, the logic instructions in the memory 330 described above can be implemented in the form of a software function unit and sold or used as an independent product, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0228] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to execute the multi-modal large model based geolocation method provided by the above-mentioned methods.

[0229] In yet another aspect, the present application also provides a non-transitory computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the multi-modal large model based geolocation method provided by the above-mentioned methods.

[0230] The device embodiments described above are only schematic, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., they can be located in one place, or distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment. Those skilled in the art can understand and implement without creative labor.

[0231] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be implemented by means of software plus necessary universal hardware platforms, and of course, can also be implemented by hardware. Based on such understanding, the above technical solutions essentially or the parts that contribute to the prior art can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments.

[0232] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit the present application; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalent features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A geolocation method based on a multimodal large model, characterized in that: include: Inputting a target image corresponding to a target geographical area into a multimodal large model, and obtaining a target description text corresponding to the target image output by the multimodal large model; The multimodal large model is used to generate the target description text based on the target image; Inputting the target description text into a text embedding model to obtain a target text embedding vector corresponding to the target description text output by the text embedding model; Matching the target text embedding vector with each text embedding vector in a preset vector library to determine target positioning information corresponding to the target geographic area; The step of matching the target text embedding vector with each text embedding vector in a preset vector library to determine target positioning information corresponding to the target geographical area includes: For each of the text embedding vectors in the preset vector library, calculating a similarity value between the target text embedding vector and the text embedding vector; Determine a first preset number of text embedding vectors sorted in descending order according to similarity values ​​as matching vectors, and determine the target positioning information based on each of the matching vectors; The determining the target positioning information according to each matching vector includes: For each of the matching vectors, determining a relative similarity value corresponding to the matching vector based on the similarity value of the matching vector, a highest matching vector with the highest similarity value among all the matching vectors, and a lowest matching vector with the lowest similarity value; Determining a confidence value corresponding to the matching vector based on a sum of the relative similarity values ​​of all the matching vectors and the relative similarity values ​​corresponding to the matching vectors; Determining the target positioning information based on a confidence threshold and a confidence value corresponding to each matching vector; the confidence threshold is determined based on the confidence value corresponding to each matching vector and a preset constant; The preset vector library is constructed by the following steps: For each panoramic image, determining the spherical coordinates corresponding to each pixel based on the original coordinates of each pixel in the panoramic image; determining the pixel color values ​​corresponding to each pixel based on the original coordinates of each pixel in the panoramic image and the pixel color values ​​of the pixel points corresponding to each pixel; wherein each panoramic image is used to construct a panoramic image database, and the panoramic image database is used to determine the target positioning information; Determining the three-dimensional Cartesian coordinates corresponding to each pixel based on the spherical coordinates corresponding to each pixel; Projecting the panoramic image onto a plurality of cubic faces, and determining the cubic face coordinates of each pixel on each cubic face based on the three-dimensional Cartesian coordinates corresponding to each pixel; For each of the cube faces, determining, based on the pixel size of the cube face and the cube face coordinates of each of the pixels on the cube face, a texture pixel coordinate of each of the pixels on the cube face; and determining, based on the texture pixel coordinates and pixel color values ​​of each of the pixels on the cube face, a local scene image corresponding to the cube face; For each of the panoramic images, multiple local scene images corresponding to the panoramic image are input into the multimodal large model to obtain description texts corresponding to each of the local scene images output by the multimodal large model; Inputting each of the description texts corresponding to the panoramic image into the text embedding model to obtain a text embedding vector corresponding to each of the description texts output by the text embedding model; The preset vector library is constructed based on the text embedding vectors corresponding to all the panoramic images.

2. The multimodal large model-based geolocation method according to claim 1, characterized in that: The determining the target positioning information based on the confidence threshold and the confidence value corresponding to each matching vector includes: Determine the confidence threshold based on the highest confidence value among the confidence values ​​and a preset constant; determine the matching vector with the confidence value greater than or equal to the confidence threshold as the first positioning vector; determine a plurality of first panoramic image data with the same identifier as each of the first positioning vectors in the panoramic image database, and determine the target positioning information corresponding to the target geographical area based on each of the first panoramic image data; or determining a matching vector corresponding to a highest confidence value among the confidence values ​​as a second positioning vector, determining second panoramic image data having the same identifier as the second positioning vector in the panoramic image database, and determining target positioning information corresponding to the target geographic area based on the second panoramic image data; The identifier of the panoramic image data in the panoramic image database corresponds to the identifier of the text embedding vector in the preset vector library.

3. The geolocation method based on a multimodal large model according to claim 2, characterized in that: The panoramic image database is constructed by the following steps: Acquire multiple panoramic images and panoramic image data corresponding to each of the panoramic images; the panoramic image data includes an identifier, positioning information, and shooting time corresponding to the panoramic image; The panoramic image database is constructed based on the identifiers, positioning information and shooting times corresponding to all the panoramic images.

4. The geolocation method based on a multimodal large model according to any one of claims 1 to 3, characterized in that: Inputting the target description text into a text embedding model to obtain a target text embedding vector corresponding to the target description text output by the text embedding model includes: Inputting the target description text into a large language model to obtain a standardized text output by the large language model; the large language model is used to perform standardization processing on the target description text; the standardization processing includes vocabulary unification processing, sentence unification processing, removal of irrelevant characters, data format unification processing and stop word filtering processing; The standardized text is input into the text embedding model to obtain the target text embedding vector output by the text embedding model; the text embedding model is used to generate the target text embedding vector after executing a long text splitting strategy, a word segmentation strategy and an embedding strategy on the standardized text.

5. A geolocation device based on a multimodal large model, characterized in that: include: A text module, configured to input a target image corresponding to a target geographical area into a multimodal large model, and obtain a target description text corresponding to the target image output by the multimodal large model; The multimodal large model is used to generate the target description text based on the target image; A vector module, configured to input the target description text into a text embedding model to obtain a target text embedding vector corresponding to the target description text output by the text embedding model; A positioning module, configured to match the target text embedding vector with each text embedding vector in a preset vector library to determine target positioning information corresponding to the target geographic area; The positioning module is specifically used for: For each of the text embedding vectors in the preset vector library, calculating a similarity value between the target text embedding vector and the text embedding vector; Determine a first preset number of text embedding vectors sorted in descending order according to similarity values ​​as matching vectors, and determine the target positioning information based on each of the matching vectors; The positioning module is further specifically used for: For each of the matching vectors, determining a relative similarity value corresponding to the matching vector based on the similarity value of the matching vector, a highest matching vector with the highest similarity value among all the matching vectors, and a lowest matching vector with the lowest similarity value; Determining a confidence value corresponding to the matching vector based on a sum of the relative similarity values ​​of all the matching vectors and the relative similarity values ​​corresponding to the matching vectors; Determining the target positioning information based on a confidence threshold and a confidence value corresponding to each matching vector; the confidence threshold is determined based on the confidence value corresponding to each matching vector and a preset constant; A second construction module is configured to determine, for each panoramic image, the spherical coordinates corresponding to each pixel in the panoramic image based on the original coordinates of each pixel in the panoramic image; and to determine the pixel color values ​​corresponding to each pixel in the panoramic image based on the original coordinates of each pixel in the panoramic image and the pixel color values ​​of the pixel points corresponding to each pixel; wherein each panoramic image is used to construct a panoramic image database, and the panoramic image database is used to determine the target positioning information; Determining the three-dimensional Cartesian coordinates corresponding to each pixel based on the spherical coordinates corresponding to each pixel; Projecting the panoramic image onto a plurality of cubic faces, and determining the cubic face coordinates of each pixel on each cubic face based on the three-dimensional Cartesian coordinates corresponding to each pixel; For each of the cube faces, determining, based on the pixel size of the cube face and the cube face coordinates of each of the pixels on the cube face, a texture pixel coordinate of each of the pixels on the cube face; and determining, based on the texture pixel coordinates and pixel color values ​​of each of the pixels on the cube face, a local scene image corresponding to the cube face; For each of the panoramic images, multiple local scene images corresponding to the panoramic image are input into the multimodal large model to obtain description texts corresponding to each of the local scene images output by the multimodal large model; Inputting each of the description texts corresponding to the panoramic image into the text embedding model to obtain a text embedding vector corresponding to each of the description texts output by the text embedding model; The preset vector library is constructed based on the text embedding vectors corresponding to all the panoramic images.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the geo-positioning method based on the multimodal large model as claimed in any one of claims 1 to 4 is implemented.

7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the geo-positioning method based on a multimodal large model as claimed in any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Image cross-modal retrieval method and device based on large language model and medium

    CN119938972A