Geographic positioning method and device based on multi-modal large model, equipment and medium

The multimodal big model generates text embedding vectors to match the preset vector library, which solves the problem of positioning instability of SLAM technology in complex environments and achieves high-precision absolute geolocation.

CN120256489AActive Publication Date: 2025-07-04北京观微科技有限公司

Patent Information

Application Number
CN202510751681.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-07-04
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

Existing SLAM technologies are difficult to perform stable and reliable positioning in complex environments, especially in large-scale, dynamic, low-texture and extreme lighting scenarios, which are susceptible to interference from moving objects and cannot obtain absolute geographical location.

Method used

The multimodal big model combines image, text and vector matching method. By inputting the target image into the multimodal big model to generate description text, the text embedding model is used to convert it into text embedding vectors, and match it with the preset vector library to determine the positioning information of the target geographical area.

Benefits of technology

Achieve stable and high-precision positioning in complex environments and obtain absolute position information, avoiding dependence on local feature points of the image and improving positioning stability and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256489A_ABST
    Figure CN120256489A_ABST
Patent Text Reader

Abstract

The invention provides a geographic positioning method, device and equipment based on a multi-modal large model and a medium, and relates to the technical field of geographic positioning, the method comprises the following steps: inputting a target picture corresponding to a target geographic area into the multi-modal large model to obtain a target description text corresponding to the target picture output by the multi-modal large model; the multi-modal large model is used for generating a target description text based on a target picture; inputting the target description text into a text embedding model to obtain a target text embedding vector corresponding to the target description text output by the text embedding model; and matching the target text embedding vector with each text embedding vector in a preset vector library, and determining target positioning information corresponding to the target geographic area. The method does not depend on local feature points of the image, so that the method is not easily influenced by a dynamic environment, positioning can be stably carried out in a complex environment, and the positioning precision is high. And compared with a synchronous positioning and mapping technology, absolute position information can be obtained through a matching process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of geolocation, and particularly to a geolocation method, device, equipment and medium based on a multimodal large model. Background Art

[0002] Geolocation technology is widely used in fields such as navigation, autonomous driving, augmented reality, robot positioning, and geographic information systems.

[0003] The existing geolocation mainly adopts the following two methods: (1) Simultaneous Localization and Mapping (SLAM), that is, enabling devices such as robots to construct an environmental map in real time through sensor data and determine their own positions; (2) an image matching and positioning method based on a panoramic image database, that is, taking a photo and performing feature matching with the street view images in the database to determine the shooting position.

[0004] However, SLAM technology still has many limitations in complex scenarios such as large-scale, dynamic, low-texture, and extreme lighting. First of all, it relies on local feature points of images, has poor adaptability to dynamic environments, and is easily interfered by moving objects such as pedestrians and vehicles, resulting in tracking drift or failure. Secondly, in low-texture or scenes with a large number of repetitive structures (such as white walls, glass curtain walls, corridors, etc.), due to the lack of stable feature points, it is difficult to perform reliable tracking and positioning. In addition, SLAM technology obtains relative position information based on relative motion and cannot directly obtain absolute geographical locations (such as longitude and latitude). Summary of the Invention

[0005] The present invention provides a geolocation method, device, equipment and medium based on a multimodal large model to solve the defect that the existing SLAM technology in the prior art cannot perform stable and reliable positioning in complex environments. The technical solution of the present invention enables stable positioning in complex environments with relatively high positioning accuracy through the connection of images, texts, and vectors, as well as the matching of a preset vector library and the target text embedding vector.

[0006] The present invention provides a geolocation method based on a multimodal large model, including the following steps.

[0007] Input a target picture corresponding to a target geographical area into the multimodal large model to obtain a target description text corresponding to the target picture output by the multimodal large model; the multimodal large model is used to generate the target description text based on the target picture; Input the target description text into a text embedding model to obtain a target text embedding vector corresponding to the target description text output by the text embedding model; Match the target text embedding vector with each text embedding vector in the preset vector library to determine the target location information corresponding to the target geographical area.

[0008] According to a geographical location method based on a multimodal large model provided by the present invention, the matching the target text embedding vector with each text embedding vector in the preset vector library to determine the target location information corresponding to the target geographical area includes: For each text embedding vector in the preset vector library, calculate the similarity value between the target text embedding vector and the text embedding vector; Determine the first preset number of text embedding vectors sorted in descending order of similarity value as matching vectors, and determine the target location information according to each of the matching vectors.

[0009] According to a geographical location method based on a multimodal large model provided by the present invention, the determining the target location information according to each of the matching vectors includes: For each of the matching vectors, determine the relative similarity value corresponding to the matching vector based on the similarity value of the matching vector, the highest matching vector with the highest similarity value among all the matching vectors, and the lowest matching vector with the lowest similarity value; Based on the sum of the relative similarity values of all the matching vectors and the relative similarity value corresponding to the matching vector, determine the confidence value corresponding to the matching vector; Determine the target location information based on the confidence values corresponding to each of the matching vectors.

[0010] According to a geographical location method based on a multimodal large model provided by the present invention, the determining the target location information based on the confidence values corresponding to each of the matching vectors includes: Based on the highest confidence value among the confidence values and a preset constant, determine a confidence threshold; determine the matching vectors with confidence values greater than or equal to the confidence threshold as the first positioning vectors; determine a plurality of first panoramic image data with the same identifier as each of the first positioning vectors in the panoramic image database, and determine the target location information corresponding to the target geographical area based on each of the first panoramic image data; or, Determine the matching vector corresponding to the highest confidence value among the confidence values as the second positioning vector, determine the second panoramic image data with the same identifier as the second positioning vector in the panoramic image database, and determine the target location information corresponding to the target geographical area based on the second panoramic image data; Wherein, the identifier of the panoramic image data in the panoramic image database corresponds to the identifier of the text embedding vector in the preset vector library.

[0011] A geolocation method based on a multimodal large model provided by the present invention, the panoramic image database is constructed through the following steps: Obtain multiple panoramic images, and the panoramic image data corresponding to each of the panoramic images; the panoramic image data includes the identifier, location information, and shooting time corresponding to the panoramic image; Construct the panoramic image database based on the identifiers, location information, and shooting times corresponding to all the panoramic images.

[0012] A geolocation method based on a multimodal large model provided by the present invention, the panoramic image includes multiple pixels; the preset vector library is constructed through the following steps: For each of the panoramic images, determine the spherical coordinates corresponding to each of the pixels based on the original coordinates of the pixels in the panoramic image; determine the pixel color values corresponding to each of the pixels based on the original coordinates of the pixels in the panoramic image and the pixel color values of the pixel points corresponding to the pixels; Determine the three-dimensional Cartesian coordinates corresponding to each of the pixels based on the spherical coordinates corresponding to each of the pixels; Project the panoramic image onto multiple cubic surfaces, and determine the cubic surface coordinates of each of the pixels on each of the cubic surfaces based on the three-dimensional Cartesian coordinates corresponding to each of the pixels; For each of the cubic surfaces, determine the mapped pixel coordinates of each of the pixels on the cubic surface based on the pixel size of the cubic surface and the cubic surface coordinates of each of the pixels on the cubic surface; determine the local scene image corresponding to the cubic surface based on the mapped pixel coordinates and pixel color values of each of the pixels on the cubic surface; For each of the panoramic images, input the multiple local scene images corresponding to the panoramic image into the multimodal large model to obtain the description text corresponding to each of the local scene images output by the multimodal large model; Input the description texts corresponding to the panoramic image into the text embedding model to obtain the text embedding vectors corresponding to the description texts output by the text embedding model; Construct the preset vector library based on the text embedding vectors corresponding to all the panoramic images.

[0013] A geolocation method based on a multimodal large model provided by the present invention, the inputting the target description text into the text embedding model to obtain the target text embedding vector corresponding to the target description text output by the text embedding model includes: Input the target description text into a large language model to obtain the standardized text output by the large language model; the large language model is used to perform standardization processing on the target description text; the standardization processing includes vocabulary unification processing, sentence pattern unification processing, removal of irrelevant characters, data format unification processing, and stop word filtering processing; Input the standardized text into the text embedding model to obtain the target text embedding vector output by the text embedding model; the text embedding model is used to generate the target text embedding vector after performing a long text splitting strategy, a word segmentation strategy, and an embedding strategy on the standardized text.

[0014] The present invention also provides a geographic positioning device based on a multimodal large model, including the following modules: A text module, configured to input a target picture corresponding to a target geographic area into a multimodal large model to obtain the target description text corresponding to the target picture output by the multimodal large model; the multimodal large model is used to generate the target description text based on the target picture; A vector module, configured to input the target description text into a text embedding model to obtain the target text embedding vector corresponding to the target description text output by the text embedding model; A positioning module, configured to match the target text embedding vector with each text embedding vector in a preset vector library to determine the target positioning information corresponding to the target geographic area.

[0015] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, it implements the geographic positioning method based on a multimodal large model as described in any one of the above.

[0016] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the geographic positioning method based on a multimodal large model as described in any one of the above.

[0017] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the geographic positioning method based on a multimodal large model as described in any one of the above.

[0018] The geolocation method, device, equipment and medium based on a multimodal large model provided by the present invention input a target picture corresponding to a target geographical area into the multimodal large model to obtain a target description text corresponding to the target picture output by the multimodal large model; the multimodal large model is used to generate the target description text based on the target picture; input the target description text into a text embedding model to obtain a target text embedding vector corresponding to the target description text output by the text embedding model; match the target text embedding vector with each text embedding vector in a preset vector library to determine the target location information corresponding to the target geographical area. The technical solution of the present invention, through the connection of images, texts and vectors, and the matching of the preset vector library with the target text embedding vector, does not rely on local feature points of the image, so it is not easily affected by the dynamic environment, and can still stably perform positioning in a complex environment with relatively high positioning accuracy. In addition, compared with the SLAM technology, the present invention can also obtain absolute position information through the matching process. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 is a schematic flowchart of the geolocation method based on a multimodal large model provided by the present invention.

[0021] Figure 2 is a schematic structural diagram of the geolocation device based on a multimodal large model provided by the present invention.

[0022] Figure 3 is a schematic structural diagram of the electronic equipment provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] To make the objectives, technical solutions and advantages of the present invention clearer, the following clearly and completely describes the technical solutions in the present invention with reference to the drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.

[0024] In view of the above problems in the prior art, the present invention provides a geolocation method based on a multimodal large model. Figure 1 is a schematic flowchart of the geolocation method based on a multimodal large model provided by the present invention, as Figure 1As shown, the method includes the following steps 110 to 130.

[0025] Step 110: Input the target picture corresponding to the target geographical area into the multimodal large model to obtain the target description text corresponding to the target picture output by the multimodal large model; the multimodal large model is used to generate the target description text based on the target picture.

[0026] Specifically, the target geographical area can be any geographical area, for example, it can be Square A in City A. The target picture is a picture taken of the target geographical area. For example, it can be a picture taken with a mobile phone, a camera, or other devices with a camera function, and the target picture should be horizontal. Exemplarily, when it is necessary to locate a certain square, a photo of the square can be taken as the target picture. The multimodal large model can be, for example, GLM-4V, or other multimodal large models that can generate description text according to pictures. The embodiments of the present invention do not make specific limitations here. After inputting the target picture into the multimodal large model, the multimodal large model can generate the target description text based on the target picture and the preset prompt word template, and then obtain the target description text corresponding to the target picture output by the multimodal large model. This target description text is used to describe the characteristics of the target picture.

[0027] It should be noted that due to the following problems in the multimodal large model: due to the randomness of language generation, the description methods are different; due to differences in focus, key information may be omitted; due to lexical diversity, there are different expressions for the same thing; and due to environmental factors (such as time, weather, and lighting changes), etc., the description content is different. Therefore, the description text generated by the multimodal large model for the same picture may be different. Therefore, the preset prompt word template can be set in advance according to needs. Further, in the application process, the preset prompt word template and the target picture can be input into the multimodal large model synchronously.

[0028] Exemplarily, the preset prompt word template can be as follows: (1) Buildings: Appearance description: {The shape, material, and color of the building, such as "a three-story building with a red brick structure"}.

[0029] Building use: {The function of the building, such as "the ground floor is a store, and the upper floors are residences"}.

[0030] Building location: {The relative position of the building in the picture, such as "located in the center of the picture, adjacent to a main road on the left"}.

[0031] (2) Signboards: Signboard content: {The complete text on the signboard, such as "100 meters of XX Road"}.

[0032] Location: {The relative position of the signboard in the picture, e.g., "Located in the lower right corner, near the sidewalk".}

[0033] (3)Roads and Greenery: Road Description: {The type, width, material, and markings of the road, e.g., "Two-way four-lane asphalt road with a yellow center line".}

[0034] Green Belt: {The type of vegetation in the green belt, e.g., "There is a row of plane trees beside the sidewalk".}

[0035] (4)Other Fixed Facilities: Infrastructure: {Such as bus stops, benches, and mailboxes, e.g., "Bus stop with a sign saying 'XX Station'".}

[0036] Please strictly follow the above format for describing the picture to ensure the content is complete and key information is not omitted. If a certain item is missing, please fill in "None".

[0037] The target text description A generated by the multimodal large model based on the target picture and the preset prompt template of the above example can be as follows: (1)Buildings: Appearance Description: The buildings are mainly composed of glass and concrete. The building on the left is cylindrical, and the building on the right is rectangular. The glass windows reflect the yellow light, giving a modern and technological feeling.

[0038] Building Use: Judging from the lights on inside the windows, these buildings may be office buildings or commercial spaces.

[0039] Building Location: The buildings are located in an open square, surrounded by other buildings, and it seems to be in the commercial area of a city.

[0040] (2)Signboards: Signboard Content: There are no obvious signboards in the picture.

[0041] Location: There are no signboards in the picture.

[0042] (3)Roads and Greenery: Road Description: The roads in the square are paved with gray bricks, of moderate width, and there are no obvious traffic markings.

[0043] Green Belt: There is no obvious green belt in the square.

[0044] (4)Other Fixed Facilities: Other Infrastructure: There are several street lights and several trash cans in the square. On the right side of the picture, a parked white van can be seen.

[0045] Step 120: Input the target description text into the text embedding model to obtain the target text embedding vector corresponding to the target description text output by the text embedding model.

[0046] Specifically, after obtaining the target description text, the target description text can be input into the text embedding model. Then, the text embedding model can process the target description text and output the target text embedding vector corresponding to the target description text. For example, the text embedding model can be the bge-base-1.5v-zh embedding model, which is a Chinese text embedding model developed by the Beijing Academy of Artificial Intelligence. The text embedding model can also be other embedding models, and the embodiments of the present invention do not make specific limitations here.

[0047] In one embodiment, the step of inputting the target description text into the text embedding model to obtain the target text embedding vector corresponding to the target description text output by the text embedding model includes: Input the target description text into the large language model to obtain the standardized text output by the large language model; the large language model is used to perform standardization processing on the target description text; the standardization processing includes vocabulary unification processing, sentence pattern unification processing, removal of irrelevant characters, data format unification processing, and stop word filtering processing; Input the standardized text into the text embedding model to obtain the target text embedding vector output by the text embedding model; the text embedding model is used to generate the target text embedding vector after performing long text splitting strategy, word segmentation strategy, and embedding strategy on the standardized text.

[0048] Specifically, to improve the consistency of the target description text, standardization processing can also be performed on the target description text. For example, the target description text can be input into the large language model, and the large language model performs standardization processing on the target description text to obtain the standardized text output by the large language model. The standardization processing can include: (1) Vocabulary unification processing: Replace synonyms, simplify redundant descriptions, and make the expression more consistent. For example, "tall building" → "building", "office building" → "office building".

[0049] (2) Sentence pattern unification processing: Unify the description structure to make different texts have the same grammatical format. For example, "This is a building with a glass curtain wall" → "Building: Glass curtain wall structure".

[0050] (3) Removal of irrelevant characters: Delete redundant punctuation, spaces, etc. to ensure the cleanliness of the text. For example, "Building: Glass curtain wall." → "Building: Glass curtain wall".

[0051] (4)Unified processing of data format: Unify the measurement units to ensure data consistency. For example, "5m" → "5 meters", "ten meters" → "10 meters".

[0052] (5)Stop word filtering: Remove words that affect semantic matching, such as fuzzy words like "may", "seem", etc. For example, "This may be an office building" → "Building: Office building".

[0053] Exemplarily, the standardized text A obtained after standardizing the target text description A in the above example can be as follows: (1)Building: Appearance description: Glass and concrete structure. The left building is cylindrical, and the right building is rectangular. The glass windows reflect yellow light, and the overall style is modern and technological.

[0054] Purpose of the building: There are lights on inside the windows, presumably an office building or commercial space.

[0055] Location of the building: Located in an open square, surrounded by other buildings, probably belonging to the urban commercial area.

[0056] (2)Sign: Content of the sign: No obvious sign.

[0057] Location: No sign.

[0058] (3)Road and greening: Road description: The square road is paved with gray bricks, of moderate width, without obvious traffic markings.

[0059] Green belt: No obvious green belt.

[0060] (4)Other fixed facilities: Infrastructure: The square is equipped with street lights and trash cans. There is a parked white van on the right side of the picture.

[0061] The specific adjustment content from the target description text A to the standardized text A in the above example is as follows: Unified processing of vocabulary: such as "The building is mainly composed of glass and concrete" → "Glass and concrete structure"; Unified processing of sentence patterns: Remove redundant modifiers, such as "Giving a modern and technological feeling" → "The overall style is modern and technological"; Remove irrelevant characters: Delete irrelevant words such as "Looks like", "Can be seen"; Unified processing of data format: No numerical values need to be adjusted; Stop word filtering: Remove fuzzy words such as "may", "seem".

[0062] Further, after obtaining the standardized text, the standardized text can be input into a text embedding model, which can perform long text splitting strategy, word segmentation strategy, and embedding strategy on the standardized text, and then generate a target text embedding vector.

[0063] Among them, in the long text splitting strategy, it can be split according to semantic units, fixed length, or combined with the theme. For example, in the embodiments of the present invention, it can be split by fixed length, and every 800 words are split into a text block. The last 200 words of the previous text block and the first 200 words of the next text block are the overlapping information of the context. Usually, the text description of a picture is about 350 words, and basically does not exceed 400 words. Therefore, splitting every 800 words into a text block ensures that each text block can independently cover all the information of a picture.

[0064] In the word segmentation strategy, the context features of the text can be learned through a neural network to dynamically adjust the word segmentation strategy. The sub-model in the word segmentation strategy can use SentencePiece, which is a word segmentation tool that can decompose text into smaller sub-word units, and these sub-word units have certain frequencies and statistical meanings in the dictionary. The advantages of SentencePiece are: suitable for long Chinese texts, can handle complex syntactic structures, support out-of-vocabulary words, can avoid information loss, and can also maintain semantic integrity and improve the quality of downstream tasks (such as retrieval, dialogue generation).

[0065] The embedding strategy can include direct embedding, segmented embedding and aggregation, key content extraction and embedding, and multi-granularity embedding fusion. For example, in the embodiments of the present invention, direct embedding can be performed, and the text is input into bge-base-1.5v-zh based on the Transformer structure. bge-base-1.5v-zh adopts a multi-layer self-attention mechanism to globally model the text and can capture long-distance dependencies.

[0066] In the above embodiments, by standardizing the target description text, the description accuracy of the text is higher. Further, through the text embedding model, the standardized text can be converted into a text embedding vector that retains the corresponding features, so that subsequent matching can be performed according to the text embedding vector, improving the accuracy of matching and positioning. And by analyzing the core visual information such as buildings, signs, and roads, the system can accurately infer the actual geographical location of the photo and achieve high-precision positioning in urban environments or unmarked areas.

[0067] Step 130: Match the target text embedding vector with each text embedding vector in the preset vector library to determine the target positioning information corresponding to the target geographical area.

[0068] Specifically, the target text embedding vector may be matched with each text embedding vector in a preset vector library, and the target positioning information corresponding to the target geographic area may be determined according to the text embedding vector that matches the target text embedding vector.

[0069] In one embodiment, matching the target text embedding vector with each text embedding vector in a preset vector library to determine the target positioning information corresponding to the target geographical area includes: For each of the text embedding vectors in the preset vector library, calculating a similarity value between the target text embedding vector and the text embedding vector; A first preset number of text embedding vectors sorted in descending order according to similarity values ​​are determined as matching vectors, and the target positioning information is determined according to each of the matching vectors.

[0070] Specifically, for each text embedding vector in the preset vector library, the cosine similarity can be used to calculate the similarity value between the target text embedding vector and the text embedding vector. The similarity value corresponding to the text embedding vector It can be calculated by the following formula:

[0071] in, represents the target text embedding vector, Indicates the number The text embedding vector of .

[0072] Furthermore, all text embedding vectors may be sorted from large to small according to the corresponding similarity values, and then a first preset number of text embedding vectors with a top ranking may be determined as matching vectors, and the first preset number may be set as required, for example, 5 to 10. It should be noted that each text embedding vector has a corresponding identifier, and each matching vector also has a corresponding identifier. The positioning information corresponding to each matching vector may be determined according to the corresponding identifier, and then the target positioning information of the target geographic area may be determined according to the positioning information corresponding to each matching vector.

[0073] In the above embodiment, the target text embedding vector is matched with the text embedding vector in the preset vector library according to the similarity value, and then the target positioning information can be determined according to the matched text embedding vector, thereby getting rid of the dependence on the global positioning system in the traditional positioning method, and reliable positioning can be performed even in a signal-limited environment (indoors or densely populated urban areas).

[0074] In one embodiment, determining the target positioning information according to each of the matching vectors includes: For each of the matching vectors, a relative similarity value corresponding to the matching vector is determined based on the similarity value of the matching vector, the highest matching vector with the highest similarity value among all the matching vectors, and the lowest matching vector with the lowest similarity value; Based on the sum of the relative similarity values of all the matching vectors and the relative similarity value corresponding to the matching vector, a confidence value corresponding to the matching vector is determined; Based on the confidence values corresponding to the respective matching vectors, the target location information is determined.

[0075] Specifically, for each matching vector, the relative similarity value of the matching vector can be determined based on the similarity value of the matching vector, the highest matching vector corresponding to the highest similarity value, and the lowest matching vector corresponding to the lowest similarity value. The process of solving the relative similarity value is essentially a process of normalizing the similarity value. The relative similarity value can ensure comparability between different matching vectors. The relative similarity value corresponding to the matching vector numbered can be expressed by the following formula: can be represented by the following formula:

[0076] where, represents the lowest matching vector, represents the highest matching vector.

[0077] Furthermore, the relative similarity value can measure the matching degree between two vectors, but cannot measure the reliability of the result. Therefore, the confidence value can be further calculated based on the relative similarity value. The confidence value corresponding to the matching vector numbered can be expressed by the following formula: can be represented by the following formula:

[0078] where, represents the sum of the relative similarity values of all the matching vectors, ranges from (0, 1]. It is easy to understand that the larger is, the larger the proportion of the matching vector numbered

[0079] After determining the confidence values corresponding to the respective matching vectors, the target location information can be determined based on the confidence values corresponding to the matching vectors.

[0080] In the above embodiment, calculating the relative similarity value can ensure comparability between different matching vectors. The further determined confidence value can reduce the risk of misjudgment and improve the matching accuracy in complex scenarios (such as image matching and long text retrieval).

[0081] In one embodiment, determining the target positioning information based on the confidence values corresponding to the respective matching vectors includes: Determining a confidence threshold based on the highest confidence value among the respective confidence values and a preset constant; determining the matching vectors with confidence values greater than or equal to the confidence threshold as first positioning vectors; determining a plurality of first panoramic image data with the same identifier as each of the first positioning vectors in the panoramic image database, and determining the target positioning information corresponding to the target geographical area based on the respective first panoramic image data; or, Determining the matching vector corresponding to the highest confidence value among the respective confidence values as a second positioning vector, determining second panoramic image data with the same identifier as the second positioning vector in the panoramic image database, and determining the target positioning information corresponding to the target geographical area based on the second panoramic image data; Wherein, the identifier of the panoramic image data in the panoramic image database corresponds to the identifier of the text embedding vector in the preset vector library.

[0082] Specifically, the target positioning information corresponding to the target geographical area can be determined in two ways. In the first way: A confidence threshold can be determined based on the highest confidence value among the respective confidence values and a preset constant, and the confidence threshold can be determined by the following formula:

[0083] Wherein, represents the highest confidence value among the respective confidence values, represents the preset constant, and the preset constant can be set as needed, for example, it can be 0.15.

[0084] Furthermore, the matching vectors can be screened by the confidence threshold, the matching vectors with confidence values less than the confidence threshold can be excluded, and the matching vectors with confidence values greater than or equal to the confidence threshold can be determined as first positioning vectors. Further, the identifier of the first positioning vector can be obtained and matched with the identifier of the panoramic image data in the panoramic image database to determine a plurality of first panoramic image data with the same identifier as each of the first positioning vectors. It should be noted that one first positioning vector corresponds to only one first panoramic image data. Further, the target positioning information can be determined according to the first panoramic image data corresponding to the first positioning vector with the highest confidence value, or the target positioning information can be determined according to the first panoramic image data corresponding to any first positioning vector. It should be noted that if the confidence values of all matching vectors are less than the confidence threshold, the positioning fails.

[0085] In the second method, the matching vector corresponding to the highest confidence value among the confidence values can be determined as the second positioning vector, and the remaining matching vectors can be completely removed. The second panoramic image data with the same identifier as the second positioning vector is determined in the panoramic image database, and then the target positioning information corresponding to the target geographical area is determined based on the second panoramic image data.

[0086] Among them, the identifier of the panoramic image data in the panoramic image database corresponds to the identifier of the text embedding vector in the preset vector library. The panoramic image data can include the identifier corresponding to the panoramic image, the positioning information, and the shooting time. Therefore, when the identifier of the positioning vector is the same as that of the panoramic image data, the positioning information corresponding to the panoramic image data can be determined as the target positioning information.

[0087] In the above embodiment, through efficient vector matching, the most similar vector is searched from the preset vector library, and the positioning information is queried in the panoramic image database based on the identifier of the matching result, so that the target positioning information corresponding to the target geographical area can be accurately determined.

[0088] In one embodiment, the panoramic image database is constructed through the following steps: Obtain multiple panoramic images and the panoramic image data corresponding to each of the panoramic images; the panoramic image data includes the identifier corresponding to the panoramic image, the positioning information, and the shooting time; Construct the panoramic image database based on the identifiers, positioning information, and shooting times corresponding to all the panoramic images.

[0089] Specifically, multiple panoramic images can be obtained from the open-source panoramic image libraries on the Internet. For example, multiple panoramic images can be obtained from the Google Street View library. The panoramic images use the equirectangular projection, so the panoramic images have the following characteristics: Horizontally (X-axis) in the image represents a 360° panorama, covering the azimuth angle (Yaw) from -180° to 180°. Vertically (Y-axis) in the image represents a 180° view angle, covering the pitch angle (Pitch) from -90° (zenith) to 90° (ground). Horizontally, the angle corresponding to each pixel is uniform, but there is distortion in the angle distribution vertically. The areas near the zenith and the ground will be stretched.

[0090] The panoramic image data corresponding to each panoramic image can also be obtained. The panoramic image data includes the identifier corresponding to the panoramic image, the positioning information, and the shooting time. The positioning information therein can include longitude and latitude.

[0091] The constructed panoramic image database can be shown in Table 1 below, for example: Table 1 Example Table of Panoramic Image Database

[0092] As shown in Table 1 above, for the panorama labeled 123, its longitude is 112.34, its latitude is 45.67, and the shooting time is 12:00 on March 15, 2025.

[0093] In the above embodiment, a panorama database is constructed according to the label, positioning information, and shooting time. The label enables the panorama to correspond to the text embedding vector. Then, according to this correspondence, the target text embedding vector corresponding to the text embedding vector can be determined based on the positioning information in the panorama database, and then the target positioning information can be determined, improving the flexibility and accuracy of positioning. The shooting time can improve the user experience.

[0094] In one embodiment, the panorama includes a plurality of pixels; the preset vector library is constructed through the following steps: For each of the panoramas, based on the original coordinates of each of the pixels in the panorama, the spherical coordinates corresponding to each of the pixels are determined; based on the original coordinates of each of the pixels in the panorama and the pixel color value of the pixel point corresponding to each of the pixels, the pixel color value corresponding to each of the pixels is determined; Based on the spherical coordinates corresponding to each of the pixels, the three-dimensional Cartesian coordinates corresponding to each of the pixels are determined; The panorama is projected onto a plurality of cube faces, and based on the three-dimensional Cartesian coordinates corresponding to each of the pixels, the cube face coordinates of each of the pixels on each of the cube faces are determined; For each of the cube faces, based on the pixel size of the cube face and the cube face coordinates of each of the pixels on the cube face, the texture pixel coordinates of each of the pixels on the cube face are determined; based on the texture pixel coordinates and pixel color values of each of the pixels on the cube face, the local scene map corresponding to the cube face is determined; For each of the panoramas, the plurality of local scene maps corresponding to the panorama are input into the multimodal large model, and the description texts corresponding to each of the local scene maps output by the multimodal large model are obtained; The description texts corresponding to the panorama are input into the text embedding model, and the text embedding vectors corresponding to each of the description texts output by the text embedding model are obtained; The preset vector library is constructed based on the text embedding vectors corresponding to all the panoramas.

[0095] Specifically, if directly performing semantic description on the panoramic image, the following problems will occur: ①Projection distortion causes the shape of objects to be distorted, and the shape and proportion of objects are distorted, affecting the understanding of the multimodal large model. For example, text (such as street signs) may become unrecognizable in the stretched area. ②Uneven perspectives reduce the multimodal large model's perception ability of vertical information. The semantic description may overly focus on horizontal information and ignore important features in the vertical direction. Information in the vertical direction may be misinterpreted. For example, the stretched part of the ground may be misrecognized as a large flat area. ③Detail loss may cause important positioning information (such as signs, logos) to be ignored. For example, if there is a store sign in the lower right corner of a panoramic image, but due to projection, it is small and blurred, the model may ignore it.

[0096] Therefore, it is necessary to convert the equidistant cylindrical projection panoramic image to a cube face projection. First, the spherical coordinates corresponding to each pixel can be determined based on the original coordinates of each pixel in the panoramic image. The spherical coordinates include the azimuth coordinate and the elevation coordinate, and the spherical coordinates corresponding to the pixel can be solved by the following formula:

[0097]

[0098] where represents the azimuth coordinate, with a range of , represents pi, represents the elevation coordinate, with a range of , represents the original coordinates of the pixel, represents the width of the panoramic image, represents the height of the panoramic image.

[0099] Then, the three-dimensional Cartesian coordinates corresponding to each pixel can be determined based on the spherical coordinates corresponding to each pixel. The three-dimensional Cartesian coordinates can be solved by the following formula:

[0100]

[0101]

[0102] where, it is easy to understand that controls the horizontal direction and determines the X-axis and Z-axis coordinates, controls the vertical direction and determines the Y-axis coordinate.

[0103] After obtaining the three-dimensional Cartesian coordinates corresponding to each pixel, the panoramic image can be projected onto multiple cube faces. It is easy to understand that a cube has six faces, namely the front, back, left, right, top, and bottom. Each face is a square, and the center of the cube can be set as the origin. The side length can be set to 2, for example. Then the range of pixel coordinates on the cube face is [-1, 1]. To facilitate the projection, it is also necessary to define the central direction vector of each cube face. The central direction vectors of the front (+Z), back (-Z), left (-X), right (+X), top (+Y), and bottom (-Y) are (0, 0, 1), (0, 0, -1), (-1, 0, 0), (1, 0, 0), (0, 1, 0), and (0, -1, 0) respectively. When projecting the panoramic image onto the cube face, it can be projected according to its largest component (maxComponen). Let maxComponent = max(|X|, |Y|, |Z|). If maxComponent = Z, then it is projected onto the front. If maxComponent = , then it is projected onto the back. If maxComponent = , then it is projected onto the left. If maxComponent = X, then it is projected onto the right. If maxComponent = Y, then it is projected onto the top. If maxComponent = , then it is projected onto the bottom.

[0104] After projecting the panoramic image, the cube face coordinates of each pixel on each cube face can be determined by the following formula: Front (+Z): ,

[0105] Back (-Z): ,

[0106] Left (-X): ,

[0107] Right (+X): ,

[0108] Top (+Y): ,

[0109] Bottom (-Y): ,

[0110] Among them, it is easy to understand that is in the range of [-1, 1].

[0111] For each cubic face, based on the pixel size of the cubic face and the cubic face coordinates of each pixel on the cubic face, the texture pixel coordinates of each pixel on the cubic face can be determined. The texture pixel coordinates can be solved by the following formula:

[0112]

[0113] Among them, represents the pixel size of the cubic face, that is, the cubic face includes pixels.

[0114] In addition, for each pixel in the panoramic image, the four pixel points corresponding to it, the pixel color value corresponding to each pixel can be determined based on the original coordinates of each pixel and the pixel color values of the pixel points corresponding to each pixel. The pixel corresponds to the upper left pixel point , the upper right pixel point , the lower left pixel point , and the lower right pixel point . Among them, is the value of rounding down, is the value of rounding down. The pixel color values corresponding to the four pixel points are respectively , , , .

[0115] Furthermore, the pixel color value can be solved by the bilinear interpolation method. Solving the pixel color value by the bilinear interpolation method can reduce distortion and obtain a smoother visual effect. In the bilinear interpolation method, first interpolation can be performed in the direction to obtain , . Among them, represents the weight in the direction. Then interpolation can be performed in the direction to obtain . Among them, represents

[0116] Finally, based on the texture pixel coordinates of each pixel on the cubic face and the pixel color value , the local scene image corresponding to the cubic face can be determined.

[0117] Furthermore, for each panoramic image, multiple local scene images corresponding to the panoramic image can be input into the multi-modal large model to obtain description texts respectively corresponding to the local scene images output by the multi-modal large model. It should be noted that only the local scene images corresponding to the front, back, left, and right sides can be input into the multi-modal large model. The top view is usually the sky, without obvious landmarks, buildings, or street signs, making it difficult to provide meaningful semantic descriptions. Moreover, the information such as the color of the sky and clouds changes greatly, which is not helpful for positioning. The bottom view mainly shows the ground, sidewalks, and roads, usually only with textures (such as asphalt roads and concrete floors) or shadows, lacking unique positioning features, and may contain the shadows of pedestrians or vehicles, and this shadow information is not significant for position matching. Excluding the local scene images corresponding to the top and bottom can reduce storage occupancy and computational amount, and improve the positioning efficiency.

[0118] Further, each description text corresponding to the panoramic image (after excluding the top and bottom, that is, one panoramic image corresponds to four local scene images, and one local scene image corresponds to one description text) can be input into the text embedding model to obtain text embedding vectors respectively corresponding to the description texts output by the text embedding model. Among them, the determination process of the text embedding vector can refer to step 120, which will not be elaborated in this embodiment. Finally, a preset vector library can be constructed based on the text embedding vectors corresponding to all panoramic images. Exemplarily, the preset vector library can be as shown in Table 2 below: Table 2 Example Table of Preset Vector Library

[0119] As shown in Table 2 above, the panoramic image marked as UUID-123 corresponds to four text embedding vectors. After the target text embedding vector matches a certain text embedding vector, the panoramic image database can be queried according to the UUID, and the target positioning information corresponding to the target geographical area can be determined according to the query result.

[0120] It should be noted that in the prior art, there is also an image matching and positioning method based on a panoramic image database. However, this method mainly relies on the local features of the captured panoramic image. These local features are sensitive to illumination changes, occlusion, and perspective changes, and are prone to ignoring high-level semantic information such as buildings, signs, and roads, resulting in low matching accuracy. In addition, the image matching and positioning method based on the panoramic image database usually needs to search for the most similar image in a large-scale database, which consumes a large amount of storage and computing resources. The technical solution of the present invention converts to a cube projection through equidistant cylindrical projection, reduces image deformation, and improves the accuracy of semantic extraction. At the same time, only the four most informative front, back, left, and right perspectives are selected for processing, and an efficient vector indexing technology is used to accelerate the retrieval, so as to maintain a fast and stable response ability in a large-scale data environment.

[0121] In the above embodiments, the identifier and positioning information corresponding to the panoramic image are stored. Since the panoramic image uses equidistant cylindrical projection and there is a deformation problem, it needs to be converted to cube projection to obtain pictures of the front, back, left, right, top, and bottom six perspectives. Among the six perspectives of cube projection, only the front, back, left, and right four perspectives are selected for processing to reduce redundant information and improve analysis efficiency. Subsequently, a description text is generated using a multimodal large model and standardized using a unified preset prompt template to enhance the consistency and accuracy of matching. This enables the stable provision of effective location information in complex geographical environments and improves the applicability of the navigation and positioning system.

[0122] The geographical positioning method based on a multimodal large model provided by the present invention inputs a target picture corresponding to a target geographical area into the multimodal large model to obtain a target description text corresponding to the target picture output by the multimodal large model; the multimodal large model is used to generate the target description text based on the target picture; the target description text is input into a text embedding model to obtain a target text embedding vector corresponding to the target description text output by the text embedding model; the target text embedding vector is matched with each text embedding vector in a preset vector library to determine the target positioning information corresponding to the target geographical area. The technical solution of the present invention, through the connection of images, texts, and vectors, and the matching of the preset vector library with the target text embedding vector, does not rely on local feature points of the image, so it is not easily affected by the dynamic environment and can still stably perform positioning in complex environments with relatively high positioning accuracy. In addition, compared with the SLAM technology, the present invention can also obtain absolute position information through the matching process.

[0123] The geographical positioning device based on a multimodal large model provided by the present invention will be described below. The geographical positioning device based on a multimodal large model described below can be correspondingly referred to the geographical positioning method described above.

[0124] Figure 2 is a schematic structural diagram of the geographical positioning device based on a multimodal large model provided by the present invention, as Figure 2 shown, the geographical positioning device 200 based on a multimodal large model includes the following modules: A text module 210, configured to input a target picture corresponding to a target geographical area into a multimodal large model to obtain a target description text corresponding to the target picture output by the multimodal large model; the multimodal large model is used to generate the target description text based on the target picture; A vector module 220, configured to input the target description text into a text embedding model to obtain a target text embedding vector corresponding to the target description text output by the text embedding model; A positioning module 230, configured to match the target text embedding vector with each text embedding vector in a preset vector library to determine target positioning information corresponding to the target geographical area.

[0125] In one embodiment, the vector module 220 is specifically configured to: Input the target description text into a large language model to obtain a standardized text output by the large language model; the large language model is used to perform standardization processing on the target description text; the standardization processing includes vocabulary unification processing, sentence pattern unification processing, removal of irrelevant characters, data format unification processing, and stop word filtering processing; Input the standardized text into the text embedding model to obtain the target text embedding vector output by the text embedding model; the text embedding model is used to generate the target text embedding vector after performing a long text splitting strategy, a word segmentation strategy, and an embedding strategy on the standardized text.

[0126] In one embodiment, the positioning module 230 is specifically configured to: For each text embedding vector in the preset vector library, calculate the similarity value between the target text embedding vector and the text embedding vector; Determine the first preset number of text embedding vectors sorted in descending order of similarity value as matching vectors, and determine the target positioning information according to each of the matching vectors.

[0127] In one embodiment, the positioning module 230 is specifically further configured to: For each of the matching vectors, determine the relative similarity value corresponding to the matching vector based on the similarity value of the matching vector, the highest matching vector with the highest similarity value among all the matching vectors, and the lowest matching vector with the lowest similarity value; Determine the confidence value corresponding to the matching vector based on the sum of the relative similarity values of all the matching vectors and the relative similarity value corresponding to the matching vector; Determine the target positioning information based on the confidence values corresponding to each of the matching vectors.

[0128] In one embodiment, the positioning module 230 is specifically further configured to: Determine a confidence threshold based on the highest confidence value among each of the confidence values and a preset constant; determine the matching vectors with confidence values greater than or equal to the confidence threshold as the first positioning vectors; determine a plurality of first panoramic image data with the same identifier as each of the first positioning vectors in a panoramic image database, and determine the target positioning information corresponding to the target geographical area based on each of the first panoramic image data; or, Determine the matching vector corresponding to the highest confidence value among the confidence values as the second positioning vector, determine the second panoramic image data with the same identifier as the second positioning vector in the panoramic image database, and determine the target positioning information corresponding to the target geographical area based on the second panoramic image data; Wherein, the identifier of the panoramic image data in the panoramic image database corresponds to the identifier of the text embedding vector in the preset vector library.

[0129] In one embodiment, the geographical positioning device based on the multi-modal large model further includes a first construction module, and the first construction module is specifically used for: Obtain multiple panoramic images and the panoramic image data corresponding to each of the panoramic images; the panoramic image data includes the identifier, positioning information, and shooting time corresponding to the panoramic image; Construct the panoramic image database based on the identifiers, positioning information, and shooting times corresponding to all the panoramic images.

[0130] In one embodiment, the panoramic image includes multiple pixels; the geographical positioning device based on the multi-modal large model further includes a second construction module, and the second construction module is specifically used for: For each of the panoramic images, determine the spherical coordinates corresponding to each of the pixels based on the original coordinates of the pixels in the panoramic image; determine the pixel color values corresponding to each of the pixels based on the original coordinates of the pixels in the panoramic image and the pixel color values of the pixel points corresponding to the pixels; Determine the three-dimensional Cartesian coordinates corresponding to each of the pixels based on the spherical coordinates corresponding to each of the pixels; Project the panoramic image onto multiple cube faces, and determine the cube face coordinates of each of the pixels on each of the cube faces based on the three-dimensional Cartesian coordinates corresponding to each of the pixels; For each of the cube faces, determine the texture pixel coordinates of each of the pixels on the cube face based on the pixel size of the cube face and the cube face coordinates of each of the pixels on the cube face; determine the local scene image corresponding to the cube face based on the texture pixel coordinates and pixel color values of each of the pixels on the cube face; For each of the panoramic images, input the multiple local scene images corresponding to the panoramic image into the multi-modal large model to obtain the description texts corresponding to each of the local scene images output by the multi-modal large model; Input the description texts corresponding to each of the panoramic images into the text embedding model to obtain the text embedding vectors corresponding to each of the description texts output by the text embedding model; Construct the preset vector library based on the text embedding vectors corresponding to all the panoramic images.

[0131] The geolocation device based on a multimodal large model provided by the present invention inputs a target picture corresponding to a target geographical area into the multimodal large model to obtain a target description text corresponding to the target picture output by the multimodal large model; the multimodal large model is used to generate the target description text based on the target picture; the target description text is input into a text embedding model to obtain a target text embedding vector corresponding to the target description text output by the text embedding model; the target text embedding vector is matched with each text embedding vector in a preset vector library to determine the target location information corresponding to the target geographical area. The technical solution of the present invention, through the connection of images, texts, and vectors, and the matching between the preset vector library and the target text embedding vector, does not rely on local feature points of the image, so it is not easily affected by the dynamic environment, can still perform stable positioning in complex environments, and has a high positioning accuracy. In addition, compared with the SLAM technology, the present invention can also obtain absolute position information through the matching process.

[0132] Figure 3 An example of the physical structure diagram of an electronic device is shown as Figure 3 As shown, the electronic device may include: a processor 310, a communication interface 320, a memory 330, and a communication bus 340. Among them, the processor 310, the communication interface 320, and the memory 330 complete mutual communication through the communication bus 340. The processor 310 can call the logical instructions in the memory 330 to execute the geolocation method based on the multimodal large model.

[0133] In addition, when the logical instructions in the above-mentioned memory 330 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0134] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the geolocation method based on a multimodal large model provided by the above-mentioned various methods.

[0135] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the geolocation method based on a multimodal large model provided by the above-mentioned various methods.

[0136] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0137] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0138] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A geolocation method based on a multimodal large model, characterized in that, Including: Input the target picture corresponding to the target geographical area into the multi-modal large model to obtain the target description text corresponding to the target picture output by the multi-modal large model; The multi-modal large model is used to generate the target description text based on the target picture; Input the target description text into the text embedding model to obtain the target text embedding vector corresponding to the target description text output by the text embedding model; Match the target text embedding vector with each text embedding vector in the preset vector library to determine the target location information corresponding to the target geographical area; The matching of the target text embedding vector with each text embedding vector in the preset vector library to determine the target location information corresponding to the target geographical area includes: For each text embedding vector in the preset vector library, calculate the similarity value between the target text embedding vector and the text embedding vector; Determine the first preset number of text embedding vectors sorted by similarity value from large to small as the matching vectors, and determine the target location information according to each matching vector.

2. The geolocation method based on a multimodal large model according to claim 1, wherein The determining the target location information according to each matching vector includes: For each matching vector, determine the relative similarity value corresponding to the matching vector based on the similarity value of the matching vector, the highest matching vector with the highest similarity value among all the matching vectors, and the lowest matching vector with the lowest similarity value; Based on the sum of the relative similarity values of all the matching vectors and the relative similarity value corresponding to the matching vector, determine the confidence value corresponding to the matching vector; Determine the target location information based on the confidence values corresponding to each matching vector.

3. The geolocation method based on a multimodal large model according to claim 2, wherein The determining the target location information based on the confidence values corresponding to each matching vector includes: Based on the highest confidence value among the confidence values and a preset constant, determine a confidence threshold; determine the matching vectors with confidence values greater than or equal to the confidence threshold as the first positioning vectors; determine a plurality of first panoramic image data with the same identifier as each of the first positioning vectors in the panoramic image database, and determine the target location information corresponding to the target geographical area based on each of the first panoramic image data; or, Determine the matching vector corresponding to the highest confidence value among the confidence values as the second positioning vector, determine the second panoramic image data with the same identifier as the second positioning vector in the panoramic image database, and determine the target location information corresponding to the target geographical area based on the second panoramic image data; Wherein, the identifier of the panoramic image data in the panoramic image database corresponds to the identifier of the text embedding vector in the preset vector library.

4. The geolocation method based on a multimodal large model according to claim 3, characterized in that The panoramic image database is constructed through the following steps: Obtain multiple panoramic images and the panoramic image data corresponding to each of the panoramic images; the panoramic image data includes the identifier, location information, and shooting time corresponding to the panoramic image; Construct the panoramic image database based on the identifiers, location information, and shooting times corresponding to all the panoramic images.

5. The geolocation method based on a multimodal large model according to claim 4, wherein The panoramic image includes multiple pixels; the preset vector library is constructed through the following steps: For each of the panoramic images, determine the spherical coordinates corresponding to each pixel based on the original coordinates of each pixel in the panoramic image; determine the pixel color value corresponding to each pixel based on the original coordinates of each pixel in the panoramic image and the pixel color value of the pixel point corresponding to each pixel; Determine the three-dimensional Cartesian coordinates corresponding to each pixel based on the spherical coordinates corresponding to each pixel; Project the panoramic image onto multiple cube faces, and determine the cube face coordinates of each pixel on each cube face based on the three-dimensional Cartesian coordinates corresponding to each pixel; For each cube face, determine the texture pixel coordinates of each pixel on the cube face based on the pixel size of the cube face and the cube face coordinates of each pixel on the cube face; determine the local scene image corresponding to the cube face based on the texture pixel coordinates and pixel color values of each pixel on the cube face; For each panoramic image, input the multiple local scene images corresponding to the panoramic image into the multimodal large model, and obtain the description text corresponding to each local scene image output by the multimodal large model; Input the description texts corresponding to each panoramic image into the text embedding model, and obtain the text embedding vectors corresponding to each description text output by the text embedding model; Construct the preset vector library based on the text embedding vectors corresponding to all the panoramic images.

6. The geolocation method based on a multimodal large model according to any one of claims 1 to 5, characterized in that, The step of inputting the target description text into the text embedding model to obtain the target text embedding vector corresponding to the target description text output by the text embedding model includes: Input the target description text into a large language model to obtain the standardized text output by the large language model; the large language model is used to perform standardization processing on the target description text; the standardization processing includes vocabulary unification processing, sentence pattern unification processing, removal of irrelevant characters, data format unification processing, and stop word filtering processing; Input the standardized text into the text embedding model to obtain the target text embedding vector output by the text embedding model; the text embedding model is used to generate the target text embedding vector after performing a long text splitting strategy, a word segmentation strategy, and an embedding strategy on the standardized text.

7. A geolocation device based on a multimodal large model, characterized in that, It includes: A text module for inputting a target image corresponding to a target geographical area into a multimodal large model to obtain the target description text corresponding to the target image output by the multimodal large model; The multimodal large model is used to generate the target description text based on the target image; A vector module for inputting the target description text into a text embedding model to obtain the target text embedding vector corresponding to the target description text output by the text embedding model; A positioning module for matching the target text embedding vector with each text embedding vector in the preset vector library to determine the target positioning information corresponding to the target geographical area; The positioning module is specifically used for: For each text embedding vector in the preset vector library, calculate the similarity value between the target text embedding vector and the text embedding vector; Determine the first preset number of text embedding vectors sorted from largest to smallest similarity value as matching vectors, and determine the target location information according to each of the matching vectors.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the geographic location method based on the multimodal large model according to any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the geographic location method based on the multimodal large model according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Video coding method, server, device and computer readable storage medium

    CN113014924A

  • Layout-controllable three-dimensional scene characterization and generation method based on large language model

    CN117409140A

  • Street view image geographic positioning reasoning method based on visual language large model

    CN118551066A

  • Panorama processing method, server, storage medium and program product

    CN118608653A

  • Image geographic positioning method and system based on street view image super-radius cube projection three-dimensional reconstruction

    CN118918185A

Cited By

  • Geospatial position positioning method, device, equipment and product

    CN121740086A