Visual positioning method and device, electronic equipment and storage medium

By combining the visual positioning method of multi-view two-dimensional images and three-dimensional point cloud features, and using the target positioning inference model to decode the target position information, the problem of insufficient visual positioning accuracy in the existing technology is solved, and higher positioning accuracy is achieved.

CN120655709APending Publication Date: 2025-09-16PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510713100.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing visual positioning technologies have difficulty accurately understanding natural language expressions with complex semantics and implicit intentions, resulting in insufficient positioning accuracy.

Method used

By acquiring multi-view two-dimensional images and three-dimensional point cloud features of the target environment, combined with target positioning prompt text, and using the target positioning inference model to perform position reasoning, a structured reasoning chain is formed, and target position information is generated through feature decoding.

Benefits of technology

The accuracy of visual positioning in complex scenes is significantly improved, and the position of the target object in the environment can be intuitively indicated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655709A_ABST
    Figure CN120655709A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a visual positioning method and device, electronic equipment and a storage medium, belongs to the technical field of artificial intelligence, and is suitable for financial science and technology scenes and medical science and technology scenes. The method comprises the following steps: acquiring a multi-view two-dimensional image and a three-dimensional point cloud feature of a target environment; obtaining a target positioning prompt text; wherein the target positioning prompt text is used for indicating the position of the target object found in the target environment; performing position reasoning on the target positioning prompt text and the multi-view two-dimensional image through a preset target positioning reasoning model to obtain initial positioning reasoning data; performing visual positioning according to the initial positioning reasoning data and the three-dimensional point cloud features to obtain target positioning features; performing feature decoding based on the target positioning features to obtain target position information; wherein the target position information is used for indicating the position of the target object in the target environment. According to the embodiment of the invention, the accuracy of visual positioning can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and is applicable to financial technology scenarios and medical technology scenarios, and in particular to a visual positioning method and device, electronic equipment, and storage medium. Background Art

[0002] Visual positioning is a technology that analyzes three-dimensional visual data based on natural language descriptions to determine the position and orientation of target objects in space. Visual positioning can be applied in a variety of scenarios. For example, in financial technology, bank branches can use visual positioning technology to locate the corresponding business processing location in the branch's three-dimensional space based on customer voice commands. In medical technology, for example, visual positioning technology can be used to quickly locate medical images to assist medical staff in diagnosis and analysis.

[0003] At present, visual positioning mainly relies on clear text descriptions for positioning. It lacks the ability to deeply understand natural language expressions containing complex semantics, implicit intentions or multiple conditional constraints, making it difficult to accurately understand the user's real needs, thus affecting the accuracy of visual positioning.

[0004] Therefore, how to improve the accuracy of visual positioning has become a technical problem that needs to be solved urgently. Summary of the Invention

[0005] The main purpose of the embodiments of the present application is to provide a visual positioning method and device, an electronic device and a storage medium, aiming to improve the accuracy of visual positioning.

[0006] To achieve the above objectives, a first aspect of an embodiment of the present application provides a visual positioning method, the method comprising:

[0007] Get multiple views of the target environment Figure 2 3D image and 3D point cloud features;

[0008] Obtaining a target positioning prompt text; wherein the target positioning prompt text is used to indicate the location of the target object in the target environment;

[0009] The target positioning prompt text and the multi-view Figure 2 3D image to perform position reasoning and obtain initial positioning reasoning data;

[0010] Perform visual positioning based on the initial positioning inference data and the three-dimensional point cloud features to obtain target positioning features;

[0011] Feature decoding is performed based on the target positioning feature to obtain target position information; wherein the target position information is used to indicate the position of the target object in the target environment.

[0012] In some embodiments, the target positioning prompt text and the multi-view Figure 2 dimensional image to perform position reasoning and obtain initial positioning reasoning data, including:

[0013] The multi-view Figure 2 dimensional image to encode the image and obtain the multi-view Figure 2 dimensional features;

[0014] Performing text encoding on the target positioning prompt text to obtain prompt text encoding features;

[0015] The target positioning inference model is used to encode the prompt text features and the multi-view Figure 2 dimensional features to perform positioning reasoning to obtain the initial positioning reasoning data.

[0016] In some embodiments, the target positioning inference model is used to encode the prompt text features and the multi-view Figure 2 Performing positioning reasoning based on the dimensional features to obtain the initial positioning reasoning data includes:

[0017] Performing text decomposition on the prompt text encoding feature to obtain text encoding sub-features;

[0018] The text encoding sub-features and the multi-view Figure 2 Perform chain reasoning based on dimensional features to obtain original positioning reasoning data;

[0019] Structural processing is performed based on the coding features of the prompt text and the original positioning inference data to obtain the initial positioning inference data.

[0020] In some embodiments, performing visual positioning based on the initial positioning inference data and the three-dimensional point cloud features to obtain target positioning features includes:

[0021] Performing feature encoding on the three-dimensional point cloud features to obtain three-dimensional scene features;

[0022] Extracting position information from the initial positioning inference data to obtain a two-dimensional positioning feature;

[0023] Performing context extraction on the initial positioning inference data to obtain location context features;

[0024] Performing attention calculation based on the three-dimensional scene feature, the two-dimensional positioning feature, and the position context feature to obtain an initial positioning feature;

[0025] Feature screening is performed on the initial positioning features to obtain the target positioning features.

[0026] In some embodiments, performing attention calculation based on the three-dimensional scene feature, the two-dimensional positioning feature, and the position context feature to obtain the initial positioning feature includes:

[0027] Calculating similarity between the three-dimensional scene features and the two-dimensional positioning features to obtain scene matching data;

[0028] Normalizing the scene matching data to obtain a scene matching weight;

[0029] The location context feature and the scene matching weight are weighted and summed to obtain the initial positioning feature.

[0030] In some embodiments, the performing feature screening on the initial positioning feature to obtain the target positioning feature includes:

[0031] A similarity heat map is constructed based on the initial positioning features to obtain a target feature map; wherein the target feature map is used to represent the feature similarity between each region in the three-dimensional scene feature and the two-dimensional positioning feature;

[0032] The initial positioning feature is subjected to regional screening based on the target feature map to obtain the target positioning feature.

[0033] In some embodiments, the target positioning prompt text and the multi-view Figure 2 Before performing position reasoning on the dimensional image to obtain initial positioning reasoning data, the method further includes:

[0034] Acquire visual positioning sample data, wherein the visual positioning sample data includes a sample two-dimensional scene image, a sample positioning prompt text, and sample position information;

[0035] Performing position reasoning on the sample two-dimensional scene image and the sample positioning prompt text using a preset original positioning reasoning model to obtain sample positioning reasoning data;

[0036] Performing recognition loss calculation based on the sample position information and the sample positioning inference data to obtain a position inference loss function;

[0037] The parameters of the original positioning reasoning model are adjusted based on the position reasoning loss function to obtain the target positioning reasoning model.

[0038] To achieve the above-mentioned purpose, a second aspect of an embodiment of the present application provides a visual positioning device, comprising:

[0039] Environmental information acquisition module, used to obtain multiple views of the target environment Figure 2 3D image and 3D point cloud features;

[0040] A prompt text acquisition module is used to acquire a target positioning prompt text; wherein the target positioning prompt text is used to indicate the location of the target object in the target environment;

[0041] Position reasoning module, used for positioning the target prompt text and the multi-view Figure 2 3D image to perform position reasoning and obtain initial positioning reasoning data;

[0042] A visual positioning module is used to perform visual positioning based on the initial positioning inference data and the three-dimensional point cloud features to obtain target positioning features;

[0043] A feature decoding module is used to perform feature decoding based on the target positioning feature to obtain target position information; wherein the target position information is used to indicate the position of the target object in the target environment.

[0044] To achieve the above-mentioned purpose, the third aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the method described in the first aspect when executing the computer program.

[0045] To achieve the above-mentioned purpose, the fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the method described in the first aspect.

[0046] The visual positioning method and device, electronic device and storage medium proposed in this application obtain multiple visual positions of the target environment. Figure 2 3D image and 3D point cloud features, and obtain target positioning prompt text; wherein, the target positioning prompt text is used to indicate the location of the target object in the target environment. Then, the target positioning prompt text and multi-view are compared using the target positioning inference model. Figure 2 The system uses the initial positioning inference data and the 3D point cloud features to perform position reasoning on the 3D image, forming a structured reasoning chain, namely the initial positioning inference data, thereby converting the ambiguous text description into an intermediate representation of the implicit spatial relationship. Furthermore, visual positioning is performed based on the initial positioning inference data and the 3D point cloud features to obtain the target positioning features. Finally, the target positioning features are feature decoded to generate target position information. The target position information can intuitively indicate the position of the target object in the target environment, significantly improving the accuracy of visual positioning in complex scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1is a flow chart of the visual positioning method provided in an embodiment of the present application;

[0048] Figure 2 is a flow chart of a visual positioning method provided by another embodiment of the present application;

[0049] Figure 3 yes Figure 1 Flowchart of step S103 in FIG.

[0050] Figure 4 yes Figure 3 Flowchart of step S303 in FIG.

[0051] Figure 5 yes Figure 1 Flowchart of step S104 in FIG.

[0052] Figure 6 yes Figure 5 Flowchart of step S504 in FIG.

[0053] Figure 7 yes Figure 5 Flowchart of step S505 in FIG.

[0054] Figure 8 Schematic diagram of the structure of the visual positioning device provided in the embodiment of the present application;

[0055] Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0056] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0057] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.

[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0059] First, let’s analyze some of the terms used in this application:

[0060] Artificial intelligence (AI) is a new technical discipline that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. A branch of computer science, AI seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thinking. It also encompasses the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results.

[0061] Natural language processing (NLP): NLP uses computers to process, understand, and apply human languages ​​(such as Chinese and English). A branch of artificial intelligence, NLP is an interdisciplinary field between computer science and linguistics, often referred to as computational linguistics. Natural language processing encompasses grammatical analysis, semantic analysis, and discourse comprehension. Natural language processing is commonly used in technical fields such as machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information intent recognition, information extraction and filtering, text classification and clustering, public opinion analysis, and opinion mining. It encompasses data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research related to language processing, and linguistics research related to language computing.

[0062] Information Extraction: A text processing technology that extracts specified types of entity, relationship, event, and other factual information from natural language text and forms structured data output. Information extraction is a technology that extracts specific information from text data. Text data is composed of some specific units, such as sentences, paragraphs, and chapters. Text information is composed of some small specific units, such as characters, words, phrases, sentences, paragraphs, or a combination of these specific units. Extracting noun phrases, names, place names, etc. from text data is all text information extraction. Of course, the information extracted by text information extraction technology can be of various types.

[0063] Visual Localization refers to the process of determining the precise position and posture (position and orientation) of a device or target in a known scene through computer vision technology, using environmental images or video data collected by cameras or other visual sensors, based on natural language descriptions combined with algorithmic processing (such as feature extraction, matching, 3D reconstruction, etc.). The core of visual positioning is to compare visual information with pre-built maps or scene models to achieve real-time spatial positioning. Visual positioning can be applied to a variety of application scenarios. For example, in financial technology scenarios, bank branches can use visual positioning technology to locate the corresponding business processing location in the three-dimensional space of the branch based on the voice commands input by customers; for example, in medical technology scenarios, visual positioning technology can be used to quickly locate medical images to assist medical staff in diagnosis and analysis.

[0064] A 3D point cloud is a collection of a large number of discrete three-dimensional spatial points. Each point contains at least one set of coordinates (X, Y, Z) and can also include information such as color (RGB), intensity, or normal vector. 3D point clouds are acquired using technologies such as LiDAR, depth cameras, or multi-view image reconstruction. They can accurately characterize the geometric structure and spatial distribution of object surfaces and are widely used in fields such as autonomous driving, robotic navigation, and 3D modeling. Compared to traditional two-dimensional images, point clouds directly reflect the three-dimensional topological relationships of the real world. However, they are unstructured, sparse, and noisy, requiring optimization and analysis through point cloud processing algorithms (such as filtering, segmentation, and registration).

[0065] The Attention Mechanism is a computational technique that mimics human cognitive focusing, allowing the model to dynamically assign weights when processing input data, prioritizing the parts most relevant to the current task. The core idea of ​​the Attention Mechanism is to generate a weighted contextual representation by calculating similarity scores between the query, key, and value.

[0066] Feature-level Reasoning Activation Heatmap, Feature-level Reasoning Activation Heatmap is an implicit spatial attention representation generated in a multimodal 3D visual positioning task, which is used to reflect the semantic association strength between different regions in a 3D scene and the positioning target in natural language instructions. The core principle is to match the semantic features of the language reasoning chain with the geometric features of the 3D point cloud through a cross-attention mechanism, calculate the response weight of each 3D local feature to the key positioning token in the language description, and finally form a heat map. The high-value area represents the potential target position that is highly matched with the instruction, and the low-value area is suppressed. This heat map is different from the traditional pixel-level heat map. It is based on the semantic alignment of the high-level feature space. It can not only integrate language logic, but also adapt to the sparsity and perspective changes of the 3D point cloud, providing an interpretable intermediate representation for subsequent precise positioning.

[0067] At present, visual positioning mainly relies on clear text descriptions for positioning. It lacks the ability to deeply understand natural language expressions containing complex semantics, implicit intentions or multiple conditional constraints, making it difficult to accurately understand the user's real needs, thus affecting the accuracy of visual positioning.

[0068] Based on this, the embodiments of the present application provide a visual positioning method and device, an electronic device, and a storage medium, aiming to improve the accuracy of visual positioning.

[0069] The visual positioning method and device, electronic device and storage medium provided in the embodiments of the present application are specifically illustrated through the following embodiments. First, the visual positioning method in the embodiments of the present application is described.

[0070] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.

[0071] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0072] The visual positioning method provided in the embodiment of the present application relates to the field of artificial intelligence technology. The visual positioning method provided in the embodiment of the present application can be applied to a terminal, can be applied to a server side, or can be software running in a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the visual positioning method, etc., but is not limited to the above forms.

[0073] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.

[0074] It should be noted that in each specific embodiment of the present application, when it comes to the need to perform relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.

[0075] Figure 1 This is an optional flowchart of the visual positioning method provided in the embodiment of the present application. Figure 1 The method may include but is not limited to steps S101 to S105.

[0076] Step S101: Acquire multiple views of the target environment Figure 2 3D image and 3D point cloud features;

[0077] Step S102: obtaining a target location prompt text; wherein the target location prompt text is used to indicate the location of the target object in the target environment;

[0078] Step S103: Using a preset target location inference model to locate the target prompt text and the multi-view Figure 2 3D image to perform position reasoning and obtain initial positioning reasoning data;

[0079] Step S104: Perform visual positioning based on the initial positioning inference data and the three-dimensional point cloud features to obtain target positioning features;

[0080] Step S105 , performing feature decoding based on the target positioning feature to obtain target position information; wherein the target position information is used to indicate the position of the target object in the target environment.

[0081] In the embodiment of the present application, steps S101 to S105 are shown by obtaining multiple views of the target environment. Figure 2 3D image and 3D point cloud features, and obtain target positioning prompt text; wherein, the target positioning prompt text is used to indicate the location of the target object in the target environment. Then, the target positioning prompt text and multi-view are compared using the target positioning inference model. Figure 2 The system uses the initial positioning inference data and the 3D point cloud features to perform position reasoning on the 3D image, forming a structured reasoning chain, namely the initial positioning inference data, thereby converting the ambiguous text description into an intermediate representation of the implicit spatial relationship. Furthermore, visual positioning is performed based on the initial positioning inference data and the 3D point cloud features to obtain the target positioning features. Finally, the target positioning features are feature decoded to generate target position information. The target position information can intuitively indicate the position of the target object in the target environment, significantly improving the accuracy of visual positioning in complex scenes.

[0082] In step S101 of some embodiments, the target environment is determined based on the actual application scenario. For example, in a fintech scenario, the target environment may be a bank branch, a trading hall, a conference room, etc. In a medical technology scenario, the target environment may be a hospital environment, the physiological environment of biological tissue, etc. However, this is not limiting.

[0083] Among them, the multi-view of the target environment Figure 2 A three-dimensional image is an image of the target environment captured from different perspectives. For example, in a square room, the room is captured at the four corners.

[0084] The 3D point cloud features of the target environment refer to structured representations extracted from the target environment's point cloud data (e.g., a set of 3D spatial coordinates collected by a laser radar or depth camera). These features describe the scene's geometry, semantics, and context, providing precise 3D coordinates, shape, and scale information to compensate for the lack of depth in 2D images. 3D point cloud data can be collected using technologies such as laser radar (LiDAR) and depth cameras.

[0085] In some embodiments, in medical technology scenarios, the three-dimensional point cloud data of the physiological environment of biological tissues can be obtained by computed tomography (CT), magnetic resonance imaging (MRI), ultrasonic three-dimensional reconstruction and other technologies to achieve a comprehensive scan of the physiological environment of biological tissues. Figure 2 Multi-dimensional images can be collected through medical imaging equipment (CT equipment\MRI equipment\ultrasound equipment), endoscopes / microscopes and other technologies to obtain multi-view images of the front, side and surrounding tissues of the physiological environment in which the biological tissue is located.

[0086] In some embodiments, in step S102, the target location prompt text is manually input and is used to assign a reasoning task to the target location reasoning model, instructing it to find the location of the target object in the target environment, or to find a suitable location for the target object to stay or be placed. The target object can be, but is not limited to, a person, an object, an animal, or a plant.

[0087] For example, the target positioning prompt text may be: "Please find a trash can", "Please find a table near the window", "Please locate the ATM at the door of the bank branch", "I am looking for a comfortable position to watch TV, which position is most suitable for me", "Please find the lesion area", etc. It is understandable that the target positioning prompt text contains target objects, such as "trash can", "table", "ATM", "I" (i.e., a person), "lesion area"; in addition, the target positioning prompt text may also include a location description of the target object, i.e., "table near the window", "ATM at the door of the bank branch", etc.

[0088] See also Figure 2 In some embodiments, before step S103, the visual positioning method may further include but is not limited to steps S201 to S204:

[0089] Step S201: Obtain visual positioning sample data, wherein the visual positioning sample data includes a sample two-dimensional scene image, a sample positioning prompt text, and sample position information;

[0090] Step S202: performing position reasoning on the sample two-dimensional scene image and the sample positioning prompt text using a preset original positioning reasoning model to obtain sample positioning reasoning data;

[0091] Step S203, performing recognition loss calculation based on the sample position information and the sample positioning inference data to obtain a position inference loss function;

[0092] Step S204: Adjust the parameters of the original positioning reasoning model based on the position reasoning loss function to obtain a target positioning reasoning model.

[0093] Steps S201 to S204 shown in the embodiment of the present application are used to train the original positioning reasoning model by obtaining visual positioning sample data including sample two-dimensional scene images, sample positioning prompt texts and sample position information. Then, the original positioning reasoning model is used to perform position reasoning on the sample two-dimensional scene images and the sample positioning prompt texts, extract visual-language association features from the multimodal input, establish a mapping relationship between semantics and space, find the position of the sample object, and obtain sample positioning reasoning data. Furthermore, by performing loss calculation on the sample positioning reasoning data obtained by model reasoning and the pre-labeled sample position information, a position reasoning loss function is obtained to quantify the deviation of the model in the positioning reasoning process. Finally, based on the position reasoning loss function, the parameters of the original positioning reasoning model are adjusted to optimize the model's capabilities and obtain a target positioning reasoning model, which improves the accuracy and generalization ability of the positioning reasoning model and can more accurately convert text instructions into specific positioning information.

[0094] In step S201 of some embodiments, the visual positioning sample data is a data set pre-collected in an actual application scenario and obtained by manual annotation, and includes a sample two-dimensional scene image, a sample positioning prompt text and a sample position information.

[0095] Wherein, the sample two-dimensional scene image is a two-dimensional image of multiple views of the sample environment;

[0096] The sample location prompt text is used to indicate the location of the sample object in the sample environment. The details are basically the same as the above-mentioned "target location prompt text" and will not be repeated here.

[0097] The sample location information is manually labeled to represent the exact location of the sample object in the sample environment.

[0098] In step S202 of some embodiments, the preset original positioning inference model can adopt a multimodal large language model, a graph neural network-based model, a reinforcement learning model, etc., and the specific selection needs to be based on the actual application scenario, but is not limited to this.

[0099] In some embodiments, the original positioning reasoning model is a multimodal large language model, which can fully utilize the reasoning capabilities of the large language model to implement chained positioning reasoning. The multimodal large language model can be, but is not limited to, an LLMA, GPT model, etc.

[0100] For example:

[0101] The sample positioning prompt text is "Find the red fire hydrant on the right side of the conference room", and the sample 2D scene image is a multi-view image of the area inside the conference room and at the entrance of the conference room. Figure 2 dimensional image.

[0102] The chain positioning reasoning can be specifically:

[0103] Step 1: Decompose the sample location prompt text into multiple sub-descriptions, including: "find the meeting room", "right side of the meeting room", "find the red object", "find the red fire hydrant";

[0104] In the second step, based on the sub-description "find the conference room", the sample 2D scene image is located at the conference room, and the 2D image corresponding to the conference room is used as image 1;

[0105] Step 3: Based on the sub-description "right side of the conference room", locate the right side of the conference room from image 1 and use the corresponding two-dimensional image as image 2;

[0106] Step 4: Based on the sub-description "find red objects", find the red objects in image 2 and use the corresponding two-dimensional image as image 3;

[0107] Step 5: Find the fire hydrant in image 3 based on the sub-description "find the red fire hydrant" and output the sample positioning inference data, for example: the red fire hydrant is on the left side of the conference room entrance.

[0108] It can be understood that through the chain positioning reasoning shown above, each step is based on the reasoning result of the previous step, and combined with the language description and scene information, the scope of the object is gradually narrowed, the positioning result is optimized, and the positioning accuracy is improved.

[0109] In step S203 of some embodiments, a corresponding loss function is selected according to the original positioning inference model, and the recognition loss is calculated for the sample position information and the sample positioning inference data to obtain a position inference loss function; wherein, the loss function can be selected from a mean square error loss function, a cross entropy loss function, etc., which is not limited in the embodiments of the present application.

[0110] In step S204 of some embodiments, the back propagation method, gradient descent method, momentum update method, etc. can be used to adjust the parameters of the original positioning reasoning model based on the position reasoning loss function, so as to optimize the learning ability and generalization ability of the model and obtain the target positioning reasoning model.

[0111] In some embodiments, if the original positioning inference model is a multimodal large language model, technologies such as LoRA and Adapter can be used to fine-tune the original positioning inference model, such as adjusting the learning rate and batch size.

[0112] See also Figure 3 In some embodiments, step S103 may include but is not limited to steps S301 to S303:

[0113] Step S301: multi-view Figure 2 dimensional image to encode the image and obtain the multi-view Figure 2 dimensional features;

[0114] Step S302: encoding the target positioning prompt text to obtain the prompt text encoding feature;

[0115] Step S303: Use the target location inference model to encode the prompt text features and multi-view Figure 2 dimensional features to perform positioning reasoning and obtain initial positioning reasoning data.

[0116] In the embodiment of the present application, steps S301 to S303 are shown by Figure 2 dimensional image to encode the image and obtain the multi-view Figure 2 dimensional features, encode the target positioning prompt text, obtain the prompt text encoding features, and convert the image and text into a structured representation that the model can process. Then, the target positioning inference model projects the prompt text encoding features to the multi-view Figure 2 dimensional features, implicitly learning the correspondence between language descriptions and visual areas, thereby generating initial positioning reasoning data, which can improve the accuracy of visual positioning in complex scenes.

[0117] In step S301 of some embodiments, a pre-trained image encoder is used to decode the multi-view image. Figure 2 dimensional image for image encoding, wherein the image encoder can adopt a ResNet model, a ViT model, a CLIP image encoder, etc., but is not limited thereto.

[0118] In step S302 of some embodiments, the target positioning prompt text is text-encoded by a pre-trained text encoder, wherein the text encoder may adopt a BERT model, a CLIP text encoder, a FastText model, etc., but is not limited thereto.

[0119] See also Figure 4 In some embodiments, step S303 may include but is not limited to steps S401 to S403:

[0120] Step S401, performing text decomposition on the prompt text encoding feature to obtain text encoding sub-features;

[0121] Step S402: The text encoding sub-features and multi-view Figure 2 Perform chain reasoning based on dimensional features to obtain original positioning reasoning data;

[0122] Step S403: Perform structural processing based on the prompt text encoding features and the original positioning inference data to obtain initial positioning inference data.

[0123] In the embodiment of the present application, steps S401 to S403 are performed by decomposing the prompt text encoding features to obtain text encoding sub-features, and the text encoding sub-features and the multi-view are compared through the target positioning inference model. Figure 2 Chain reasoning based on dimensional features can gradually decouple the fine-grained semantics of language instructions, reduce semantic ambiguity in the positioning reasoning process, and obtain original positioning reasoning data. Finally, structured processing is performed based on the prompt text encoding features and the original positioning reasoning data to obtain initial positioning reasoning data, improving the accuracy of visual positioning.

[0124] Specifically, the specific implementation of step S401 to step S402 is basically the same as the specific implementation of the above-mentioned step S202, and will not be repeated here.

[0125] In step S403 of some embodiments, the prompt text encoding features and the original positioning inference data may be structured based on a preset template to obtain initial positioning inference data.

[0126] For example:

[0127] Example 1:

[0128] The target positioning prompt text is: "I am looking for a comfortable position to watch TV. Which position is best for me?"

[0129] The original positioning inference data is: "The sofa is suitable for watching TV";

[0130] The structured initial positioning inference data is: "I am willing to serve you, the location is <loc>You can watch TV while sitting on the sofa, which is the best viewing position". <loc>is the target location information to be determined.

[0131] Example 2:

[0132] The target positioning prompt text is: "Find the lesion";

[0133] The original location inference data is: "The lesion is located in the left breast lobule";

[0134] The structured initial positioning inference data is: "I am willing to serve you, and my location is <loc>The lesion is located in the left breast lobule. <loc>is the target location information to be determined.

[0135] See also Figure 5 In some embodiments, step S104 may also include but is not limited to steps S501 to S505:

[0136] Step S501, performing feature encoding on the 3D point cloud features to obtain 3D scene features;

[0137] Step S502: extracting position information from the initial positioning inference data to obtain a two-dimensional positioning feature;

[0138] Step S503: performing context extraction on the initial positioning inference data to obtain location context features;

[0139] Step S504: performing attention calculation based on the three-dimensional scene features, the two-dimensional positioning features, and the position context features to obtain initial positioning features;

[0140] Step S505 , performing feature screening on the initial positioning features to obtain target positioning features.

[0141] In the steps S501 to S505 shown in the embodiment of the present application, three-dimensional scene features are obtained by feature encoding the three-dimensional point cloud features, and position information and context information are extracted from the initial positioning reasoning data to obtain two-dimensional positioning features and position context features. Then, the attention mechanism is used to calculate the attention of the three-dimensional scene features, two-dimensional positioning features, and position context features. The depth information of the three-dimensional point cloud features is used to make up for the accuracy defects of positioning based on two-dimensional images, and the potential area of ​​the target object can be focused more accurately, so that the positioning results are consistent with both local feature matching and global semantic logic. Finally, the initial positioning features are subjected to feature screening to extract more accurate target positioning features, thereby improving the accuracy of visual positioning.

[0142] In step S501 of some embodiments, a pre-trained three-dimensional visual encoder may be used to perform feature encoding on the three-dimensional point cloud features, wherein the three-dimensional visual encoder may adopt a PointNet model, a VoxNet model, a DGCNN model, etc., but is not limited thereto.

[0143] Specifically, three-dimensional scene features may include shape information, spatial layout information, etc. of objects in the scene.

[0144] In some embodiments, position information is extracted from the initial positioning inference data to obtain a two-dimensional positioning feature; wherein the two-dimensional positioning feature is the initial position information of the target object.

[0145] By performing context extraction on the initial positioning inference data, location context features are obtained; among them, location context features contain information about spatial relationships and object attributes.

[0146] For example:

[0147] If the initial positioning inference data is: "I am willing to serve you, the location is <loc>You can sit on the sofa and watch TV, which is the best viewing position. The two-dimensional positioning feature obtained by position information extraction is "sofa"; the position context feature obtained by context extraction is "sitting on the sofa and watching TV".

[0148] If the initial positioning inference data is: "I am willing to serve you, the location is <loc>The lesion is located in the left mammary lobule. The two-dimensional positioning feature obtained by position information extraction is "left mammary lobule"; the position context feature obtained by context extraction is "the lesion is located in the left mammary lobule".

[0149] In step S504 of some embodiments, a cross-attention mechanism is used to perform attention calculation on three-dimensional scene features, two-dimensional positioning features, and position context features.

[0150] Specifically, the three-dimensional scene feature is used as the query feature (Query), the two-dimensional positioning feature is used as the key feature (Key), and the location context feature is used as the value feature (Value).

[0151] See also Figure 6 In some embodiments, step S504 includes but is not limited to steps S601 to S603:

[0152] Step S601, performing similarity calculation on the three-dimensional scene features and the two-dimensional positioning features to obtain scene matching data;

[0153] Step S602: normalize the scene matching data to obtain a scene matching weight;

[0154] Step S603 : performing weighted sum processing on the location context feature and the scene matching weight to obtain an initial positioning feature.

[0155] In the embodiment of the present application, steps S601 to S603 are performed by calculating the similarity between the three-dimensional scene features and the two-dimensional positioning features to obtain scene matching data to measure the strength of the association between the three-dimensional scene features and the two-dimensional positioning features. Then, the scene matching data is normalized to obtain scene matching weights, which can highlight key areas and suppress noise, highlighting local features with high correlation. Finally, the position context features and the scene matching weights are weighted and summed to retain semantic information and implicitly learn multi-scale spatial dependencies (such as distance and occlusion relationships) through weight distribution, thereby achieving more accurate visual positioning in complex scenes.

[0156] In step S601 of some embodiments, the dot product or cosine similarity between the three-dimensional scene features and the two-dimensional positioning features can be calculated to obtain scene matching data; wherein the scene matching data is used to represent the degree of matching between the scene area in the three-dimensional scene features and the two-dimensional positioning features.

[0157] In step S602 of some embodiments, the scene matching data can be normalized using methods such as softmax function and Z-score normalization to convert the scene matching data into scene matching weights, where the scene matching weights are used to characterize the confidence of the scene area as the target location in the three-dimensional scene features, thereby highlighting the scene area with high response.

[0158] It can be understood that the higher the response of a scene region, the more closely the scene region matches the two-dimensional positioning feature.

[0159] In step S603 of some embodiments, by performing a weighted sum calculation on the location context features and the scene matching weights, features of the scene regions with high weights are retained, and features of the scene regions with low weights are weakened.

[0160] For example, in the visual localization task of "the table to the left of the sofa," the positional context features include the geometric properties of the "table" and its relative positional relationship with the "sofa." By calculating the weighted sum of the positional context features and the scene matching weights, the final output initial localization features are both semantically relevant and spatially reasonable.

[0161] See also Figure 7 In some embodiments, step S505 may include but is not limited to steps S701 to S702:

[0162] Step S701: constructing a similarity heat map based on the initial positioning features to obtain a target feature map; wherein the target feature map is used to represent the feature similarity between each region in the three-dimensional scene feature and the two-dimensional positioning feature;

[0163] Step S702: performing regional screening on the initial positioning features based on the target feature map to obtain target positioning features.

[0164] In steps S701 and S702 of the present embodiment, a target feature map is generated by constructing a similarity heat map based on the initial positioning features. This map can characterize the feature similarity between each region representing the three-dimensional scene features and the two-dimensional positioning features. The target feature map is then subjected to regional screening of the initial positioning features to obtain target positioning features, thereby improving the accuracy of visual positioning.

[0165] In step S701 of some embodiments, the target feature map is a feature-level inference activation heat map, wherein the numerical value of each region in the target feature map can represent the degree of matching between the region in the three-dimensional scene feature and the two-dimensional positioning feature, i.e., the feature similarity.

[0166] It should be noted that the larger the value of the region in the target feature map, the more likely it is that the target object is in the region, thereby reflecting the potential location of the target object.

[0167] In step S702 of some embodiments, the target feature map can be filtered by setting a threshold to obtain high-response areas, or the top-K high-response areas can be selected from the target feature map. Based on the high-response areas, features of corresponding areas can be extracted from the initial positioning features to obtain target positioning features. This can reduce redundancy and interference, retain the most relevant features, and improve positioning accuracy.

[0168] It is understandable that the number of regions included in the target positioning feature is at least one, that is, it may include multiple regions. Therefore, it is necessary to further filter through a preset target decoding model to obtain the final target position information.

[0169] In step S105 of some embodiments, the target positioning feature is decoded using a preset target decoding model to obtain target position information.

[0170] The preset target decoding model is a Transformer decoder, which consists of M Transformer decoding layers and a prediction head. Each decoding layer contains two cross-attention layers: a text feature cross-attention layer and a scene feature cross-attention layer.

[0171] Among them, the text feature cross-attention layer is used to process the relationship between target positioning features and initial positioning inference data, and the scene feature cross-attention layer is used to process the relationship between target positioning features and three-dimensional scene features.

[0172] Specifically, the target localization feature is first updated through a cross-attention layer with text features and combined with the initial localization inference data. The updated target localization feature is then further updated through a cross-attention layer with scene features and combined with the target localization feature. Finally, the prediction head uses the updated target localization feature as input to predict the 3D position of the target object and a matching score. The matching score indicates how closely the prediction matches the target object. It should be noted that the 3D position of the target object is a coordinate value (X, Y, Z).

[0173] Finally, the three-dimensional position of the target object is combined with the initial positioning inference data to obtain the target position information. Specifically, the three-dimensional position of the target object is used as <loc>content.

[0174] The visual positioning method provided in the embodiments of the present application can be applied to multiple reference scenarios, for example:

[0175] In medical technology scenarios, visual positioning methods can be applied to intelligent diagnosis of medical images. Combined with the descriptions of medical staff (such as "ground-glass nodules in the left lower lobe of the lung"), they can automatically locate specific lesion areas in lung CT point cloud data, assisting medical staff in rapid labeling and diagnosis, reducing manual screening time, and are suitable for scenarios such as early screening of lung cancer.

[0176] In addition, visual positioning methods can also be applied to surgical navigation. By parsing the doctor's natural language instructions (such as "locate the tumor edge"), key anatomical structures or lesion areas in the patient's CT / MRI point cloud data can be identified and highlighted in real time, thereby improving surgical accuracy and safety.

[0177] In fintech scenarios, visual positioning methods can be applied to bank branches, creating high-precision spatial maps of bank branches through 3D point cloud scanning. When customers or staff use voice queries (such as "locate the ATM" or "find the receipt printer on the shelf"), the system can mark the target location on the high-precision spatial map of the bank branch in real time, improving service efficiency.

[0178] See also Figure 8 The present application also provides a visual positioning device that can implement the above-mentioned visual positioning method. The device includes:

[0179] Environmental information acquisition module 801, used to obtain multiple views of the target environment Figure 2 3D image and 3D point cloud features;

[0180] The prompt text acquisition module 802 is used to obtain a target positioning prompt text; wherein the target positioning prompt text is used to indicate the location of the target object in the target environment;

[0181] Position reasoning module 803, used for positioning the target prompt text and multi-view through the preset target positioning reasoning model Figure 2 3D image to perform position reasoning and obtain initial positioning reasoning data;

[0182] The visual positioning module 804 is used to perform visual positioning based on the initial positioning inference data and the three-dimensional point cloud features to obtain target positioning features;

[0183] The feature decoding module 805 is used to perform feature decoding based on the target positioning feature to obtain target position information; wherein the target position information is used to indicate the position of the target object in the target environment.

[0184] The specific implementation of the visual positioning device is basically the same as the specific embodiment of the above-mentioned visual positioning method, and will not be repeated here.

[0185] The present application also provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned visual positioning method when executing the computer program. The electronic device can be any smart terminal including a tablet computer, an in-vehicle computer, or the like.

[0186] See also Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:

[0187] The processor 901 can be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0188] The memory 902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program codes are stored in the memory 902 and are called by the processor 901 to execute the visual positioning method of the embodiments of this application.

[0189] Input / output interface 903, used to implement information input and output;

[0190] Communication interface 904, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);

[0191] Bus 905 , which transmits information between various components of the device (e.g., processor 901 , memory 902 , input / output interface 903 , and communication interface 904 );

[0192] The processor 901 , the memory 902 , the input / output interface 903 and the communication interface 904 are connected to each other in communication within the device via a bus 905 .

[0193] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned visual positioning method is implemented.

[0194] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0195] The visual positioning method and device, electronic device and storage medium provided in the embodiments of the present application obtain multiple visual positions of the target environment. Figure 2 3D image and 3D point cloud features, and obtain target positioning prompt text; wherein, the target positioning prompt text is used to indicate the location of the target object in the target environment. Then, the target positioning prompt text and multi-view are compared using the target positioning inference model. Figure 2 The system uses the initial positioning inference data and the 3D point cloud features to perform position reasoning on the 3D image, forming a structured reasoning chain, namely the initial positioning inference data, thereby converting the ambiguous text description into an intermediate representation of the implicit spatial relationship. Furthermore, visual positioning is performed based on the initial positioning inference data and the 3D point cloud features to obtain the target positioning features. Finally, the target positioning features are feature decoded to generate target position information. The target position information can intuitively indicate the position of the target object in the target environment, significantly improving the accuracy of visual positioning in complex scenes.

[0196] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0197] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0198] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0199] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.

[0200] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0201] It should be understood that in this application, "at least one (item)" means one or more, and "more" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or plural.

[0202] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0203] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0204] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0205] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0206] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.< / loc> < / loc> < / loc> < / loc> < / loc> < / loc> < / loc>

Claims

1. A visual positioning method, characterized in that: The method comprises: Acquire multi-view 2D images and 3D point cloud features of the target environment; Obtaining a target positioning prompt text; wherein the target positioning prompt text is used to indicate the location of the target object in the target environment; Performing position reasoning on the target positioning prompt text and the multi-view two-dimensional image using a preset target positioning reasoning model to obtain initial positioning reasoning data; Perform visual positioning based on the initial positioning inference data and the three-dimensional point cloud features to obtain target positioning features; Feature decoding is performed based on the target positioning feature to obtain target position information; wherein the target position information is used to indicate the position of the target object in the target environment.

2. The method according to claim 1, characterized in that The performing position reasoning on the target positioning prompt text and the multi-view two-dimensional image by using a preset target positioning reasoning model to obtain initial positioning reasoning data includes: performing image encoding on the multi-view two-dimensional image to obtain multi-view two-dimensional features; Performing text encoding on the target positioning prompt text to obtain prompt text encoding features; The target positioning reasoning model is used to perform positioning reasoning on the prompt text encoding features and the multi-view two-dimensional features to obtain the initial positioning reasoning data.

3. The method according to claim 2, characterized in that The performing positioning reasoning on the prompt text encoding feature and the multi-view two-dimensional feature by the target positioning reasoning model to obtain the initial positioning reasoning data includes: Performing text decomposition on the prompt text encoding feature to obtain text encoding sub-features; Performing chain reasoning on the text encoding sub-features and the multi-view two-dimensional features through the target positioning reasoning model to obtain original positioning reasoning data; Structural processing is performed based on the coding features of the prompt text and the original positioning inference data to obtain the initial positioning inference data.

4. The method according to claim 1, wherein The performing visual positioning based on the initial positioning inference data and the three-dimensional point cloud features to obtain target positioning features includes: Performing feature encoding on the three-dimensional point cloud features to obtain three-dimensional scene features; Extracting position information from the initial positioning inference data to obtain a two-dimensional positioning feature; Performing context extraction on the initial positioning inference data to obtain location context features; Performing attention calculation based on the three-dimensional scene feature, the two-dimensional positioning feature, and the position context feature to obtain an initial positioning feature; Feature screening is performed on the initial positioning features to obtain the target positioning features.

5. The method according to claim 4, characterized in that The performing attention calculation based on the three-dimensional scene feature, the two-dimensional positioning feature, and the position context feature to obtain the initial positioning feature includes: Calculating similarity between the three-dimensional scene features and the two-dimensional positioning features to obtain scene matching data; Normalizing the scene matching data to obtain a scene matching weight; The location context feature and the scene matching weight are weighted and summed to obtain the initial positioning feature.

6. The method according to claim 4, characterized in that The performing feature screening on the initial positioning feature to obtain the target positioning feature includes: A similarity heat map is constructed based on the initial positioning features to obtain a target feature map; wherein the target feature map is used to represent the feature similarity between each region in the three-dimensional scene feature and the two-dimensional positioning feature; The initial positioning feature is subjected to regional screening based on the target feature map to obtain the target positioning feature.

7. The method according to any one of claims 1 to 6, characterized in that Before performing position reasoning on the target positioning prompt text and the multi-view two-dimensional image using a preset target positioning reasoning model to obtain initial positioning reasoning data, the method further includes: Acquire visual positioning sample data, wherein the visual positioning sample data includes a sample two-dimensional scene image, a sample positioning prompt text, and sample position information; Performing position reasoning on the sample two-dimensional scene image and the sample positioning prompt text using a preset original positioning reasoning model to obtain sample positioning reasoning data; Performing recognition loss calculation based on the sample position information and the sample positioning inference data to obtain a position inference loss function; The parameters of the original positioning reasoning model are adjusted based on the position reasoning loss function to obtain the target positioning reasoning model.

8. A visual positioning device, characterized in that: The device comprises: Environmental information acquisition module, used to obtain multi-view two-dimensional images and three-dimensional point cloud features of the target environment; A prompt text acquisition module is used to acquire a target positioning prompt text; wherein the target positioning prompt text is used to indicate the location of the target object in the target environment; A position reasoning module is used to perform position reasoning on the target positioning prompt text and the multi-view two-dimensional image using a preset target positioning reasoning model to obtain initial positioning reasoning data; A visual positioning module is used to perform visual positioning based on the initial positioning inference data and the three-dimensional point cloud features to obtain target positioning features; A feature decoding module is used to perform feature decoding based on the target positioning feature to obtain target position information; wherein the target position information is used to indicate the position of the target object in the target environment.

9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.