Positioning model training method and target area positioning method and device

By constructing semantic appearance maps and semantic transient maps and training them in combination with spatial model and network model architectures, the problem of low positioning accuracy in the positioning model in the existing technology in complex environments is solved, and higher positioning accuracy and robustness are achieved.

CN120125959APending Publication Date: 2025-06-10BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510154816.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In the prior art, when the positioning model processes pictures with complex environments and rich details, it is difficult to fully capture the complex features of the scenes in the picture, resulting in a reduced positioning accuracy.

Method used

By acquiring the original image and multiple rendered images, semantic appearance maps and semantic transient maps are extracted and constructed, and model training is combined with preset spatial models and network model architectures to generate positioning models.

Benefits of technology

By introducing semantic appearance maps and semantic transient maps, the positioning model can better understand and distinguish semantic information changes caused by appearance differences and occlusion, thereby improving positioning accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125959A_ABST
    Figure CN120125959A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a training method of a positioning model and a target area positioning method and device, and relates to the technical field of computer vision. The training method of the positioning model comprises the following steps: acquiring an original image and a rendered image; performing feature extraction on the original image and the rendered image to obtain original image features and rendered image features; respectively constructing a semantic appearance graph and a semantic transient graph according to the original image features and the rendered image features; respectively carrying out model training according to the original image features, the rendered image features, the semantic appearance graph, the semantic transient graph, the spatial model architecture and the network model architecture to obtain a spatial model and a network model; and generating a positioning model according to the space model and the network model. Semantic appearance maps may capture semantic information variations caused by appearance differences, and semantic transient maps may capture semantic information variations caused by occlusions or transient objects. And by introducing a semantic appearance graph and a semantic transient graph, the accuracy of the positioning model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular, to a method for training a positioning model, a method and device for target area positioning. Background Art

[0002] With the development of Internet technology, users can access many pictures with different styles on the Internet. These pictures can be used to construct a three-dimensional model of the scene through deep learning technology. Then, with the help of the constructed three-dimensional model, users can intuitively locate and understand the details of the items in the picture.

[0003] In the prior art, a large number of pictures are usually collected and the items in the pictures are labeled, and a deep learning algorithm is used to train a positioning model to accurately locate the items in the pictures; after the positioning model is trained, it can be applied to new pictures to achieve the positioning of the items in the new pictures.

[0004] However, the prior art has the problem of low positioning accuracy. Although the prior art trains a positioning model through a deep learning algorithm and can locate the items in the pictures, when dealing with pictures with complex environments and rich details, the positioning models in the prior art often have difficulty fully capturing the complex features of the scenes in the pictures, resulting in a decrease in the positioning accuracy of the models. Summary of the Invention

[0005] Embodiments of this application provide a method for training a positioning model, a method and device for target area positioning, so as to solve the problem of low positioning accuracy in the prior art.

[0006] In a first aspect, an embodiment of this application provides a method for training a positioning model, including:

[0007] Obtain an original image and a plurality of rendered images; wherein, the plurality of rendered images are obtained by rendering the original image;

[0008] Extract features from the original image to obtain a plurality of original image features;

[0009] Extract features from the plurality of rendered images to obtain a plurality of rendered image features;

[0010] Construct a semantic appearance graph according to the plurality of rendered image features; wherein, the semantic appearance graph is used to represent the variability of semantic information caused by appearance differences in the original image;

[0011] Construct a semantic transient graph according to the plurality of original image features and the plurality of rendered image features; wherein, the semantic transient graph is used to represent the variability of semantic information caused by occluders in the original image;

[0012] Model training is performed according to the multiple original image features, the multiple rendered image features, the semantic appearance map, the semantic transient map, and a preset spatial model architecture to obtain a spatial model; wherein, the spatial model is used to represent the three-dimensional scene where the original image is located.

[0013] Model training is performed according to the multiple original image features, the semantic transient map, and a preset network model architecture to obtain a network model; wherein, the network model is used to compress and decompress the multiple original image features and the multiple rendered image features.

[0014] A positioning model is generated according to the spatial model and the network model.

[0015] In one possible design, the model training according to the multiple original image features, the multiple rendered image features, the semantic appearance map, the semantic transient map, and a preset spatial model architecture to obtain a spatial model includes:

[0016] Construct spatial constraint conditions according to the semantic appearance map and the semantic transient map;

[0017] Model training is performed according to the spatial constraint conditions, the multiple original image features, the multiple rendered image features, and the spatial model architecture to obtain the spatial model.

[0018] In one possible design, the model training according to the multiple original image features, the semantic transient map, and a preset network model architecture to obtain a network model includes:

[0019] Construct network constraint conditions according to the semantic transient map;

[0020] Model training is performed according to the network constraint conditions, the multiple original image features, and the network model architecture to obtain the network model.

[0021] In one possible design, the obtaining of the original image and the multiple rendered images includes:

[0022] Obtain the original image and the camera pose corresponding to the original image;

[0023] Perform image rendering according to the camera pose and a preset scene radiance field to obtain the multiple rendered images.

[0024] In one possible design, the feature extraction of the original image to obtain multiple original image features includes:

[0025] Segment the original image to obtain multiple segmentation regions; wherein, the multiple segmentation regions refer to multiple pixel blocks segmented from the original image;

[0026] Crop the multiple segmentation regions to obtain multiple cropped regions; wherein, the multiple cropped regions refer to multiple pixel blocks cropped from the multiple segmentation regions;

[0027] Input the multiple cropped regions into a preset feature extraction model to obtain the multiple original image features.

[0028] In a second aspect, an embodiment of the present application provides a method for target region localization, including:

[0029] Obtain the camera pose and the query text;

[0030] Input the camera pose into the localization model to obtain multiple compressed features; wherein, the localization model is trained by the method according to any item in the first aspect;

[0031] Calculate the correlation between the query text and the multiple compressed features to obtain multiple correlation scores;

[0032] Generate a correlation score map according to the multiple correlation scores;

[0033] Extract the region in the correlation score map where the correlation score is greater than a preset threshold as the target region; wherein, the target region is the region queried by the query text.

[0034] In a third aspect, the present application provides a training device for a localization model, and the device includes:

[0035] An image acquisition module, configured to acquire an original image and multiple rendered images; wherein, the multiple rendered images are obtained by rendering the original image;

[0036] An original image feature extraction module, configured to extract features from the original image to obtain multiple original image features;

[0037] A rendered image feature extraction module, configured to extract features from the multiple rendered images to obtain multiple rendered image features;

[0038] A semantic appearance map construction module, configured to construct a semantic appearance map according to the multiple rendered image features; wherein, the semantic appearance map is used to represent the semantic information variability caused by appearance differences in the original image;

[0039] A semantic transient graph construction module, configured to construct a semantic transient graph according to the multiple original image features and the multiple rendered image features; wherein, the semantic transient graph is used to represent the semantic information variability caused by occluders in the original image;

[0040] A spatial model training module, configured to perform model training according to the multiple original image features, the multiple rendered image features, the semantic appearance graph, the semantic transient graph, and a preset spatial model architecture to obtain a spatial model; wherein, the spatial model is used to represent the three-dimensional scene where the original image is located;

[0041] A network model training module, configured to perform model training according to the multiple original image features, the semantic transient graph, and a preset network model architecture to obtain a network model; wherein, the network model is used to compress and decompress the multiple original image features and the multiple rendered image features;

[0042] A positioning model generation module, configured to generate a positioning model according to the spatial model and the network model.

[0043] In a possible design, the spatial model training module includes:

[0044] A spatial constraint condition construction unit, configured to construct spatial constraint conditions according to the semantic appearance graph and the semantic transient graph;

[0045] A spatial model training unit, configured to perform model training according to the spatial constraint conditions, the multiple original image features, the multiple rendered image features, and the spatial model architecture to obtain the spatial model.

[0046] In a possible design, the network model training module includes:

[0047] A network constraint condition construction unit, configured to construct network constraint conditions according to the semantic transient graph;

[0048] A network model training unit, configured to perform model training according to the network constraint conditions, the multiple original image features, and the network model architecture to obtain the network model.

[0049] In a possible design, the image acquisition module includes:

[0050] A camera pose acquisition unit, configured to acquire the original image and the camera pose corresponding to the original image;

[0051] An image rendering unit, configured to perform image rendering according to the camera pose and a preset scene radiance field to obtain the multiple rendered images.

[0052] In a possible design, the original image feature extraction module includes:

[0053] A segmentation unit for segmenting the original image to obtain a plurality of segmentation regions; wherein, the plurality of segmentation regions refer to a plurality of pixel blocks segmented from the original image;

[0054] A cropping unit for cropping the plurality of segmentation regions to obtain a plurality of cropped regions; wherein, the plurality of cropped regions refer to a plurality of pixel blocks cropped from the plurality of segmentation regions;

[0055] A feature extraction unit for inputting the plurality of cropped regions into a preset feature extraction model to obtain the plurality of original image features.

[0056] In a fourth aspect, the present application provides a target area positioning device, and the device includes:

[0057] A pose text acquisition module for acquiring a camera pose and a query text;

[0058] A feature calculation module for inputting the camera pose into a positioning model to obtain a plurality of compressed features; wherein, the positioning model is trained by the device described in the third aspect;

[0059] A correlation calculation module for calculating the correlation between the query text and the plurality of compressed features to obtain a plurality of correlation scores;

[0060] A score map generation module for generating a correlation score map according to the plurality of correlation scores;

[0061] A target area determination module for extracting the area in the correlation score map where the correlation score is greater than a preset threshold as the target area; wherein, the target area is the area queried by the query text.

[0062] In a fifth aspect, the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;

[0063] The memory stores computer execution instructions;

[0064] When the processor executes the computer execution instructions stored in the memory, it is used to implement the training method of the positioning model described in any item of the first aspect, or to implement the target area positioning method described in the second aspect.

[0065] Sixth aspect, the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the training method of the positioning model as described in any item of the first aspect, or to implement the target area positioning method as described in the second aspect.

[0066] A training method of a positioning model, a target area positioning method and device provided by the present application. The training method of the positioning model includes: obtaining an original image and a plurality of rendered images; extracting features from the original image to obtain a plurality of original image features; extracting features from the plurality of rendered images to obtain a plurality of rendered image features; constructing a semantic appearance graph according to the plurality of rendered image features; constructing a semantic transient graph according to the plurality of original image features and the plurality of rendered image features; performing model training according to the plurality of original image features, the plurality of rendered image features, the semantic appearance graph, the semantic transient graph and a preset spatial model architecture to obtain a spatial model; performing model training according to the plurality of original image features, the semantic transient graph and a preset network model architecture to obtain a network model; generating a positioning model according to the spatial model and the network model. In the training method of the positioning model of the present application, a semantic appearance graph and a semantic transient graph are introduced. The semantic appearance graph can capture the semantic information changes caused by appearance differences in the original image. This enables the positioning model to understand and distinguish feature differences caused by illumination, color or texture changes, thereby reducing semantic ambiguity; the semantic transient graph can capture the semantic information changes caused by occluders or transient objects in the original image. By quantifying these changes, the positioning model can identify and filter out the interference caused by instantaneous occluders, thereby avoiding semantic contamination. By introducing the semantic appearance graph and the semantic transient graph, the accuracy of the positioning model is improved. Description of the Drawings

[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0068] Figure 1 It is a schematic diagram of the application scenario of the training method of the positioning model provided by the embodiment of the present application;

[0069] Figure 2 It is a schematic flowchart of the training method of the positioning model provided by the embodiment of the present application;

[0070] Figure 3 It is a schematic flowchart of the target area positioning method provided by the embodiment of the present application;

[0071] Figure 4 Schematic diagram of the structure of the training device for the positioning model provided by the embodiment of the present application;

[0072] Figure 5 Schematic diagram of the structure of the target area positioning device provided by the embodiment of the present application;

[0073] Figure 6 Schematic diagram of the structure of the electronic device provided by the embodiment of the present application. Detailed implementation manners

[0074] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of the devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0075] In the embodiments of the present application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. Those skilled in the art can understand that the terms "first" and "second" do not limit the quantity and execution order, and the terms "first" and "second" do not necessarily limit to be different. It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific manner. In the embodiments of the present application, "at least one" means one or more, and "a plurality of" means two or more.

[0076] It should be noted that "when... " in the embodiments of the present application can be at the instant when a certain situation occurs, or within a period of time after a certain situation occurs. The embodiments of the present application do not make specific limitations on this. In addition, the training method, target area positioning method and device provided by the embodiments of the present application are only examples, and the training method, target area positioning method and device may also include more or less content.

[0077] To facilitate a clear description of the technical solutions of the embodiments of the present application, the following briefly introduces some terms and technologies involved in the embodiments of the present application:

[0078] Rendered Image: An image generated through computer image generation techniques. These images can typically simulate real-world lighting, material, and shadow effects to realistically display virtual scenes or objects.

[0079] Semantic Appearance Map: A graphical structure used to represent the semantic information changes in an image caused by appearance differences, such as lighting variations and different materials. By analyzing and encoding the changes in image features, it helps machine learning models understand and distinguish the semantic relationships and appearance features of different elements in the image.

[0080] Semantic Transient Map: A graphical structure used to represent the semantic information changes in an image caused by temporary change factors, such as occluders and dynamic objects. By capturing and encoding the features of these transient changes, it helps the model understand the transient and unstable semantic relationships in the image, thereby enhancing the analysis and recognition capabilities for dynamic scenes.

[0081] Relevance Score Map: A graphical representation used to visualize and quantify the relationship between query text and image features. By calculating the relevance scores of each image region with the query text and mapping these scores onto the image, it highlights the regions most relevant to the query content, thus helping to identify and locate the target regions in the image.

[0082] Here, exemplary embodiments will be described in detail, and their examples are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.

[0083] The technical solution of the present invention will be described in detail below with specific embodiments. These several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will be described below in conjunction with the accompanying drawings.

[0084] To clearly understand the technical solution of this application, the solutions of the prior art will be introduced in detail first. A vast amount of image resources provides a rich data foundation for 3D modeling technology. By using these images for 3D modeling, realistic 3D models of scenes can be generated. Then, with the help of positioning technology, users can not only identify the shape and position of the items in the image but also gain in-depth understanding of their detailed information. This ability has important application values in many fields. For example, in virtual reality and augmented reality, accurate 3D models and item positioning can enhance the immersive experience of users.

[0085] In the prior art, a large amount of image data is usually collected, and the objects in these images are detailedly annotated to create a rich training dataset. Then, using these annotated data, a localization model is trained through deep learning algorithms. The deep learning algorithms can automatically learn and extract features in the images, thereby constructing a model capable of recognizing and localizing objects. During the training process, the model continuously adjusts its internal parameters to improve the recognition accuracy and generalization ability for the training data. Once the localization model is trained, it can be applied to new images to automatically identify and localize the objects therein. However, although the prior art uses deep learning algorithms to train a localization model to achieve the recognition and localization of objects in images, there is still a problem of insufficient localization accuracy. Especially when dealing with images with complex environments and rich details, existing models often have difficulty fully capturing and analyzing the complex features in these images. This is because there may be various interference factors in complex scenes, such as illumination changes, occlusions, cluttered backgrounds, and similarities between objects. The existing localization models in the prior art often have difficulty fully capturing the complex features in such images, which leads to a decrease in the localization accuracy of the localization model.

[0086] Therefore, aiming at the problem of low localization accuracy of the localization model in the prior art, it is found in the research that to solve this problem, the localization accuracy of the model for three-dimensional scenes and object localization can be improved by extracting various feature information in the images and combining semantic changes and occlusion effects in complex scenes for modeling: ① Collect and generate multi-view image data, simulate different scenes, and enrich the training dataset. This method can help the model learn highly robust features and improve its ability to recognize and localize objects in complex environments. ② Feature alignment technology can be introduced to correct the feature deviation caused by appearance differences and occlusions by comparing the features of the original image and the rendered image. This can ensure the consistency of the features extracted by the model under different conditions, thereby improving the localization accuracy. ③ A multi-task learning framework can be adopted, taking appearance changes and occlusion processing as auxiliary tasks and jointly training them with the main localization task. By sharing features and parameters, the model can better capture the correlations between different tasks during the learning process and improve the overall localization performance.

[0087] Specifically:

[0088] First, the training dataset can be enriched by generating diverse rendered images, enabling the model to better adapt to various environmental changes, such as lighting and texture differences. Second, develop specialized algorithms to identify occlusion effects for more accurate prediction of the features of occluded objects. At the same time, apply feature alignment techniques to correct the feature deviations caused by appearance differences and occlusions, ensuring the consistency of feature extraction. In addition, design a dynamic update mechanism to enable the semantic appearance map and semantic transient map to be adjusted according to new inputs, improving the model's real-time adaptability. Finally, adopt a multi-task learning strategy, treating the handling of appearance changes and occlusions as auxiliary tasks and jointly training them with the main localization task to enhance the model's performance in complex scenarios.

[0089] The training method of the localization model in the embodiments of this application introduces a semantic appearance map and a semantic transient map. The semantic appearance map can capture the semantic information changes caused by appearance differences in the original image. This enables the localization model to understand and distinguish feature differences caused by lighting, color, or texture changes, thereby reducing semantic ambiguity. The semantic transient map can capture the semantic information changes caused by occluders or transient objects in the original image. By quantifying these changes, the localization model can identify and filter out the interference caused by instantaneous occluders, thus avoiding semantic contamination. By introducing the semantic appearance map and the semantic transient map, the accuracy of the localization model is improved.

[0090] Based on the above creative findings, the technical solution of this application is proposed.

[0091] The application scenarios of the training method of the localization model provided in the embodiments of the present invention will be introduced below. Figure 1 It is a schematic diagram of the application scenario of the training method of the localization model provided in the embodiments of this application. As Figure 1 shown, this application scenario includes a mobile terminal 101 and a server 102. The mobile terminal 101 collects the original image and sends the original image to the server 102. The server 102 performs image rendering on the original image to obtain multiple rendered images. The server 102 performs model training based on the original image, multiple rendered images, and a preset model architecture to obtain a localization model. The mobile terminal 101 sends a query text to the server 102 again. The server 102 performs localization based on the query text and the trained localization model and sends the localization result to the mobile terminal 101.

[0092] The embodiments of the present invention will be introduced below in conjunction with the accompanying drawings of the specification.

[0093] Figure 2 It is a schematic flowchart of the training method of the localization model provided in the embodiments of this application. As Figure 2 shown, in this embodiment, the execution subject of the embodiments of the present invention is the server. Then the training method of the localization model provided in this embodiment includes the following steps:

[0094] S201. Obtain the original image and multiple rendered images.

[0095] Specifically, the original image of the real scene can be obtained by using an image acquisition device such as a camera or a sensor, and these original images are rendered using computer graphics technology to generate multiple rendered images with different appearance features. By generating diverse rendered images, the robustness of the positioning model to different appearance changes and occlusion situations can be enhanced, and the positioning accuracy and generalization ability of the positioning model in different environments can be improved. Among them, the multiple rendered images are obtained by rendering the original image.

[0096] S202. Extract features from the original image to obtain multiple original image features.

[0097] Specifically, a deep learning model such as a convolutional neural network can be used to process the original image and extract representative feature vectors or feature maps therefrom. These features can capture key information in the image, such as edges, textures, and shapes. By extracting the features of the original image, a rich information basis can be provided for subsequent model training, enabling the positioning model to accurately understand and represent the semantic content in the image and improving the performance and accuracy of the positioning model.

[0098] S203. Extract features from the multiple rendered images to obtain multiple rendered image features.

[0099] Specifically, a deep learning algorithm such as a convolutional neural network can be applied to analyze each rendered image and extract feature vectors or feature maps. These features capture the key visual information in the rendered image, including changes in color, texture, and shape. By extracting the features of the rendered images, the model can be helped to understand the semantic information changes caused by appearance changes such as lighting and material changes, enhance the adaptability of the model to different visual conditions, and thus improve the robustness and accuracy of the positioning model in diverse scenarios.

[0100] For example, when extracting features from multiple rendered images to obtain multiple rendered image features, the formula for feature extraction can be:

[0101]

[0102] Among them, is the rendered image feature, is the rendered image, is the feature extraction model.

[0103] S204. Construct a semantic appearance graph based on the multiple rendered image features.

[0104] Specifically, by analyzing and integrating the features extracted from multiple rendered images, and using a graph structure or a graph neural network to represent the relationships and differences between these features, a semantic appearance graph can be obtained. The semantic appearance graph captures the changes in semantic information caused by appearance changes in the form of nodes and edges. This processing can effectively express and quantify the semantic changes caused by appearance differences, enabling the model to better understand and adapt to different visual conditions, and improving the performance and robustness of the localization model in diverse environments. Among them, the semantic appearance graph is used to represent the variability of semantic information caused by appearance differences in the original image.

[0105] For example, the semantic appearance variability can be calculated based on the features of multiple rendered images, and then a semantic appearance graph can be constructed according to the semantic appearance variability. The formula for calculating the semantic information variability can be:

[0106]

[0107] Where, is the semantic variability, is the feature of the rendered image, is the average value of the features of the rendered images, and N is the upper limit of the number of features of the rendered images.

[0108] S205. Construct a semantic transient graph based on multiple original image features and multiple rendered image features.

[0109] Specifically, by comparing and analyzing the differences between the original image features and the rendered image features, and using a graph structure or a graph neural network to represent the semantic changes caused by these differences, a semantic transient graph can be obtained. The semantic transient graph captures the changes in semantic information caused by occluders or dynamic changes in the form of nodes and edges. This processing can effectively identify and express the changes in semantic information caused by occlusion or dynamic changes, enabling the model to better handle the dynamic elements in the scene, and improving the accuracy and robustness of the localization model in complex environments. Among them, the semantic transient graph is used to represent the variability of semantic information caused by occluders in the original image.

[0110] For example, the semantic transient variability can be calculated based on multiple original image features and multiple rendered image features, and then a semantic transient graph can be constructed according to the semantic transient variability. The formula for calculating the semantic transient variability can be:

[0111]

[0112] Where, is the semantic transient variability, is the feature of the rendered image, is the feature of the original image, and T is the number of original images.

[0113] S206. Perform model training based on multiple original image features, multiple rendered image features, semantic appearance maps, semantic transient maps, and a preset spatial model architecture to obtain a spatial model.

[0114] Specifically, these features and graph structures can be input into a preset deep learning architecture, which may include convolutional neural networks, graph neural networks, and other network layers suitable for processing spatial information. During the training process, by optimizing the loss function, the model can learn how to map the input features and graph structures to the representation of three-dimensional spatial information. The generated spatial model can comprehensively consider the impacts of appearance changes and dynamic changes on scene understanding, thereby accurately representing the three-dimensional scene where the original image is located, and improving the accuracy and robustness of the localization model in different environments. Among them, the spatial model is used to represent the three-dimensional scene where the original image is located.

[0115] S207. Perform model training based on multiple original image features, semantic transient maps, and a preset network model architecture to obtain a network model.

[0116] Specifically, the original image features and semantic transient maps can be used as input data, and combined with a preset network model architecture, such as a convolutional neural network or a graph neural network, for training. During the training process, the network model learns how to effectively compress and decompress these feature information to capture the key semantic changes in the image. By integrating the original image features and semantic transient maps, the network model can accurately understand and process the semantic information changes caused by occlusion or dynamic changes, thereby improving the performance of the model in complex scenes and its adaptability to dynamic elements. Among them, the network model is used to compress and decompress multiple original image features and multiple rendered image features.

[0117] S208. Generate a localization model based on the spatial model and the network model.

[0118] Specifically, the outputs of the spatial model and the network model can be fused to construct a comprehensive localization model by combining the strengths of both. The spatial model provides an understanding of the three-dimensional scene, while the network model effectively compresses and decompresses feature information. By integrating the capabilities of these two models, the localization model can accurately identify and locate targets in diverse and complex environments. The generated localization model can comprehensively utilize spatial information and semantic features to improve the localization accuracy and robustness under different visual conditions and scene dynamic changes.

[0119] A training method for a positioning model provided in this embodiment includes: obtaining an original image and a plurality of rendered images; extracting features from the original image to obtain a plurality of original image features; extracting features from the plurality of rendered images to obtain a plurality of rendered image features; constructing a semantic appearance graph according to the plurality of rendered image features; constructing a semantic transient graph according to the plurality of original image features and the plurality of rendered image features; performing model training according to the plurality of original image features, the plurality of rendered image features, the semantic appearance graph, the semantic transient graph, and a preset spatial model architecture to obtain a spatial model; performing model training according to the plurality of original image features, the semantic transient graph, and a preset network model architecture to obtain a network model; and generating a positioning model according to the spatial model and the network model. A training method for a positioning model achieves the following technical effects: The training method for the positioning model introduces a semantic appearance graph and a semantic transient graph. The semantic appearance graph can capture semantic information changes caused by appearance differences in the original image. This enables the positioning model to understand and distinguish feature differences caused by lighting, color, or texture changes, thereby reducing semantic ambiguity; the semantic transient graph can capture semantic information changes caused by occluders or transient objects in the original image. By quantifying these changes, the positioning model can identify and filter out interference caused by instantaneous occluders, thereby avoiding semantic contamination. By introducing the semantic appearance graph and the semantic transient graph, the accuracy of the positioning model is improved.

[0120] In a possible design, S206, performing model training according to the plurality of original image features, the plurality of rendered image features, the semantic appearance graph, the semantic transient graph, and a preset spatial model architecture to obtain a spatial model, includes:

[0121] S2061, constructing a spatial constraint condition according to the semantic appearance graph and the semantic transient graph.

[0122] Specifically, by analyzing the relationships between nodes and edges in the semantic appearance graph and the semantic transient graph, patterns and rules of semantic information changes caused by appearance changes and occlusions can be identified. These patterns and rules can be transformed into spatial constraint conditions, serving as prior knowledge or rules for model training. These constraint conditions can help the model understand and adapt to complex semantic changes in the three-dimensional scene during training. By introducing spatial constraint conditions, the model can accurately capture and represent semantic information in the three-dimensional scene, improving the accuracy and robustness of the positioning model in complex environments.

[0123] Among them, the spatial constraint condition can be expressed as:

[0124]

[0125] Among them, is an encoder, is a decoder, is the t-th original image feature, is the semantic transient variability for the t-th

[0126] S2062. Perform model training according to the spatial constraint conditions, multiple original image features, multiple rendered image features, and the spatial model architecture to obtain a spatial model.

[0127] Specifically, the spatial constraint conditions can be combined with the features of the original and rendered images and input into a preset spatial model architecture for training. The spatial constraint conditions provide prior information about semantic changes in the scene, while the image features provide rich visual information. During the training process, the model learns how to accurately represent and reconstruct the 3D scene under these conditions through optimization algorithms. By combining spatial constraints and image features, the spatial model can precisely understand and describe the 3D environment, improving the localization ability and robustness in diverse and complex scenes.

[0128] The technical effect of this solution in this embodiment is that by constructing spatial constraint conditions using the semantic appearance map and semantic transient map, and combining these constraint conditions with the original image features, rendered image features, and the spatial model architecture for model training, the understanding and representation ability of the spatial model for the 3D scene are enhanced. This method effectively integrates appearance changes and occlusion information, providing more accurate scene reconstruction and semantic information capture capabilities, thereby improving the accuracy and robustness of the localization model in complex and dynamic environments.

[0129] In a possible design, S207. Perform model training according to multiple original image features, the semantic transient map, and a preset network model architecture to obtain a network model, including:

[0130] S2071. Construct network constraint conditions according to the semantic transient map.

[0131] Specifically, by analyzing the relationships between nodes and edges in the semantic transient map, the patterns and rules of semantic information changes caused by occluders can be identified. These patterns and rules can be transformed into network constraint conditions, serving as prior knowledge or rules for network model training. These constraint conditions can help the network model understand and adapt to dynamic changing scene elements during the training process. By introducing network constraint conditions, the network model can accurately process and represent semantic information changes caused by occlusion or dynamic changes, thereby improving the adaptability and processing accuracy of the model in complex environments.

[0132] For example, the network constraint conditions can be expressed as:

[0133]

[0134] where are the compressed original image features and rendered image features is the average of the compressed original image features and the rendered image features, is the semantic map, is the t-th semantic variability, is the t-th transient variability.

[0135] S2072. Perform model training according to network constraint conditions, multiple original image features, and the network model architecture to obtain a network model.

[0136] Specifically, the network constraint conditions can be combined with the original image features and input into a preset network model architecture for training. The network constraint conditions provide prior knowledge about semantic information changes caused by occlusion or dynamic changes, while the original image features provide rich visual information. During the training process, the model learns how to effectively compress and decompress image features under these conditions through optimization algorithms to accurately capture semantic changes in dynamic scenes. By combining network constraints and image features, the network model can accurately understand and process complex semantic information in dynamic environments, improving the performance and robustness of the model in diverse scenarios.

[0137] The technical effect of this solution in this embodiment is that by using the semantic transient map to construct network constraint conditions and combining the constraint conditions with the original image features for training in a preset network model architecture, the understanding and processing ability of the network model for dynamic change scenes is enhanced. This method effectively integrates semantic information changes caused by occlusion or dynamic changes, provides accurate feature compression and decompression capabilities, thereby improving the adaptability and accuracy of the positioning model in complex and dynamic environments.

[0138] In a possible design, S201. Obtain the original image and multiple rendered images, including:

[0139] S2011. Obtain the original image and the camera pose corresponding to the original image.

[0140] Specifically, image data can be captured by using a sensor device, and sensor fusion technology or computer vision algorithms, such as visual inertial odometry, can be used to estimate the pose information of the camera, including the camera position and the camera orientation. This information is usually represented in the form of a pose matrix or a quaternion. Accurately obtained camera pose information can provide the necessary geometric reference for subsequent image rendering and feature extraction, enabling the rendered image to truly reflect the perspective and structure of the original scene, thereby improving the training accuracy and effect of the positioning model.

[0141] S2012. Perform image rendering according to the camera pose and the preset scene radiation field to obtain multiple rendered images.

[0142] Specifically, by using computer graphics technology, the camera pose can be applied to the scene radiation field to simulate the effects of observing the scene from different perspectives. The scene radiation field usually contains information about the lighting, materials, and geometric structures in the scene. By combining the position information of the camera with the radiation field, multiple rendered images are generated, which simulate the appearance of the scene under different conditions. The generated rendered images can provide diverse perspectives and appearance information, helping to train the model to better understand and adapt to changes in the scene, and improving the robustness and accuracy of the model in practical applications.

[0143] For example, the formula for calculating multiple rendered images can be:

[0144]

[0145] Where, is the rendered image, is the rendering algorithm, is the original image, is the camera pose of the original image, is the radiation field.

[0146] The technical effect of this solution in this embodiment is: by obtaining the original image and its corresponding camera pose, and using this information to combine with the preset scene radiation field for image rendering, multiple rendered images are generated. This process provides rich perspective and appearance change data for model training, enabling the trained localization model to comprehensively understand and adapt to different scene conditions. The localization model shows higher robustness and accuracy in practical applications and can effectively handle diverse and dynamically changing environments.

[0147] In a possible design, in S202, feature extraction is performed on the original image to obtain multiple original image features, including:

[0148] S2021: Segment the original image to obtain multiple segmented regions.

[0149] Specifically, it can be achieved by using image segmentation algorithms, such as deep learning-based semantic segmentation networks or traditional segmentation methods, threshold segmentation, region growing, or graph cut algorithms. These algorithms divide the original image into multiple regions with similar features or semantics. By segmenting the image, different parts of the image can be analyzed and processed more precisely, enabling the subsequent feature extraction steps to accurately capture the feature information of each region, thereby improving the localization model's understanding and recognition ability of complex scenes. Among them, the multiple segmented regions refer to multiple pixel blocks segmented from the original image.

[0150] S2022: Crop the multiple segmented regions to obtain multiple cropped regions.

[0151] Specifically, by setting specific cropping rules or criteria, the segmented regions can be further processed to extract more representative pixel blocks. This may involve removing edge redundancy, focusing on regions of interest, or adjusting the region size to fit subsequent processing steps. Through the cropping operation, unnecessary background information and noise can be removed, enhancing the representativeness and effectiveness of image features, thereby improving the accuracy and efficiency of the feature extraction model and ultimately enhancing the performance of the localization model in complex scenarios. Among them, multiple cropping regions refer to multiple pixel blocks cropped from multiple segmented regions.

[0152] S2023. Input the multiple cropping regions into a preset feature extraction model to obtain multiple original image features.

[0153] Specifically, it can be achieved by using a trained deep learning model, such as a convolutional neural network. These cropping regions are passed as input data to the feature extraction model, which automatically learns and extracts important features in the image, such as edges, textures, and shapes, through its hierarchical structure. The feature extraction model can extract high-dimensional and meaningful feature representations from the cropping regions, which can more accurately describe the image content, thereby providing rich information for subsequent model training and improving the recognition accuracy and robustness of the localization model.

[0154] The technical effect of this solution in this embodiment is that by segmenting and cropping the original image, more representative pixel regions are extracted, and these cropping regions are input into a preset feature extraction model to obtain multiple original image features. This process can effectively remove redundant information and noise, focus on the most informative parts of the image, and make the extracted features more accurate and efficient. The localization model can better understand and recognize key elements in the image when dealing with complex scenarios, improving the overall performance and accuracy of the model.

[0155] Figure 3 It is a schematic flowchart of the target region localization method provided by the embodiment of the present application. As Figure 3 shown, in this embodiment, the execution subject of the embodiment of the present invention is a server. Then, the target region localization method provided by this embodiment includes:

[0156] S301. Obtain the camera pose and the query text.

[0157] Specifically, the camera pose can be obtained by using sensor devices that can provide the position information and orientation information of the camera; the query text can be input by the user or extracted from a preset database. By accurately obtaining the camera pose, the system can determine the specific position and viewing angle of the camera in three-dimensional space, while the query text provides the target information that the user is interested in. This combination enables the positioning model to effectively identify and locate the target area specified by the user in complex scenarios, improving the interactivity and practicality of the system.

[0158] S302. Input the camera pose into the positioning model to obtain multiple compressed features.

[0159] Specifically, the obtained camera pose data can be used as input parameters and passed to a pre-trained positioning model. The model uses its internal neural network structure and learned weights to fuse the camera pose with the image features and generate multiple compressed features. These compressed features are compact representations of high-dimensional information and can effectively describe the spatial characteristics and semantic information of the scene. By generating compressed features, the positioning model can reduce the computational complexity while retaining the key spatial and semantic information, thereby improving the efficiency and accuracy of the model in real-time applications. Among them, the positioning model is trained by the Figure 2 method.

[0160] S303. Calculate the correlation between the query text and the multiple compressed features to obtain multiple correlation scores.

[0161] Specifically, the query text can be converted into a text feature vector through a natural language processing model, and then the similarity calculation is performed with the compressed features of the image in a common feature space. Common methods include cosine similarity or dot product operation. By calculating the correlation scores, the system can quantify the semantic matching degree between the query text and the image features, thereby identifying the image region most relevant to the user's query. This method improves the performance of the model in multimodal data fusion, enabling the system to more accurately locate the target area of interest to the user.

[0162] For example, the formula for calculating the correlation score can be:

[0163]

[0164] where is the correlation score, is the query text, is the compressed feature.

[0165] S304. Generate a correlation score map based on the multiple correlation scores.

[0166] Specifically, a correlation score map can be generated by mapping the calculated correlation scores to the spatial layout of the image. The correlation scores can be regarded as weights, which are applied to each feature position of the image to generate a two-dimensional correlation score map, where the value of each pixel or region represents the strength of its correlation with the query text. The correlation score map intuitively shows which regions in the image are most relevant to the query text, enabling the system to quickly identify and locate the regions of interest to the user. This visualization method enhances the interpretability of the model and the user's interaction experience, and helps to accurately locate the target region in complex scenarios.

[0167] For example, a weight map can be calculated first, and then a correlation score map can be generated based on the weight map and multiple correlation scores. The formula for calculating the weight map can be:

[0168]

[0169] where, is the weight map, is the correlation score, is the mean value of the correlation scores.

[0170] The formula for calculating the correlation score map can be:

[0171]

[0172] where, is the correlation score map, is the weight map, is the correlation score.

[0173] S305. Extract the regions in the correlation score map where the correlation scores are greater than a preset threshold as the target regions.

[0174] Specifically, by setting a threshold, the correlation scores of each pixel or region in the correlation score map can be compared with the threshold. For those regions with scores higher than the threshold, the system marks them as target regions. This process can be achieved through image processing techniques such as threshold segmentation or connected region analysis. By filtering out the regions with low correlation, the system can accurately identify and locate the parts of the image that best match the query text. This method effectively filters out irrelevant information, improves the accuracy and reliability of target region localization, and meets the user's query requirements. Among them, the target region is the region queried by the query text.

[0175] In this embodiment, the technical effect of the solution is as follows: By combining the camera pose and the query text provided by the user, and using the trained positioning model to generate multiple compressed features and calculate the correlation between these features and the query text, the system can generate a correlation score map and extract the target area highly relevant to the query text. This process realizes accurate positioning from text to image, enabling users to quickly find the areas of interest in complex 3D scenes, improving the accuracy and efficiency of the system in multi-modal information processing, and enhancing the user interaction experience and the practicality of the application.

[0176] Figure 4 FIG. is a schematic structural diagram of a training device for a positioning model provided by an embodiment of the present application. As Figure 4 shown, the training device for the positioning model includes:

[0177] An image acquisition module 401, configured to acquire an original image and multiple rendered images; wherein, the multiple rendered images are obtained by rendering the original image.

[0178] An original image feature extraction module 402, configured to extract features from the original image to obtain multiple original image features.

[0179] A rendered image feature extraction module 403, configured to extract features from the multiple rendered images to obtain multiple rendered image features.

[0180] A semantic appearance map construction module 404, configured to construct a semantic appearance map according to the multiple rendered image features; wherein, the semantic appearance map is used to represent the variability of semantic information caused by appearance differences in the original image.

[0181] A semantic transient map construction module 405, configured to construct a semantic transient map according to the multiple original image features and the multiple rendered image features; wherein, the semantic transient map is used to represent the variability of semantic information caused by occluders in the original image.

[0182] A spatial model training module 406, configured to perform model training according to the multiple original image features, the multiple rendered image features, the semantic appearance map, the semantic transient map, and a preset spatial model architecture to obtain a spatial model; wherein, the spatial model is used to represent the 3D scene where the original image is located.

[0183] A network model training module 407, configured to perform model training according to the multiple original image features, the semantic transient map, and a preset network model architecture to obtain a network model; wherein, the network model is used to compress and decompress the multiple original image features and the multiple rendered image features.

[0184] A positioning model generation module 408, configured to generate a positioning model according to the spatial model and the network model.

[0185] In a possible design, the spatial model training module 406 includes:

[0186] A spatial constraint condition construction unit, configured to construct spatial constraint conditions according to the semantic appearance map and the semantic transient map.

[0187] A spatial model training unit, configured to perform model training according to the spatial constraint conditions, multiple original image features, multiple rendered image features, and the spatial model architecture to obtain a spatial model.

[0188] In a possible design, the network model training module 407 includes:

[0189] A network constraint condition construction unit, configured to construct network constraint conditions according to the semantic transient map.

[0190] A network model training unit, configured to perform model training according to the network constraint conditions, multiple original image features, and the network model architecture to obtain a network model.

[0191] In a possible design, the image acquisition module 401 includes:

[0192] A camera pose acquisition unit, configured to acquire the original image and the camera pose corresponding to the original image.

[0193] An image rendering unit, configured to perform image rendering according to the camera pose and the preset scene radiation field to obtain multiple rendered images.

[0194] In a possible design, the original image feature extraction module 402 includes:

[0195] A segmentation unit, configured to segment the original image to obtain multiple segmentation regions; wherein, the multiple segmentation regions refer to multiple pixel blocks segmented from the original image.

[0196] A cropping unit, configured to crop the multiple segmentation regions to obtain multiple cropped regions; wherein, the multiple cropped regions refer to multiple pixel blocks cropped from the multiple segmentation regions.

[0197] A feature extraction unit, configured to input the multiple cropped regions into a preset feature extraction model to obtain multiple original image features.

[0198] The training device of the positioning model provided in this embodiment can execute Figure 2 the technical solutions of the method embodiment shown in Figure 2 The implementation principle and technical effects are similar to those of the method embodiment shown in

[0199] Figure 5 It is a schematic structural diagram of the target area positioning device provided in the embodiment of the present application. AsFigure 5 As shown in the figure, the target area positioning device includes:

[0200] A pose text acquisition module 501, configured to acquire the camera pose and the query text.

[0201] A feature calculation module 502, configured to input the camera pose into the positioning model to obtain a plurality of compressed features; wherein, the positioning model is trained by a Figure 4 device.

[0202] A correlation calculation module 503, configured to calculate the correlation between the query text and the plurality of compressed features to obtain a plurality of correlation scores.

[0203] A score map generation module 504, configured to generate a correlation score map according to the plurality of correlation scores.

[0204] A target area determination module 505, configured to extract the area in the correlation score map where the correlation score is greater than a preset threshold as the target area; wherein, the target area is the area queried by the query text.

[0205] The target area positioning device provided in this embodiment can execute Figure 3 the technical solutions of the method embodiment shown in the figure, and its implementation principle and technical effects are similar to those of Figure 3 the method embodiment shown in the figure, and will not be elaborated here one by one.

[0206] Figure 6 This is a schematic structural diagram of an electronic device provided in an embodiment of the present application. As Figure 6 shown in the figure, the electronic device includes: at least one processor 610 and a memory 620. The electronic device further includes a communication component 630. Among them, the processor 610, the memory 620, and the communication component 630 are connected through a bus 640.

[0207] In a specific implementation process, at least one processor 610 executes the computer execution instructions stored in the memory 620, so that at least one processor 610 is used to implement the training method of the positioning model and the target area positioning method in the above embodiments.

[0208] For the specific implementation process of the processor 610, reference can be made to the above method embodiment, and its implementation principle and technical effects are similar, and will not be elaborated here in this embodiment.

[0209] In the above embodiments, it should be understood that the processor 610 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly implemented by the execution of the hardware processor, or can be implemented by the combination of hardware and software modules in the processor.

[0210] The memory 620 may include a high-speed RAM memory and may also include non-volatile storage NVM, such as at least one disk memory.

[0211] The bus 640 may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus 640 may be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, the bus 640 in the drawings of the present application is not limited to only one bus or one type of bus.

[0212] The functions implemented for the electronic device and the main control device are described above for the solution provided by the embodiments of the present invention. It can be understood that in order for the electronic device or the main control device to implement the above functions, it includes the corresponding hardware structures and / or software modules for executing each function. Combining the units and algorithm steps of each example described in the embodiments disclosed in the embodiments of the present invention, the embodiments of the present invention can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the technical solution of the embodiments of the present invention.

[0213] The embodiments of the present application further provide a computer-readable storage medium. Computer-executable instructions are stored in the computer-readable storage medium. When the computer-executable instructions are executed by a processor, they are used to implement the training method of the positioning model and the target area positioning method in the above embodiments. Among them, in the specific implementation of the foregoing training method of the positioning model and the target area positioning method, each module may be implemented as a processor.

[0214] The above-readable storage medium may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk. The readable storage medium may be any available medium accessible by a general-purpose or special-purpose computer.

[0215] An exemplary readable storage medium is coupled to the processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium may also be a component of the processor. The processor and the readable storage medium may be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium may also exist as discrete components in an electronic device or a master control device.

[0216] So far, the technical solutions of the present application have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present application is obviously not limited to these specific embodiments. The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for training a positioning model, characterized in that: include: Acquire an original image and a plurality of rendered images; wherein the plurality of rendered images are obtained by rendering the original image; Extracting features from the original image to obtain multiple original image features; Performing feature extraction on the multiple rendered images to obtain multiple rendered image features; Constructing a semantic appearance graph according to the plurality of rendered image features; wherein the semantic appearance graph is used to describe the variability of semantic information caused by appearance differences in the original image; Constructing a semantic transient graph according to the plurality of original image features and the plurality of rendered image features; wherein the semantic transient graph is used to describe the variability of semantic information caused by the occluder in the original image; Performing model training according to the multiple original image features, the multiple rendered image features, the semantic appearance graph, the semantic transient graph and a preset spatial model architecture to obtain a spatial model; wherein the spatial model is used to describe the three-dimensional scene in which the original image is located; Performing model training according to the multiple original image features, the semantic transient graph and a preset network model architecture to obtain a network model; wherein the network model is used to compress and decompress the multiple original image features and the multiple rendered image features; A positioning model is generated according to the space model and the network model.

2. The method for training a positioning model according to claim 1, characterized in that: The performing model training according to the plurality of original image features, the plurality of rendered image features, the semantic appearance graph, the semantic transient graph and a preset spatial model architecture to obtain a spatial model comprises: Constructing spatial constraints according to the semantic appearance graph and the semantic transient graph; Model training is performed according to the spatial constraint conditions, the multiple original image features, the multiple rendered image features and the spatial model architecture to obtain the spatial model.

3. The training method of the positioning model according to claim 1, characterized in that: The performing model training according to the plurality of original image features, the semantic transient graph and a preset network model architecture to obtain a network model comprises: constructing network constraints according to the semantic transient graph; Model training is performed according to the network constraints, the multiple original image features and the network model architecture to obtain the network model.

4. The method for training a positioning model according to claim 1, characterized in that: The obtaining of the original image and the plurality of rendered images comprises: Acquire the original image and the camera posture corresponding to the original image; Image rendering is performed according to the camera posture and a preset scene radiation field to obtain the multiple rendered images.

5. The method for training a positioning model according to claim 1, characterized in that: The extracting features of the original image to obtain multiple original image features includes: Segmenting the original image to obtain a plurality of segmented regions; wherein the plurality of segmented regions refer to a plurality of pixel blocks segmented from the original image; Cropping the multiple segmented areas to obtain multiple cropped areas; wherein the multiple cropped areas refer to multiple pixel blocks cropped from the multiple segmented areas; The multiple cropped areas are input into a preset feature extraction model to obtain the multiple original image features.

6. A method for locating a target area, characterized in that: include: Get the camera pose and query text; Inputting the camera pose into a positioning model to obtain a plurality of compressed features; wherein the positioning model is trained by the method according to any one of claims 1 to 5; Calculating the correlation between the query text and the plurality of compression features to obtain a plurality of correlation scores; generating a correlation score graph based on the plurality of correlation scores; The area in the correlation score graph whose correlation score is greater than a preset threshold is extracted as the target area; wherein the target area is the area queried by the query text.

7. A training device for a positioning model, characterized in that: include: An image acquisition module, used to acquire an original image and a plurality of rendered images; wherein the plurality of rendered images are obtained by rendering according to the original image; An original image feature extraction module is used to extract features from the original image to obtain multiple original image features; A rendering image feature extraction module, used to extract features from the plurality of rendering images to obtain a plurality of rendering image features; A semantic appearance graph construction module, used to construct a semantic appearance graph according to the plurality of rendered image features; wherein the semantic appearance graph is used to express the variability of semantic information caused by appearance differences in the original image; A semantic transient graph construction module, used to construct a semantic transient graph according to the plurality of original image features and the plurality of rendered image features; wherein the semantic transient graph is used to express the variability of semantic information caused by occluders in the original image; A spatial model training module, used for performing model training according to the multiple original image features, the multiple rendered image features, the semantic appearance graph, the semantic transient graph and a preset spatial model architecture to obtain a spatial model; wherein the spatial model is used to describe the three-dimensional scene in which the original image is located; A network model training module, used for performing model training according to the plurality of original image features, the semantic transient graph and a preset network model architecture to obtain a network model; wherein the network model is used for compressing and decompressing the plurality of original image features and the plurality of rendered image features; The positioning model generation module is used to generate a positioning model according to the space model and the network model.

8. A target area positioning device, characterized in that: include: The pose and text acquisition module is used to obtain the camera pose and query text; A feature calculation module, used for inputting the camera posture into a positioning model to obtain a plurality of compressed features; wherein the positioning model is obtained by training the device according to claim 7; A correlation calculation module, used to calculate the correlation between the query text and the plurality of compression features to obtain a plurality of correlation scores; A score graph generating module, used for generating a correlation score graph according to the multiple correlation scores; The target region determination module is used to extract the region whose correlation score is greater than a preset threshold in the correlation score map as the target region; wherein the target region is the region queried by the query text.

9. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; When the processor executes the computer-executable instructions stored in the memory, it is used to implement the training method of the positioning model as described in any one of claims 1 to 5, or to implement the target area positioning method as described in claim 6.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the training method for the positioning model as described in any one of claims 1 to 5, or to implement the target area positioning method as described in claim 6.