Spatial reasoning device, spatial encoder training device and electronic equipment

By generating viewpoint-independent global spatial features, this study addresses the insufficient accuracy of existing visual language models in 3D spatial understanding and reasoning tasks, achieving higher accuracy and robustness in 3D spatial understanding and reasoning.

CN121999342APending Publication Date: 2026-05-08SHENZHEN SWEET POTATO ROBOT CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN SWEET POTATO ROBOT CO LTD
Filing Date
2026-01-29
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing visual language models lack accuracy and robustness in 3D spatial understanding and reasoning tasks, mainly because visual encoders extract features based on 2D images, which cannot effectively model the depth information, 3D geometric structure, and spatial topological relationships of the scene.

Method used

By using a spatial reasoning device and an encoder training device, visual features and local spatial features of the image are acquired. A pre-trained spatial encoder is used to generate global spatial features, thereby realizing an abstract representation of three-dimensional space and generating a global spatial description of the target object in three-dimensional space.

Benefits of technology

It improves the accuracy and robustness of 3D spatial understanding and reasoning, generates viewpoint-independent global spatial features, and can more accurately describe the geometric structure, positional relationships and semantic content of target objects in 3D space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121999342A_ABST
    Figure CN121999342A_ABST
Patent Text Reader

Abstract

The invention discloses a spatial reasoning device, a spatial encoder training device and electronic equipment. The spatial reasoning device comprises one or more processors, and the one or more processors are configured to obtain a task instruction and a to-be-processed image corresponding to a first view angle; wherein the task instruction is used for acquiring global space description of a target object in the to-be-processed image; processing the to-be-processed image to obtain a visual feature and a local spatial feature corresponding to the first visual angle; processing the local spatial features through a pre-trained spatial encoder to obtain global spatial features of the to-be-processed image in a three-dimensional space; and processing the visual features, the global spatial features and the task instruction to obtain the global spatial description of the target object in the three-dimensional space. According to the invention, the accuracy and robustness of three-dimensional space understanding and reasoning can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a spatial reasoning device, a spatial encoder training device, and an electronic device. Background Technology

[0002] In recent years, visual language models (VLMs) have made significant progress in two-dimensional visual reasoning tasks such as image description, visual question answering, and cross-modal retrieval.

[0003] First, the visual language model extracts visual features from the input image through a visual encoder and semantic features from the input text through a large language model. Then, through a multimodal fusion mechanism, the visual language model maps visual and semantic features to a unified latent space, achieving cross-modal feature alignment and interaction. Afterward, the visual language model undergoes end-to-end optimization through fine-tuning based on pre-training objectives or instructions from downstream tasks, thereby performing two-dimensional visual reasoning tasks such as image description, visual question answering, and cross-modal retrieval.

[0004] However, existing visual language models primarily rely on two-dimensional images for feature extraction, focusing on visual representations at the appearance and semantic levels. Their ability to model spatial features such as scene depth, 3D geometry, and spatial topological relationships remains limited. Therefore, in tasks involving 3D spatial understanding and reasoning, the accuracy and robustness of spatial reasoning in existing visual language models are still insufficient. Summary of the Invention

[0005] To address the aforementioned technical problems, this disclosure provides a spatial reasoning device, a spatial encoder training device, and an electronic device to improve the accuracy and robustness of three-dimensional spatial understanding and reasoning.

[0006] A first aspect of this disclosure provides a spatial reasoning apparatus, including one or more processors, said one or more processors being configured to: Obtain the task instruction and the image to be processed corresponding to the first-person perspective; wherein, the task instruction is used to obtain the global spatial description of the target object in the image to be processed; The image to be processed is processed to obtain the visual features and local spatial features corresponding to the first viewpoint; The local spatial features are processed by a pre-trained spatial encoder to obtain the global spatial features of the image to be processed in three-dimensional space. The visual features, the global spatial features, and the task instructions are processed to obtain a global spatial description of the target object in the three-dimensional space.

[0007] A second aspect of this disclosure provides a spatial encoder training apparatus, including one or more processors, said one or more processors being configured to: Obtain the first image corresponding to the reference viewpoint and the second image corresponding to the target viewpoint; The first image and the second image are processed to obtain the first spatial feature corresponding to the reference viewpoint and the second spatial feature corresponding to the target viewpoint; The spatial encoder to be trained performs feature fusion processing on the preset spatial query features and the first spatial features to obtain the global spatial features of the first image in three-dimensional space. Camera pose estimation is performed on the second spatial features to obtain the camera pose corresponding to the target viewpoint; The global spatial features and the camera pose are processed to obtain the third spatial features corresponding to the target viewpoint; Based on the second spatial feature and the third spatial feature, the model parameters of the spatial encoder to be trained are optimized to obtain the trained spatial encoder.

[0008] A third aspect of this disclosure provides a spatial reasoning method, comprising: Obtain the task instruction and the image to be processed corresponding to the first-person perspective; wherein, the task instruction is used to obtain the global spatial description of the target object in the image to be processed; The image to be processed is processed to obtain the visual features and local spatial features corresponding to the first viewpoint; The local spatial features are processed by a pre-trained spatial encoder to obtain the global spatial features of the image to be processed in three-dimensional space. The visual features, the global spatial features, and the task instructions are processed to obtain a global spatial description of the target object in the three-dimensional space.

[0009] A fourth aspect of this disclosure provides a spatial encoder training method, comprising: Obtain the first image corresponding to the reference viewpoint and the second image corresponding to the target viewpoint; The first image and the second image are processed to obtain the first spatial feature corresponding to the reference viewpoint and the second spatial feature corresponding to the target viewpoint; The spatial encoder to be trained performs feature fusion processing on the preset spatial query features and the first spatial features to obtain the global spatial features of the first image in three-dimensional space. Camera pose estimation is performed on the second spatial features to obtain the camera pose corresponding to the target viewpoint; The global spatial features and the camera pose are processed to obtain the third spatial features corresponding to the target viewpoint; Based on the second spatial feature and the third spatial feature, the model parameters of the spatial encoder to be trained are optimized to obtain the trained spatial encoder.

[0010] A fifth aspect of this disclosure provides an electronic device comprising: a device as described in the first or second aspect embodiment; or, The electronic device includes a processor and a memory for storing executable instructions of the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the method provided in the third or fourth aspect embodiment described above.

[0011] A sixth embodiment of this disclosure provides a computer-readable storage medium storing a computer program that is executed by a processor to perform the methods provided in the third or fourth aspect embodiments described above.

[0012] A seventh aspect of this disclosure provides a computer program product that, when instructions in the computer program product are executed by a processor, performs the method provided in the third or fourth aspect of the present disclosure.

[0013] This disclosure provides a spatial reasoning device, a spatial encoder training device, and an electronic device. The spatial reasoning device includes one or more processors configured to: acquire a task instruction and a to-be-processed image corresponding to a first viewpoint; wherein the task instruction is used to acquire a global spatial description of a target object in the to-be-processed image; process the to-be-processed image to obtain visual features and local spatial features corresponding to the first viewpoint; process the local spatial features using a pre-trained spatial encoder to obtain global spatial features of the to-be-processed image in three-dimensional space; and process the visual features, global spatial features, and task instruction to obtain a global spatial description of the target object in the three-dimensional space. Thus, the processor can generate global spatial features of the to-be-processed image in three-dimensional space based on the local spatial features corresponding to the first viewpoint using a pre-trained spatial encoder. These global spatial features are viewpoint-independent spatial features, not dependent on any viewpoint, but rather an abstract representation of the inherent geometric structure, positional relationships, and semantic content of the scene corresponding to the to-be-processed image (including the target object, other objects, and the surrounding environment) in three-dimensional space. Subsequently, the processor can generate a global spatial description of the target object in the three-dimensional space based on visual features, global spatial features, and task instructions, thereby improving the accuracy and robustness of three-dimensional spatial understanding and reasoning. Attached Figure Description

[0014] Figure 1 This is a schematic diagram of the structure of a spatial reasoning device provided in an exemplary embodiment of this disclosure.

[0015] Figure 2 This is a schematic diagram of a spatial reasoning model provided by an exemplary embodiment of this disclosure.

[0016] Figure 3 This is a schematic diagram of the structure of a spatial encoder training device provided in an exemplary embodiment of this disclosure.

[0017] Figure 4 This is a schematic diagram of a spatial encoder training model provided in an exemplary embodiment of this disclosure.

[0018] Figure 5 This is a flowchart illustrating a spatial reasoning method provided in an exemplary embodiment of this disclosure.

[0019] Figure 6 This is a flowchart illustrating a spatial reasoning method provided in another exemplary embodiment of this disclosure.

[0020] Figure 7 This is a schematic flowchart of a spatial encoder training method provided in an exemplary embodiment of this disclosure.

[0021] Figure 8 This is a schematic flowchart of a spatial encoder training method provided in another exemplary embodiment of this disclosure.

[0022] Figure 9 This is a schematic flowchart of a spatial encoder training method provided in another exemplary embodiment of this disclosure.

[0023] Figure 10 This is a schematic flowchart of a spatial encoder training method provided in another exemplary embodiment of this disclosure.

[0024] Figure 11 This is a structural diagram of an electronic device provided in an exemplary embodiment of this disclosure. Detailed Implementation

[0025] To explain this disclosure, exemplary embodiments of the disclosure will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the disclosure, and not all of them. It should be understood that the disclosure is not limited to exemplary embodiments.

[0026] It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of this disclosure.

[0027] Application Overview In recent years, visual language models have made significant progress in 2D visual reasoning tasks such as image captioning, visual question answering, and cross-modal retrieval. First, visual language models extract visual features from input images through visual encoders and semantic features from input text through large language models. Then, through multimodal fusion mechanisms, visual and semantic features are mapped to a unified latent space, achieving alignment and interaction of cross-modal features. Afterward, visual language models undergo end-to-end optimization through fine-tuning based on pre-training objectives or instructions from downstream tasks, thereby performing 2D visual reasoning tasks such as image captioning, visual question answering, and cross-modal retrieval. However, existing visual language models primarily extract features from 2D images using visual encoders, focusing on visual representations at the appearance and semantic levels, with limited ability to model spatial features such as scene depth information, 3D geometric structure, and spatial topological relationships. Therefore, in tasks involving 3D spatial understanding and reasoning, the accuracy and robustness of spatial reasoning in existing visual language models remain insufficient.

[0028] To compensate for the limitations of visual language models in 3D spatial reasoning, current methods primarily employ geometric prior augmentation (such as Spatial-MLLM) to extract spatial features of 3D space from 2D images acquired from a specific viewpoint, and then input these features as prompts into the visual language model. However, the spatial features extracted using these geometric prior augmentation methods are limited to the specific viewpoint corresponding to the input 2D image, exhibiting viewpoint dependence and failing to form a global understanding of the entire 3D space. This limits visual language models to spatial reasoning based solely on spatial features from a specific viewpoint, enabling them to obtain a spatial description of the target object within that specific viewpoint.

[0029] In this embodiment, the spatial reasoning device includes one or more processors configured to: acquire a task instruction and a to-be-processed image corresponding to a first viewpoint; wherein the task instruction is used to acquire a spatial description of a target object in the to-be-processed image; process the to-be-processed image to obtain visual features and local spatial features corresponding to the first viewpoint; process the local spatial features using a pre-trained spatial encoder to obtain global spatial features of the to-be-processed image in three-dimensional space; and process the visual features, global spatial features, and task instruction to obtain a global spatial description of the target object in the three-dimensional space. The processor can generate global spatial features of the to-be-processed image in three-dimensional space based on the local spatial features corresponding to the first viewpoint using a pre-trained spatial encoder. These global spatial features are viewpoint-independent spatial features, not dependent on any viewpoint, but rather an abstract representation of the inherent geometric structure, positional relationships, and semantic content of the scene corresponding to the to-be-processed image (including the target object, other objects, and the surrounding environment) in three-dimensional space. Subsequently, the processor can generate a global spatial description of the target object in the three-dimensional space based on the visual features, global spatial features, and task instruction, thereby improving the accuracy and robustness of three-dimensional spatial understanding and reasoning.

[0030] This disclosure can be applied to the fields of intelligent driving, embodied intelligence, and robotics. In these fields, it can be applied to environmental perception, environmental reconstruction, and path planning scenarios. The spatial reasoning device described can be a smart terminal, an in-vehicle terminal, an intelligent driving domain controller, a central computing platform, etc. The processor described can be a central processing unit (CPU), a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), an image signal processor (ISP), etc.

[0031] Exemplary System Figure 1 This is a schematic diagram of the structure of a spatial reasoning device provided in an exemplary embodiment of this disclosure. Figure 1 As shown, the spatial reasoning device 100 includes one or more processors 110. The one or more processors 110 are configured to: S1: Obtain the image to be processed corresponding to the task command and the first-person perspective. The task command is used to obtain the global spatial description of the target object in the image to be processed.

[0032] For example, the processor 110 can acquire task commands input by the user through an input device (such as a keyboard, touch screen, smart terminal, microphone, smart speaker, etc.) and the image to be processed corresponding to the first view captured by the image sensor in the first view.

[0033] The task instruction is used to obtain a global spatial description of the target object in the image to be processed. This task instruction can be text-based, voice-based, or other types of task instructions; this disclosure does not limit the specific types. The corresponding task instructions differ depending on the application scenario. For example, in an environment perception scenario, the task instruction can be an instruction to instruct the acquisition of spatial descriptions such as the geometric structure, spatial dimensions, and spatial distance of the target object in the three-dimensional space reflected by the image to be processed. In an environment reconstruction scenario, the task instruction can be an instruction to instruct the acquisition of spatial descriptions such as the geometric structure and positional topology of the target object in the three-dimensional space reflected by the image to be processed. In a path planning scenario, the task instruction can be an instruction to instruct the acquisition of spatial descriptions such as the target category, spatial location, and motion state of the target object in the three-dimensional space reflected by the image to be processed.

[0034] The first viewpoint can be a single viewpoint or multiple viewpoints, and this embodiment of the disclosure does not impose any limitation. Correspondingly, the image to be processed corresponding to the first viewpoint can be a single image corresponding to one viewpoint or multiple images corresponding to multiple viewpoints, and this embodiment of the disclosure does not impose any limitation. For example, when the first viewpoint comprises multiple viewpoints, it is assumed that the first viewpoint includes viewpoint A, viewpoint B, and viewpoint C. Accordingly, the image to be processed includes the image corresponding to viewpoint A, the image corresponding to viewpoint B, and the image corresponding to viewpoint C.

[0035] S2 processes the image to be processed to obtain the visual features and local spatial features corresponding to the first-person perspective.

[0036] In this disclosure, visual features are used to describe various visually related attribute features of image content. For example, visual features may include color, texture, shape, spatial relationships, semantic features, etc. Visual features can be represented in the form of high-dimensional vectors, with each attribute feature serving as information in different dimensions of the vector. Local spatial features are used to describe various spatially related attribute features of image content under the image acquisition viewpoint. For example, local spatial features may include geometric features, semantic features, spatial relationship features, etc., corresponding to objects or backgrounds in the image under the corresponding viewpoint.

[0037] For example, after the processor 110 acquires the image to be processed corresponding to the first viewpoint, if there are multiple first viewpoints, the processor 110 can perform feature extraction processing on the image to be processed corresponding to each first viewpoint to obtain the visual features corresponding to that first viewpoint. Finally, the processor 110 obtains visual features corresponding to multiple first viewpoints.

[0038] In one implementation, the processor 110 can process the image to be processed corresponding to the first viewpoint using a pre-trained visual encoder (such as a ViT model, a ResNet model, etc.) to obtain the visual features corresponding to the first viewpoint. Specifically, the processor 110 can use the visual encoder to divide the image to be processed into multiple image patches according to a preset image patch size, obtaining the image patch sequence corresponding to the image to be processed. For example, taking viewpoint A in the first viewpoint as an example, the image patch sequence Patch_a corresponding to the image to be processed is {patch_a_1, patch_a_2…patch_a_i…patch_a_n}. Here, patch_a_i represents the i-th image patch of the image to be processed corresponding to viewpoint A, and n represents the number of image patches of the image to be processed corresponding to viewpoint A. Then, the processor 110 can use the visual encoder to perform linear projection or convolution operations on the image patch sequence corresponding to the image to be processed, mapping the image patch sequence to the latent space to obtain the initial visual feature sequence corresponding to the image to be processed. Subsequently, the processor 110 can use a visual encoder to add the initial visual feature sequence, the position encoding sequence, and the view encoding corresponding to the first viewpoint to obtain the target visual feature sequence (i.e., the visual features corresponding to the first viewpoint) based on the first viewpoint. Here, the position encoding sequence corresponds one-to-one with the image patches in the image patch sequence. For example, taking viewpoint A in the first viewpoint as an example, the visual feature Token_a corresponding to viewpoint A is {token_a_1, token_a_2, ..., token_a_i, ..., token_a_n}. Here, token_a_i represents the i-th visual feature in the visual features corresponding to viewpoint A, and n represents the number of visual features in the visual features corresponding to viewpoint A.

[0039] When there are multiple first-viewpoints, after obtaining the visual features corresponding to each first-viewpoint, the processor 110 can perform feature fusion processing on the visual features corresponding to that first-viewpoint and the visual features corresponding to itself and other first-viewpoints to obtain the local visual features corresponding to that first-viewpoint. The local spatial features corresponding to the first-viewpoint can be considered as features obtained by fusing the geometric and semantic features of all other first-viewpoints while retaining the first-viewpoint's own geometric and semantic features, using that first-viewpoint as the anchor point. The local spatial features corresponding to the first-viewpoint can include the local geometric features and local semantic features of that first-viewpoint, as well as the local geometric features and local semantic features of other first-viewpoints that are spatially related to that first-viewpoint.

[0040] In one implementation, the processor 110 can perform feature fusion processing on the visual features corresponding to each first viewpoint using a pre-trained self-attention aggregator (such as a self-attention model) to obtain the local spatial features corresponding to each first viewpoint. Specifically, for each first viewpoint, the processor 110 can use the self-attention aggregator to use the visual features corresponding to that first viewpoint as query features, and concatenate the visual features corresponding to each first viewpoint into a single feature, serving as both a key feature and a value feature. Then, the processor 110 can use the self-attention aggregator to perform attention fusion on the query feature, key feature, and value feature to obtain the local spatial features corresponding to that first viewpoint. For example, when there are multiple first viewpoints, assuming the first viewpoints include viewpoint A, viewpoint B, and viewpoint C, the visual features corresponding to each first viewpoint include Token_a, Token_b, and Token_c. For viewpoint A, Token_a is used as the query feature Q_a. Token_a, Token_b, and Token_c are concatenated into a single feature, which serves as the key feature K_abc and the value feature V_abc. Attention fusion is then applied to the query feature Q_a, the key feature K_abc, and the value feature V_abc to obtain the local spatial feature Token_a_Fusion corresponding to viewpoint A. Similarly, for viewpoint B, Token_b is used as the query feature Q_b. Token_a, Token_b, and Token_c are concatenated into a single feature, which serves as the key feature K_abc and the value feature V_abc. Attention fusion is then applied to the query feature Q_b, the key feature K_abc, and the value feature V_abc to obtain the local spatial feature Token_b_Fusion corresponding to viewpoint B. Similarly, for viewpoint C, Token_c is used as the query feature Q_c, Token_a, Token_b and Token_c are concatenated into a single feature, which is then used as the key feature K_abc and the value feature V_abc. Attention is then applied to fuse the query feature Q_c, the key feature K_abc and the value feature V_abc to obtain the local spatial feature Token_c_Fusion corresponding to viewpoint C.

[0041] S3 processes local spatial features through a pre-trained spatial encoder to obtain the global spatial features of the image in three-dimensional space.

[0042] In this disclosure, the global spatial feature can characterize viewpoint-independent spatial features. This global spatial feature does not depend on any viewpoint, but rather is an abstract representation of the inherent geometric structure, positional relationships, and semantic content of the scene corresponding to the image to be processed (including target objects, other objects, and the surrounding environment) in three-dimensional space. The global spatial feature can include the global geometric features, global semantic features, and global topological relationship features of each target object, other object, and surrounding environment in the scene corresponding to the image to be processed in three-dimensional space.

[0043] For example, the processor 110 can initialize a spatial query feature. This spatial query feature is a viewpoint-independent feature, and its dimension can be the same as or different from the dimension of the local spatial features; this embodiment does not impose limitations. The dimension of the spatial query feature refers to the number of dimensions of the spatial query feature vector, such as 512 dimensions or 1024 dimensions. When there are multiple first viewpoints, after obtaining the local spatial features corresponding to each first viewpoint, the processor 110 can further perform feature fusion processing on the preset spatial query feature and the local spatial features corresponding to each first viewpoint using a pre-trained spatial encoder to obtain the global spatial features of the image to be processed in three-dimensional space. In other words, the processor 110 can extract the global spatial features of the three-dimensional space corresponding to the image to be processed from the local spatial features corresponding to each first viewpoint using the spatial query feature.

[0044] In one implementation, the pre-trained spatial encoder can employ a Transformer-based encoder. This spatial encoder may include a multi-head attention layer, a linear projection layer, a normalization layer, and a feedforward neural network. The multi-head attention layer performs feature fusion processing on the spatial query features and the local spatial features corresponding to each first viewpoint, outputting initial global spatial features. The linear projection layer compresses the dimensions of the initial global spatial features, outputting dimension-compressed global spatial features. The normalization layer normalizes the dimension-compressed global spatial features, outputting normalized global spatial features. The feedforward neural network enhances the normalized global spatial features, outputting the final global spatial features.

[0045] S4 processes visual features, global spatial features, and task instructions to obtain a global spatial description of the target object in the three-dimensional space.

[0046] For example, after receiving a task instruction, the processor 110 can perform feature extraction processing on the task instruction to obtain the semantic features corresponding to the task instruction. The semantic features corresponding to the task instruction are high-dimensional vectorized representations of the core semantic information of the task instruction. These semantic features may include the query intent information, query requirement information, and semantic logic information of the task instruction. In one implementation, the processor 110 can use a pre-trained text encoder (such as a BERT model, a MiniLM model, etc.) to perform feature extraction processing on the task instruction to obtain the semantic features corresponding to the task instruction.

[0047] When there are multiple first-viewpoints, after obtaining the visual features corresponding to each first-viewpoint and the global spatial features corresponding to the three-dimensional space, the processor 110 can perform feature fusion processing on the visual features and global spatial features corresponding to each first-viewpoint to obtain the fused features corresponding to that first-viewpoint. Ultimately, the processor 110 obtains fused features corresponding to multiple first-viewpoints. In one implementation, the processor 110 can perform feature fusion processing on the visual features and global spatial features corresponding to the first-viewpoints through a cross-attention aggregator (such as a cross-attention model) to obtain the fused features corresponding to the first-viewpoints.

[0048] When there are multiple first-viewpoints, after obtaining the fused features corresponding to each first-viewpoint and the semantic features corresponding to the task instructions, the processor 110 can further perform spatial reasoning processing on the fused features based on the semantic features to obtain a global spatial description of the target object in three-dimensional space corresponding to the semantic features. In one embodiment, the processor 110 can use a pre-trained large language model (such as GPT-4V, Gemini Pro, Qwen-VL, etc.) to perform spatial reasoning processing on the fused features corresponding to each first-viewpoint based on the semantic features corresponding to the task instructions to obtain a global spatial description of the target object in three-dimensional space corresponding to the semantic features.

[0049] In this embodiment, the processor can generate global spatial features of the image to be processed in three-dimensional space based on local spatial features corresponding to a first viewpoint, using a pre-trained spatial encoder. These global spatial features are viewpoint-independent; they do not depend on any viewpoint but are abstract representations of the inherent geometric structure, positional relationships, and semantic content of the scene corresponding to the image to be processed (including the target object, other objects, and the surrounding environment) in three-dimensional space. Subsequently, the processor can generate a global spatial description of the target object in the three-dimensional space based on visual features, global spatial features, and task instructions, thereby improving the accuracy and robustness of three-dimensional spatial understanding and reasoning.

[0050] In one embodiment, the processor 110 processes local spatial features using a pre-trained spatial encoder to obtain global spatial features of the image to be processed in three-dimensional space, specifically configured as follows: By using a pre-trained spatial encoder, query features are constructed using preset spatial query features, and key features are constructed using local spatial features. Feature fusion processing is then performed on the spatial query features and local spatial features to obtain the global spatial features of the image to be processed in three-dimensional space.

[0051] For example, when there are multiple first perspectives, after the processor 110 obtains the local spatial features corresponding to each first perspective, it can further construct query features with spatial query features and construct key-value features with the local spatial features corresponding to each first perspective through a pre-trained spatial encoder. The processor 110 then performs feature fusion processing on the spatial query features and the local spatial features corresponding to each first perspective to obtain the global spatial features of the image to be processed in three-dimensional space.

[0052] In one implementation, the processor 110 can use the spatial query feature as the query feature, concatenate all local spatial features corresponding to the first viewpoint into a single feature, and use it as both the key and value features. Then, the processor 110 can perform attention fusion on the query feature, key feature, and value feature to obtain the global spatial features of the image to be processed in three-dimensional space. This global spatial feature can represent viewpoint-independent spatial features; that is, it does not depend on any viewpoint but is an abstract representation of the inherent geometric structure, positional relationships, and semantic content of the scene corresponding to the image to be processed (including the target object, other objects, and the surrounding environment) in three-dimensional space. For example, if there are multiple first viewpoints, assuming the first viewpoints include viewpoint A, viewpoint B, and viewpoint C, the local spatial features corresponding to each first viewpoint include Token_a_Fusion, Token_b_Fusion, and Token_c_Fusion, and the spatial query feature is Token_s. Using Token_s as the query feature Q_s, concatenating Token_a_Fusion, Token_b_Fusion, and Token_c_Fusion into a single feature, and using this as the key feature K_abc_Fusion and the value feature V_abc_Fusion, then performing attention fusion on the query feature Q_s, the key feature K_abc_Fusion, and the value feature V_abc_Fusion to obtain the global spatial feature Token_g of the image to be processed in 3D space.

[0053] In this embodiment of the present disclosure, the processor 110 extracts global spatial features of the three-dimensional space that are independent of the viewpoint from the local spatial features corresponding to the first viewpoint through a spatial encoder based on a spatial query feature that is independent of the viewpoint, thereby achieving decoupling of spatial features from viewpoint.

[0054] In one embodiment, the processor 110 processes visual features, global spatial features, and task instructions to obtain a global spatial description of the target object in three-dimensional space, specifically configured as follows: S41 performs feature fusion processing on visual features and global spatial features to obtain fused features.

[0055] For example, when there are multiple first-viewpoints, after the processor 110 obtains the visual features corresponding to each first-viewpoint and the global spatial features corresponding to the three-dimensional space, for each first-viewpoint, the processor 110 can perform feature fusion processing on the visual features and global spatial features corresponding to that first-viewpoint to obtain the fused features corresponding to that first-viewpoint. Ultimately, the processor 110 obtains multiple fused features corresponding to the first-viewpoints. Each fused feature integrates the visual features corresponding to the first-viewpoint and the global spatial features independent of the viewpoint. That is, the fused features reflect the inherent geometric structure, positional relationships, and semantic content of the target object in three-dimensional space. In one implementation, the processor 110 can perform feature fusion processing on the visual features and global spatial features corresponding to the first-viewpoints through a cross-attention adapter (such as a cross-attention model) to obtain the fused features corresponding to the first-viewpoints.

[0056] S42, perform feature extraction processing on the task instructions to obtain the semantic features corresponding to the task instructions.

[0057] For example, after receiving a task instruction, the processor 110 can perform feature extraction processing on the task instruction to obtain the semantic features corresponding to the task instruction. The semantic features corresponding to the task instruction are high-dimensional vectorized representations of the core semantic information of the task instruction. These semantic features may include the query intent information, query requirement information, and semantic logic information of the task instruction.

[0058] In one implementation, for task instructions that are text-based, the processor 110 can use a pre-trained text encoder (such as a BERT model or a MiniLM model) to perform feature encoding on the task instructions and obtain the semantic features corresponding to the task instructions. For task instructions that are not text-based, the processor 110 can first convert the task instructions into text-based instructions, and then use a pre-trained text encoder to perform feature encoding on the task instructions and obtain the semantic features corresponding to the task instructions, thereby enabling the processing of non-text task instructions.

[0059] S43. Based on semantic features, spatial reasoning processing is performed on the fused features to obtain the global spatial description of the target object in three-dimensional space corresponding to the semantic features.

[0060] For example, when there are multiple first perspectives, after the processor 110 obtains the fused features corresponding to each first perspective and the semantic features corresponding to the task instructions, it can further perform spatial reasoning processing on each fused feature based on the semantic features to obtain the global spatial description of the target object in three-dimensional space corresponding to the semantic features.

[0061] In one implementation, the processor 110 can use a pre-trained large language model (such as GPT-4V, Gemini Pro, Qwen-VL, etc.) to perform spatial reasoning on the fused features corresponding to each first vision based on the semantic features corresponding to the task instructions, thereby obtaining a global spatial description of the target object in three-dimensional space corresponding to the semantic features. Specifically, the processor 110 uses the large language model to first map the semantic features and the fused features corresponding to each first vision to the same latent space using a pre-trained cross-modal alignment matrix, achieving an initial binding of spatial information with semantic requirements. Then, the processor 110 uses the large language model, with semantic features as query features and the fused features corresponding to each first vision as key-value features, calculates relevance weights through a multi-head attention mechanism, performs normalization processing, and extracts spatial features strongly correlated with the task instructions, completing a precise binding of semantics and space. Afterwards, the processor 110 performs progressive spatial reasoning through the decoder in the large language model, extracting quantized spatial parameters (such as three-dimensional dimensions and coordinate positions) from the fused features, integrating spatial information from the global perspective, constructing the three-dimensional topological relationship between the target object and its surrounding environment, and obtaining a structured spatial reasoning result. Finally, the processor 110 uses a large language model and an autoregressive generation method to transform the structured spatial reasoning results into text that conforms to the expression habits of natural language. At the same time, it filters non-spatial information, standardizes the expression format, and finally outputs a global spatial description of the target object in three-dimensional space.

[0062] In this embodiment, the processor 110 can perform feature fusion processing on visual features and global spatial features to obtain fused features, and perform feature extraction processing on task instructions to obtain semantic features corresponding to the task instructions. Subsequently, the processor 110 can perform spatial reasoning processing on the fused features based on the semantic features to obtain a global spatial description of the target object in three-dimensional space corresponding to the semantic features, thereby improving the accuracy and robustness of three-dimensional spatial understanding and reasoning.

[0063] In one embodiment, the processor 110 performs feature fusion processing on visual features and global spatial features to obtain fused features, specifically configured as follows: Query features are constructed using visual features, key-value features are constructed using global spatial features, and feature fusion processing is performed on the visual features and global spatial features to obtain fused features.

[0064] For example, when there are multiple first perspectives, after the processor 110 obtains the visual features corresponding to each first perspective and the global spatial features corresponding to the three-dimensional space, for each first perspective, the processor 110 can use the visual features corresponding to that first perspective as query features and the global spatial features as key features and value features. Then, the processor 110 can perform attention fusion on the query features, key features, and value features to obtain the fused features corresponding to that first perspective. Finally, the processor 110 obtains fused features corresponding to multiple first perspectives. For example, when there are multiple first perspectives, assuming that the first perspectives include perspective A, perspective B, and perspective C, and the visual features corresponding to each first perspective include Token_a, Token_b, and Token_c, and the global spatial feature is Token_g. For perspective A, Token_a is used as query feature Q_a, and Token_g is used as key feature K_g and value feature V_g. Attention fusion is performed on query feature Q_a, key feature K_g, and value feature V_g to obtain the fused feature Token_ag_Fusion corresponding to perspective A. For viewpoint B, Token_b is used as the query feature Q_b, and Token_g is used as the key feature K_g and value feature V_g. Attention is then applied to fuse the query feature Q_b, key feature K_g, and value feature V_g to obtain the fused feature Token_bg_Fusion for viewpoint B. For viewpoint C, Token_c is used as the query feature Q_c, and Token_g is used as the key feature K_g and value feature V_g. Attention is then applied to fuse the query feature Q_c, key feature K_g, and value feature V_g to obtain the fused feature Token_cg_Fusion for viewpoint C.

[0065] In this embodiment of the disclosure, the processor 110 achieves the ability to understand the visual appearance of a target object and to perform global spatial reasoning by fusing visual features and global spatial features.

[0066] Figure 2 This is a schematic diagram of a spatial reasoning model provided by an exemplary embodiment of this disclosure. For example... Figure 2As shown, when there are multiple first-viewpoints, in step one, the processor 110 acquires task instructions and the images to be processed corresponding to each first-viewpoint. The task instructions are used to acquire the global spatial description of the target object in the image to be processed. In step two, for each first-viewpoint, the processor 110 uses a pre-trained visual encoder to perform feature extraction processing on the image to be processed corresponding to that first-viewpoint, obtaining the visual features corresponding to that first-viewpoint. Finally, the processor 110 obtains the visual features corresponding to each first-viewpoint. In step three, the processor 110 uses a pre-trained self-attention aggregator to construct query features and key-value features based on the visual features corresponding to each first-viewpoint, and performs feature fusion processing on the visual features corresponding to each first-viewpoint to obtain the local spatial features corresponding to each first-viewpoint. In step four, the processor 110 uses a pre-trained spatial encoder to construct query features based on preset spatial query features, and constructs key-value features based on the local spatial features corresponding to each first-viewpoint, and performs feature fusion processing on the spatial query features and the local spatial features corresponding to each first-viewpoint to obtain the global spatial features of the image to be processed in three-dimensional space. Step 5: For each first-viewpoint, processor 110 constructs query features using the visual features corresponding to that first-viewpoint and key-value features using the global spatial features through a pre-trained cross-attention aggregator. It then performs feature fusion processing on the visual features and global spatial features corresponding to that first-viewpoint to obtain the fused features. Finally, processor 110 obtains the fused features corresponding to each first-viewpoint. Step 6: Processor 110 performs feature extraction processing on the task instructions using a pre-trained text encoder to obtain the semantic features corresponding to the task instructions. Step 7: Processor 110 uses a pre-trained large language model to perform spatial reasoning processing on each fused feature based on the semantic features to obtain the global spatial description of the target object in three-dimensional space corresponding to the semantic features.

[0067] Figure 3 This is a schematic diagram of the structure of a spatial encoder training apparatus provided in an exemplary embodiment of this disclosure. Figure 3 As shown, the spatial encoder training device 300 includes one or more processors 310. The one or more processors 310 are configured to: S31, acquire the first image corresponding to the reference viewpoint and the second image corresponding to the target viewpoint.

[0068] For example, the processor 310 can acquire a first image captured by an image sensor corresponding to a reference viewpoint and a second image captured by an image sensor corresponding to a target viewpoint. The reference viewpoint refers to a known viewpoint used to generate global spatial features of the three-dimensional space. The target viewpoint refers to an unknown viewpoint used to supervise and verify the accuracy of the generated global spatial features of the three-dimensional space. The target viewpoint and the reference viewpoint are different viewpoints. The reference viewpoint and the target viewpoint can be one viewpoint or multiple viewpoints; this embodiment of the disclosure does not limit this. Correspondingly, the first image corresponding to the reference viewpoint can be a single image corresponding to one viewpoint or multiple images corresponding to multiple viewpoints; this embodiment of the disclosure does not limit this. The second image corresponding to the target viewpoint can be a single image corresponding to one viewpoint or multiple images corresponding to multiple viewpoints; this embodiment of the disclosure does not limit this. For example, when there are multiple reference viewpoints and one target viewpoint, assuming the reference viewpoints may include viewpoints A, B, and C, and the target viewpoint may include viewpoint D. Accordingly, the first image may include the image corresponding to viewpoint A, the image corresponding to viewpoint B, and the image corresponding to viewpoint C, and the second image may include the image corresponding to viewpoint D.

[0069] S32, process the first image and the second image to obtain the first spatial features corresponding to the reference viewpoint and the second spatial features corresponding to the target viewpoint.

[0070] In this disclosure, visual features are used to describe various visually related attribute features of image content. For example, visual features may include features such as color, texture, shape, spatial relationships, and semantics. Visual features can be represented in the form of high-dimensional vectors, with each attribute feature serving as information in different dimensions of the vector. First spatial features and second spatial features are used to describe various spatially related attribute features of image content under the image acquisition viewpoint. For example, first spatial features and second spatial features may include geometric features, semantic features, spatial relationship features, etc., corresponding to objects or backgrounds in the image under the corresponding viewpoint.

[0071] For example, when there are multiple reference viewpoints and a single target viewpoint, after acquiring the first image corresponding to each reference viewpoint, the processor 310 can perform feature extraction processing on the first image corresponding to each reference viewpoint to obtain the first visual features corresponding to that reference viewpoint. Ultimately, the processor 310 obtains the first visual features corresponding to each reference viewpoint. Similarly, after acquiring the second image corresponding to the target viewpoint, the processor 310 can perform feature extraction processing on the second image corresponding to the target viewpoint to obtain the second visual features corresponding to that target viewpoint.

[0072] In one implementation, the processor 310 can perform feature extraction processing on the first image corresponding to the reference viewpoint using a pre-trained asymmetric viewpoint aggregator (such as a VGGT model based on asymmetric masks) to obtain the first visual features corresponding to the reference viewpoint, and perform feature extraction processing on the second image corresponding to the target viewpoint to obtain the second visual features corresponding to the target viewpoint.

[0073] After obtaining the first visual features corresponding to the reference viewpoint and the second visual features corresponding to the target viewpoint, the processor 310 needs to further perform feature fusion processing on the first and second visual features to obtain the first spatial features corresponding to the reference viewpoint and the second spatial features corresponding to the target viewpoint. The first spatial features corresponding to the reference viewpoint are the input spatial features used by the spatial encoder to be trained to learn viewpoint-independent global spatial features in 3D space. The second spatial features corresponding to the target viewpoint are the spatial features used for reverse supervision of the spatial encoder to be trained.

[0074] The processor 310 supervises the training of the spatial encoder from the target viewpoint to ensure it can accurately learn viewpoint-independent global spatial features in 3D space under known viewpoints. The processor 310 can use asymmetric masks to perform attention interactions between the second visual features corresponding to the target viewpoint and the first visual features corresponding to the reference viewpoint. In other words, for the target viewpoint, the second visual feature corresponding to that viewpoint will interact with itself and the first visual features corresponding to the reference viewpoint to obtain the second spatial feature corresponding to the target viewpoint. This second spatial feature can be considered as a fusion of the geometric and semantic features of all reference viewpoints, while retaining its own geometric and semantic features, using the target viewpoint as the anchor viewpoint. The second spatial feature corresponding to the target viewpoint can include the local geometric and semantic features of the target viewpoint, as well as the local geometric and semantic features of reference viewpoints spatially related to the target viewpoint.

[0075] In one implementation, the processor 310 can use a pre-trained asymmetric view aggregator (such as a VGGT model based on asymmetric masks) to perform feature fusion processing on the first visual features corresponding to the reference view and the second visual features corresponding to the target view based on the asymmetric mask, thereby obtaining the first spatial features corresponding to the reference view and the second spatial features corresponding to the target view.

[0076] S33, through the spatial encoder to be trained, the preset spatial query features and the first spatial features are fused to obtain the global spatial features of the first image in three-dimensional space.

[0077] For example, processor 310 can initialize a spatial query feature. This spatial query feature is a viewpoint-independent, learnable query feature. The dimension of this spatial query feature can be the same as or different from the dimension of the first spatial feature; this embodiment of the present disclosure does not impose limitations. The dimension of the spatial query feature refers to the number of dimensions of the spatial query feature vector, such as 512 dimensions or 1024 dimensions. Processor 310 can initialize the initial value of the spatial query feature through random initialization. During the training process of the spatial encoder, the feature values ​​of this spatial query feature are continuously optimized, thereby enabling it to capture global spatial features of three-dimensional space.

[0078] When there are multiple reference viewpoints, after obtaining the first spatial features corresponding to each reference viewpoint, the processor 310 can further perform feature fusion processing on the preset spatial query features and the first spatial features corresponding to each reference viewpoint through a spatial encoder to be trained, to obtain the global spatial features of the first image in three-dimensional space. In other words, the processor 310 can extract the global spatial features of the three-dimensional space corresponding to the first image from the first spatial features corresponding to each reference viewpoint through the spatial query features. These global spatial features can represent spatial features independent of the viewpoint. They do not depend on any viewpoint but are an abstract representation of the inherent geometric structure, positional relationships, and semantic content of the scene corresponding to the first image (including the target object, other objects, and the surrounding environment) in three-dimensional space. These global spatial features can include global geometric features, global semantic features, and global topological relationship features in three-dimensional space.

[0079] In one implementation, the spatial encoder may employ a Transformer-based encoder. This spatial encoder may include a multi-head attention layer, a linear projection layer, a normalization layer, and a feedforward neural network. The multi-head attention layer performs feature fusion processing on the spatial query features and the first spatial features corresponding to each reference viewpoint, outputting initial global spatial features. The linear projection layer compresses the dimensions of the initial global spatial features, outputting dimension-compressed global spatial features. The normalization layer normalizes the dimension-compressed global spatial features, outputting normalized global spatial features. The feedforward neural network enhances the normalized global spatial features, outputting the final global spatial features.

[0080] S34, perform camera pose estimation on the second spatial features to obtain the camera pose corresponding to the target viewpoint.

[0081] For example, for the same 3D space, images acquired by image sensors from different viewpoints are different. Correspondingly, the spatial features extracted based on the images from those viewpoints are also different. The camera pose corresponding to the image sensor can reflect the viewpoint of the image sensor. In order to extract the spatial features corresponding to the target viewpoint (i.e., the subsequent third spatial features) from the viewpoint-independent global spatial features, after obtaining the second spatial features corresponding to the target viewpoint, the processor 310 can further perform camera pose estimation on the second spatial features to obtain the camera pose corresponding to the target viewpoint. The camera pose can include camera intrinsic parameters and camera extrinsic parameters; the camera intrinsic parameters are the same in camera poses corresponding to different target viewpoints, but the camera extrinsic parameters are different.

[0082] In one implementation, the processor 310 can perform camera pose estimation based on the second spatial features corresponding to the target viewpoint using a pre-trained camera prediction head to obtain the camera pose corresponding to the target viewpoint. The camera prediction head can be an MLP (Multilayer Perceptron) or other types of neural networks; this disclosure does not limit the specific implementation.

[0083] S35 processes global spatial features and camera pose to obtain third spatial features corresponding to the target's viewpoint.

[0084] For example, after obtaining the camera pose corresponding to the target viewpoint, the processor 310 can generate ray query features corresponding to the target viewpoint based on the intrinsic and extrinsic parameters in the camera pose, and then obtain the third spatial features corresponding to the target viewpoint based on the ray query features. In one embodiment, the processor 310 can generate ray query features corresponding to the target viewpoint based on the camera pose using a pre-trained MLP. The third spatial features corresponding to the target viewpoint are spatial features extracted from the global spatial features.

[0085] In one implementation, the processor 310 can process ray query features and global spatial features using a spatial decoder to obtain third spatial features corresponding to the target viewpoint. The spatial decoder can be a Transformer-based decoder. This spatial decoder may include a multi-head attention layer, a linear projection layer, a normalization layer, and a feedforward neural network. The multi-head attention layer performs feature fusion processing on the ray query features and global spatial features, outputting initial third spatial features corresponding to the target viewpoint. The linear projection layer compresses the dimensions of the initial third spatial features, outputting dimension-compressed third spatial features. The normalization layer normalizes the dimension-compressed third spatial features, outputting normalized third spatial features. The feedforward neural network enhances the normalized third spatial features, outputting the final third spatial features.

[0086] S36, based on the second and third spatial features, optimizes the model parameters of the spatial encoder to be trained, and obtains the trained spatial encoder.

[0087] For example, the second spatial feature corresponding to the target viewpoint is the real spatial feature in the 3D space corresponding to the target viewpoint extracted based on the second image corresponding to the target viewpoint. The third spatial feature corresponding to the target viewpoint is the predicted spatial feature in the target viewpoint corresponding to the global spatial feature obtained by the processor 310 after feature extraction of the first image corresponding to the reference viewpoint by the spatial encoder to be trained. To verify whether the spatial encoder to be trained can extract global spatial features in the 3D space based on the first image, the similarity between the third spatial feature obtained by pose estimation based on the global spatial feature and the third spatial feature is further calculated. The similarity reflects whether the spatial encoder to be trained can learn global spatial features in the 3D space that are independent of the viewpoint under a known viewpoint. Therefore, after the processor obtains the third spatial feature corresponding to the target viewpoint, it can calculate the second spatial feature and the third spatial feature corresponding to the target viewpoint to construct a loss function, and continuously iterate and optimize the spatial encoder to be trained based on the loss function. For example, the cosine similarity between the second spatial feature and the third spatial feature can be calculated, and the model parameters of the spatial encoder to be trained can be iteratively optimized with the goal of minimizing the cosine similarity, until the cosine similarity no longer changes or the maximum number of iterations is reached, thus obtaining the trained spatial encoder.

[0088] In this embodiment, the processor first acquires a first image corresponding to a reference viewpoint and a second image corresponding to a target viewpoint. Then, it processes the first and second images to obtain first spatial features corresponding to the reference viewpoint and second spatial features corresponding to the target viewpoint. Using a spatial encoder to be trained, it performs feature fusion processing on preset spatial query features and the first spatial features to obtain global spatial features of the first image in three-dimensional space. It then performs camera pose estimation on the second spatial features to obtain the camera pose corresponding to the target viewpoint. Finally, it processes the global spatial features and the camera pose to obtain third spatial features corresponding to the target viewpoint. Finally, based on the second and third spatial features, it optimizes the model parameters of the spatial encoder to be trained to obtain the trained spatial encoder. Thus, in this disclosure, the second spatial features corresponding to the target viewpoint are used as supervision, and the third spatial features of the target viewpoint obtained based on the global spatial features output by the spatial encoder to be trained are used as prediction spatial features. The spatial encoder to be trained is iteratively trained so that the trained spatial encoder can predict the global spatial features of a target object in three-dimensional space based on an image corresponding to any viewpoint. Therefore, the spatial encoder obtained by the above training method can predict global spatial features based on single-view images, which improves the accuracy and robustness of the spatial encoder in spatial understanding and spatial reasoning.

[0089] In one embodiment, the processor 310 processes the first image and the second image to obtain a first spatial feature corresponding to the reference viewpoint and a second spatial feature corresponding to the target viewpoint, specifically configured as follows: S321, Perform feature extraction processing on the first image to obtain the first visual features corresponding to the reference viewpoint.

[0090] For example, when there are multiple reference viewpoints, after the processor 310 acquires the first image corresponding to each reference viewpoint, it can perform feature extraction processing on the first image corresponding to each reference viewpoint to obtain the first visual features corresponding to that reference viewpoint. Finally, the processor 310 obtains the first visual features corresponding to each reference viewpoint. In one implementation, the processor 310 can use a pre-trained asymmetric viewpoint aggregator (such as a VGGT model based on asymmetric masks) to perform feature extraction processing on the first image corresponding to the reference viewpoint to obtain the first visual features corresponding to the reference viewpoint. Specifically, the processor 310 can use the asymmetric viewpoint aggregator to divide the first image into multiple image patches according to a preset image patch size to obtain the image patch sequence corresponding to the first image. For example, taking viewpoint A in the reference viewpoints as an example, the image patch sequence Patch_a corresponding to the first image is {patch_a_1, patch_a_2…patch_a_i…patch_a_n}. Here, patch_a_i represents the i-th image patch of the first image corresponding to viewpoint A, and n represents the number of image patches of the first image corresponding to viewpoint A. Then, the processor 310 can perform linear projection or convolution operations on the image patch sequence corresponding to the first image using an asymmetric viewpoint aggregator to map the image patch sequence corresponding to the first image to the latent space, thereby obtaining the initial visual feature sequence corresponding to the first image. Afterwards, the processor 310 can use the asymmetric viewpoint aggregator to add the initial visual feature sequence with the position encoding sequence and the viewpoint encoding corresponding to the reference viewpoint, thereby obtaining the target visual feature sequence (i.e., the first visual feature) corresponding to the reference viewpoint. Here, the position encoding in the position encoding sequence corresponds one-to-one with the image patch in the image patch sequence. For example, taking viewpoint A in the reference viewpoint as an example, the first visual feature Token_a corresponding to viewpoint A is {token_a_1, token_a_2…token_a_i…token_a_n}. Here, token_a_i represents the i-th visual feature in the first visual feature corresponding to viewpoint A, and n represents the number of visual features in the first visual feature corresponding to viewpoint A.

[0091] S322, Perform feature extraction processing on the second image to obtain the second visual features corresponding to the target viewpoint.

[0092] For example, after acquiring the second image corresponding to the target viewpoint, the processor 310 can perform feature extraction processing on the second image corresponding to the target viewpoint to obtain the second visual features corresponding to the target viewpoint. The process by which the processor 310 performs feature extraction processing on the second image to obtain the second visual features corresponding to the target viewpoint is similar to the process by which the processor 310 performs feature extraction processing on the first image to obtain the first visual features corresponding to the reference viewpoint, and will not be described in detail here.

[0093] S323, based on a preset asymmetric mask, perform feature fusion processing on the first visual feature and the second visual feature to obtain the first spatial feature corresponding to the reference viewpoint and the second spatial feature corresponding to the target viewpoint.

[0094] For example, during the training of the spatial encoder, the spatial encoder to be trained learns viewpoint-independent global spatial features in three-dimensional space from a reference viewpoint. The processor 310 extracts the predicted spatial features (i.e., the subsequent third spatial features) corresponding to the target viewpoint from the viewpoint-independent global spatial features and compares them with the real spatial features (i.e., the subsequent second spatial features) corresponding to the target viewpoint, thereby providing reverse supervision on whether the spatial encoder to be trained can accurately learn the viewpoint-independent global spatial features in three-dimensional space under a known viewpoint.

[0095] After obtaining the first visual features corresponding to the reference viewpoint and the second visual features corresponding to the target viewpoint, the processor 310 needs to further perform feature fusion processing on the first and second visual features to obtain the first spatial features corresponding to the reference viewpoint and the second spatial features corresponding to the target viewpoint. The first spatial features corresponding to the reference viewpoint are the input spatial features used by the spatial encoder to be trained to learn viewpoint-independent global spatial features in 3D space. The second spatial features corresponding to the target viewpoint are the spatial features used for reverse supervision of the spatial encoder to be trained.

[0096] If the first spatial feature corresponding to the reference viewpoint is fused with the spatial feature corresponding to the target viewpoint, the spatial encoder to be trained will learn viewpoint-independent global spatial features in 3D space under both the reference and target viewpoints. Therefore, the processor 310 can use an asymmetric mask to limit the attention interaction between the first visual feature corresponding to the reference viewpoint and the second visual feature corresponding to the target viewpoint. That is, the processor 310 performs feature fusion processing on the first visual feature corresponding to the reference viewpoint based on a preset asymmetric mask to obtain the first spatial feature corresponding to the reference viewpoint. In other words, for each reference viewpoint, the first visual feature corresponding to that reference viewpoint only interacts with itself and the first visual features corresponding to other reference viewpoints, and will not interact with the second visual feature corresponding to the target viewpoint, thereby ensuring that the spatial encoder to be trained only learns viewpoint-independent global spatial features in 3D space from the reference viewpoint. Here, the asymmetric mask (also called an asymmetric attention mask) is used to limit the range of attention interaction during feature fusion. In one embodiment, when there are multiple reference viewpoints, for each reference viewpoint, the processor 310 uses the first visual feature corresponding to that reference viewpoint as a query feature, concatenates the first visual features corresponding to each reference viewpoint into a single feature, and uses it as both a key feature and a value feature. Then, the processor 310 can perform attention fusion on the query features, key features, and value features to obtain the first spatial features corresponding to the reference viewpoint. The first spatial features corresponding to the reference viewpoint can be considered as fusing the geometric and semantic features of all other reference viewpoints while retaining the reference viewpoint's own geometric and semantic features. The first spatial features corresponding to the reference viewpoint may include the local geometric features and local semantic features of the reference viewpoint, as well as the local geometric features and local semantic features of other reference viewpoints spatially related to the reference viewpoint. Finally, the processor 310 obtains the first spatial features corresponding to each reference viewpoint.

[0097] To avoid the second spatial features corresponding to the target viewpoint not being integrated with the spatial features corresponding to the reference viewpoint, which would cause the spatial encoder to learn only the spatial features related to the target viewpoint in 3D space under known viewpoints, and not the global spatial features independent of the viewpoint in 3D space, the processor 310 can use an asymmetric mask to perform attention interaction between the second visual features corresponding to the target viewpoint and the first visual features corresponding to the reference viewpoint. That is, the processor 310 performs feature fusion processing on the second visual features corresponding to the target viewpoint and the first visual features corresponding to the reference viewpoint based on a preset asymmetric mask to obtain the second spatial features corresponding to the target viewpoint. In other words, for the target viewpoint, the second visual features corresponding to the target viewpoint will perform attention interaction with itself and the first visual features corresponding to the reference viewpoints. In one embodiment, the processor 310 can use the second visual features corresponding to the target viewpoint as a query feature, concatenate the second visual features and the first visual features corresponding to each reference viewpoint into a single feature, and use it as the key feature and value feature. Then, the processor 310 can perform attention fusion on the query feature, key feature, and value feature to obtain the second spatial features corresponding to the target viewpoint.

[0098] In one implementation, the processor 310 can use a pre-trained asymmetric view aggregator (such as a VGGT model based on asymmetric masks) to perform feature fusion processing on the first visual features corresponding to the reference view and the second visual features corresponding to the target view based on the asymmetric mask, thereby obtaining the first spatial features corresponding to the reference view and the second spatial features corresponding to the target view.

[0099] In this embodiment of the disclosure, on the one hand, the processor 310 restricts the attention interaction between the first visual feature corresponding to the reference viewpoint and the second visual feature corresponding to the target viewpoint, thereby ensuring that the spatial encoder to be trained learns only viewpoint-independent global spatial features in three-dimensional space from the reference viewpoint. On the other hand, the processor 310 avoids the situation where the second spatial feature corresponding to the target viewpoint does not incorporate the spatial features corresponding to the reference viewpoint, causing the spatial encoder to be trained to learn only target viewpoint-related spatial features in three-dimensional space from the known viewpoint, by performing attention interaction between the second visual feature corresponding to the target viewpoint and the first visual feature corresponding to the reference viewpoint.

[0100] In one embodiment, the asymmetric mask includes a first preserved mask corresponding to the first visual feature and a second preserved mask corresponding to the second visual feature. The processor 310 performs feature fusion processing on the first visual feature and the second visual feature based on the preset asymmetric mask to obtain a first spatial feature corresponding to the reference viewpoint and a second spatial feature corresponding to the target viewpoint, specifically configured as follows: S3231, based on the first preserved mask, construct query features and key-value features with the first visual features, perform feature fusion processing on the first visual features, and obtain the first spatial features corresponding to the reference viewpoint.

[0101] For example, an asymmetric mask may include a reserved mask (represented by 0) and an occlusion mask (represented by 1). The reserved mask indicates that attentional interaction is possible, while the occlusion mask indicates that attentional interaction is not required. When the processor 310 restricts the attentional interaction between the first visual feature corresponding to the reference viewpoint and the second visual feature corresponding to the target viewpoint using an asymmetric mask, the asymmetric mask may include a first reserved mask corresponding to the first visual feature. Correspondingly, the mask corresponding to the second visual feature is an occlusion mask. For example, if the feature sequence consisting of the first and second visual features is {Token_a, Token_b, Token_c, Token_d}, then the asymmetric mask is {0, 0, 0, 1}. However, when the processor 310 uses an asymmetric mask to perform attentional interaction between the second visual feature corresponding to the target viewpoint and the first visual feature corresponding to the reference viewpoint, the asymmetric mask may include a first reserved mask corresponding to the first visual feature and a second reserved mask corresponding to the second visual feature. For example, if the feature sequence consisting of the first visual feature and the second visual feature is {Token_a, Token_b, Token_c, Token_d}, then the asymmetric mask is {0, 0, 0, 0}.

[0102] When the processor 310 restricts the attentional interaction between the first visual feature corresponding to the reference viewpoint and the second visual feature corresponding to the target viewpoint using an asymmetric mask, the asymmetric mask may include a first reserved mask corresponding to the first visual feature. Based on the first reserved mask, the processor 310 constructs query features and key-value features from the first visual features, performs feature fusion processing on the first visual features, and obtains the first spatial features corresponding to the reference viewpoint. That is, for each reference viewpoint, the first visual feature corresponding to that reference viewpoint only interacts with itself and the first visual features corresponding to other reference viewpoints, and does not interact with the second visual feature corresponding to the target viewpoint. Based on this, when there are multiple reference viewpoints, for each reference viewpoint, the processor 310 uses the first visual feature corresponding to that reference viewpoint as the query feature, concatenates the first visual features corresponding to all reference viewpoints into a single feature, and uses this as both the key and value features. Then, the processor 310 can perform attentional fusion on the query feature, key feature, and value feature to obtain the first spatial features corresponding to that reference viewpoint. Finally, the processor 310 obtains the first spatial features corresponding to each reference viewpoint. For example, when there are multiple reference viewpoints, assuming the reference viewpoints include viewpoint A, viewpoint B, and viewpoint C, the first visual features corresponding to each reference viewpoint include Token_a, Token_b, and Token_c. For viewpoint A, Token_a is used as the query feature Q_a, and Token_a, Token_b, and Token_c are concatenated into a single feature, which serves as the key feature K_abc and the value feature V_abc. Attention fusion is then performed on the query feature Q_a, the key feature K_abc, and the value feature V_abc to obtain the first spatial feature Token_a_Fusion corresponding to viewpoint A. Similarly, for viewpoint B, Token_b is used as the query feature Q_b, and Token_a, Token_b, and Token_c are concatenated into a single feature, which serves as the key feature K_abc and the value feature V_abc. Attention fusion is then performed on the query feature Q_b, the key feature K_abc, and the value feature V_abc to obtain the first spatial feature Token_b_Fusion corresponding to viewpoint B. Similarly, for viewpoint C, Token_c is used as the query feature Q_c, and Token_a, Token_b and Token_c are concatenated into a single feature, which is used as the key feature K_abc and the value feature V_abc. Attention fusion is then performed on the query feature Q_c, the key feature K_abc and the value feature V_abc to obtain the first spatial feature Token_c_Fusion corresponding to viewpoint C.

[0103] In one implementation, the processor 310 can use a pre-trained asymmetric view aggregator (such as a VGGT model based on asymmetric masks) to construct query features and key-value features based on a first preserved mask, and perform feature fusion processing on the first visual features to obtain the first spatial features corresponding to the reference view.

[0104] S3232, based on the first and second preserved masks, construct query features with second visual features, construct key-value features with the second and first visual features, perform feature fusion processing on the second and first visual features, and obtain the second spatial features corresponding to the target viewpoint.

[0105] For example, when the processor 310 performs attentional interaction between the second visual feature corresponding to the target viewpoint and the first visual feature corresponding to the reference viewpoint using an asymmetric mask, the asymmetric mask may include a first reserved mask corresponding to the first visual feature and a second reserved mask corresponding to the second visual feature. The processor 310 can construct query features based on the first and second reserved masks, construct key-value features using the second visual feature, and perform feature fusion processing on the second and first visual features to obtain the second spatial feature corresponding to the target viewpoint. In other words, for the second visual feature corresponding to the target viewpoint, this second visual feature will perform attentional interaction with itself and the first visual feature corresponding to the reference viewpoint.

[0106] In one implementation, when there are multiple reference views, the processor 310 can use the second visual feature corresponding to the target view as a query feature, and concatenate the second visual feature with the first visual features corresponding to each reference view into a single feature, serving as both a key feature and a value feature. Then, the processor 310 can perform attention fusion on the query feature, key feature, and value feature to obtain the second spatial feature corresponding to the target view. This second spatial feature can be considered as a fusion of the geometric and semantic features of all reference views, with the target view as the anchor view, while retaining its own geometric and semantic features. For example, when there are multiple reference views, assuming the reference views include view A, view B, and view C, the target view includes view D, the first visual features corresponding to each reference view include Token_a, Token_b, and Token_c, and the second visual feature corresponding to the target view includes Token_d. For viewpoint D, Token_d is used as query feature Q_d. Token_a, Token_b, Token_c and Token_d are concatenated into a single feature, which is used as key feature K_abcd and value feature V_abcd. The query feature Q_d, key feature K_abcd and value feature V_abcd are then fused with attention to obtain the second spatial feature Token_d_Fusion corresponding to viewpoint D.

[0107] In one implementation, the processor 310 can use a pre-trained asymmetric view aggregator (such as a VGGT model based on asymmetric masks) to construct query features based on a first and a second reserved mask, construct key-value features using second visual features, and perform feature fusion processing on the second and first visual features to obtain the second spatial features corresponding to the target viewpoint.

[0108] In this embodiment of the disclosure, on the one hand, the processor 310 restricts the attention interaction between the first visual feature corresponding to the reference viewpoint and the second visual feature corresponding to the target viewpoint, thereby ensuring that the spatial encoder to be trained learns only viewpoint-independent global spatial features in three-dimensional space from the reference viewpoint. On the other hand, the processor 310 avoids the situation where the second spatial feature corresponding to the target viewpoint does not incorporate the spatial features corresponding to the reference viewpoint, causing the spatial encoder to be trained to learn only target viewpoint-related spatial features in three-dimensional space from the known viewpoint, by performing attention interaction between the second visual feature corresponding to the target viewpoint and the first visual feature corresponding to the reference viewpoint.

[0109] In one embodiment, the processor 310 performs feature fusion processing on preset spatial query features and first spatial features using a spatial encoder to be trained, to obtain global spatial features of the first image in three-dimensional space, specifically configured as follows: The spatial encoder to be trained constructs query features with preset spatial query features and key features with first spatial features. The spatial query features and first spatial features are then fused to obtain the global spatial features of the first image in three-dimensional space.

[0110] For example, processor 310 can initialize a spatial query feature. This spatial query feature is a viewpoint-independent, learnable query feature. The dimension of this spatial query feature can be the same as or different from the dimension of the first spatial feature; this embodiment of the present disclosure does not impose limitations. The dimension of the spatial query feature refers to the number of dimensions of the spatial query feature vector, such as 512 dimensions or 1024 dimensions. Processor 310 can initialize the initial value of the spatial query feature through random initialization. During the training process of the spatial encoder, the feature values ​​of this spatial query feature are continuously optimized, thereby enabling it to capture global spatial features of three-dimensional space.

[0111] When there are multiple reference views, after obtaining the first spatial features corresponding to each reference view, the processor 310 can further use the spatial encoder to be trained to use the spatial query features as query features, and concatenate the first spatial features corresponding to each reference view into a single feature, which serves as both the key and value features. Then, the processor 310 can perform attention fusion on the query feature, key feature, and value feature to obtain the global spatial features of the first image in three-dimensional space. For example, when there are multiple reference views, assuming the reference views include view A, view B, and view C, the first spatial features corresponding to each reference view include Token_a_Fusion, Token_b_Fusion, and Token_c_Fusion, and the spatial query feature is Token_s. Using Token_s as the query feature Q_s, concatenating Token_a_Fusion, Token_b_Fusion, and Token_c_Fusion into a single feature, and using this as the key feature K_abc_Fusion and the value feature V_abc_Fusion, then performing attention fusion on the query feature Q_s, the key feature K_abc_Fusion, and the value feature V_abc_Fusion to obtain the global spatial feature Token_g of the first image in 3D space.

[0112] In this embodiment of the present disclosure, the processor 310 extracts the global spatial features corresponding to the three-dimensional space from the first spatial features corresponding to the reference viewpoint through the spatial encoder to be trained, based on a viewpoint-independent and learnable spatial query feature, thereby achieving decoupling of spatial features and viewpoint.

[0113] In one embodiment, the processor 310 processes global spatial features and camera pose to obtain third spatial features corresponding to the target viewpoint, specifically configured as follows: S351 generates ray query features corresponding to the target viewpoint based on the camera pose.

[0114] For example, based on the intrinsic and extrinsic parameters of the camera pose corresponding to the target viewpoint, each pixel in the second image corresponding to the target viewpoint can be back-projected with a ray originating from the camera center, passing through the pixel, and pointing towards three-dimensional space. Therefore, after obtaining the camera pose of the target viewpoint, the processor 310 can convert the pixel coordinates of each pixel in the second image into Plück coordinates based on the intrinsic and extrinsic parameters of the camera pose. Then, the processor 310 encodes the Plück coordinates of all pixels to obtain the ray query features corresponding to the target viewpoint. In one embodiment, the processor 310 can generate the ray query features corresponding to the target viewpoint based on the camera pose using a pre-trained MLP.

[0115] S352 constructs query features using ray query features and key-value features using global spatial features. It then performs feature fusion processing on the ray query features and global spatial features to obtain the third spatial features corresponding to the target viewpoint.

[0116] For example, after obtaining the ray query features corresponding to the target viewpoint, the processor 310 can use the ray query features as query features and the global spatial features as key and value features. Then, the processor 310 can perform attention fusion on the query features, key features, and value features to obtain the third spatial features corresponding to the target viewpoint. For example, if the target viewpoint is viewpoint D, the ray query features are Token_d_r, and the global spatial features are Token_g, Token_d_r is used as the query feature Q_d_r, and Token_g is used as the key feature K_g and value feature V_g. Attention fusion is then performed on the query feature Q_d_r, the key feature K_g, and the value feature V_g to obtain the third spatial feature Token_d_predict corresponding to the target viewpoint. In one embodiment, the processor 310 can process the ray query features and global spatial features through a spatial decoder to obtain the third spatial features corresponding to the target viewpoint. The spatial decoder can be a decoder based on the Transformer framework. This spatial decoder can include a multi-head attention layer, a linear projection layer, a normalization layer, and a feedforward neural network. The multi-head attention layer performs feature fusion processing on the ray query features and global spatial features, outputting the initial third spatial features corresponding to the target viewpoint. The linear projection layer compresses the dimensions of the initial third spatial features, outputting the dimension-compressed third spatial features. The normalization layer normalizes the dimension-compressed third spatial features, outputting the normalized third spatial features. The feedforward neural network enhances the normalized third spatial features, outputting the final third spatial features.

[0117] In this embodiment, the processor 310 uses camera intrinsic and extrinsic parameters to backproject the pixel coordinates of each pixel in the second image into Plück coordinates, and encodes the Plück coordinates to generate ray query features corresponding to the target viewpoint. The ray query features strictly adhere to the principles of physical imaging, thereby ensuring the accuracy of the third spatial features corresponding to the target viewpoint extracted from viewpoint-independent global spatial features.

[0118] Figure 4 This is a schematic diagram of a spatial encoder training model provided in an exemplary embodiment of this disclosure. Figure 4As shown, in step one, processor 310 acquires a first image corresponding to a reference viewpoint and a second image corresponding to a target viewpoint. In step two, processor 310 uses an asymmetric viewpoint aggregator to perform feature extraction processing on the first image to obtain first visual features corresponding to the reference viewpoint. In step three, processor 310 uses an asymmetric viewpoint aggregator to perform feature extraction processing on the second image to obtain second visual features corresponding to the target viewpoint. In step four, processor 310 uses an asymmetric viewpoint aggregator to construct query features and key-value features based on a preset first asymmetric mask, and performs feature fusion processing on the first visual features to obtain first spatial features corresponding to the reference viewpoint. Specifically, in the first asymmetric mask, the mask corresponding to the first visual features is a reserved mask, and the mask corresponding to the second visual features is an occlusion mask. In step five, processor 310 uses an asymmetric viewpoint aggregator to construct query features based on a preset second asymmetric mask, and constructs key-value features using the second visual features and the first visual features. It then performs feature fusion processing on the second visual features and the first visual features to obtain second spatial features corresponding to the target viewpoint. In the second asymmetric mask, the mask corresponding to the first visual feature is a preserved mask, and the mask corresponding to the second visual feature is also a preserved mask. Step six: The processor 310 constructs query features using preset spatial query features through the spatial encoder to be trained, constructs key-value features using the first spatial features corresponding to the reference viewpoint, and performs feature fusion processing on the spatial query features and the first spatial features to obtain global spatial features. Step seven: The processor 310 estimates the camera pose of the second spatial features corresponding to the target viewpoint using the camera prediction head to obtain the camera pose corresponding to the target viewpoint. Step eight: The processor 310 generates ray query features corresponding to the target viewpoint based on the camera pose. Step nine: The processor 310 constructs query features using ray query features and constructs key-value features using global spatial features through the spatial decoder to be trained, and performs feature fusion processing on the ray query features and global spatial features to obtain the third spatial features corresponding to the target viewpoint. Step ten: The processor 310 optimizes the model parameters of the spatial encoder and spatial decoder to be trained based on the second and third spatial features corresponding to the target viewpoint to obtain the trained spatial encoder and spatial decoder.

[0119] It should be noted that the asymmetric view aggregator can employ a VGGT (Visual Geometry Grounded Transformer) model or other types of neural network models; this disclosure does not limit the specific implementation. The spatial encoder can be a Transformer-based encoder or other types of neural network models; this disclosure does not limit the specific implementation. The camera prediction head can be an MLP or other types of neural network models; this disclosure does not limit the specific implementation. The spatial decoder can be a Transformer-based decoder or other types of neural network models; this disclosure does not limit the specific implementation.

[0120] Exemplary methods Figure 5 This is a schematic flowchart of a spatial reasoning method provided in an exemplary embodiment of this disclosure. This embodiment can be applied to a spatial reasoning device, such as... Figure 5 As shown, this spatial reasoning method may include the following steps: Step 501: Obtain the task instruction and the image to be processed corresponding to the first-person view. The task instruction is used to obtain the global spatial description of the target object in the image to be processed.

[0121] Step 502: Process the image to be processed to obtain the visual features and local spatial features corresponding to the first viewpoint.

[0122] Step 503: The local spatial features are processed by a pre-trained spatial encoder to obtain the global spatial features of the image to be processed in three-dimensional space.

[0123] Step 504: Process the visual features, global spatial features, and task instructions to obtain a global spatial description of the target object in three-dimensional space.

[0124] In one embodiment, in the above Figure 5 Based on the illustrated embodiment, step 503 may include the following steps: By using a pre-trained spatial encoder, query features are constructed using preset spatial query features, and key features are constructed using local spatial features. Feature fusion processing is then performed on the spatial query features and local spatial features to obtain the global spatial features of the image to be processed in three-dimensional space.

[0125] In one embodiment, such as Figure 6 As shown above, in the above Figure 5 Based on the illustrated embodiment, step 504 may include the following steps: Step 601: Perform feature fusion processing on visual features and global spatial features to obtain fused features.

[0126] Step 602: Perform feature extraction processing on the task instructions to obtain the semantic features corresponding to the task instructions.

[0127] Step 603: Based on semantic features, perform spatial reasoning processing on the fused features to obtain the global spatial description of the target object in three-dimensional space corresponding to the semantic features.

[0128] In one embodiment, in the above Figure 6 Based on the illustrated embodiment, step 601 may include the following steps: Query features are constructed using visual features, key-value features are constructed using global spatial features, and feature fusion processing is performed on the visual features and global spatial features to obtain fused features.

[0129] It should be noted that the spatial reasoning method of this disclosure embodiment can be executed by the above-mentioned spatial reasoning device. The specific implementation method and technical effects correspond to the spatial reasoning device of this disclosure embodiment. For details, please refer to the spatial reasoning device section. In order to reduce redundancy, it will not be described in detail.

[0130] Figure 7 This is a schematic flowchart of a spatial encoder training method provided in an exemplary embodiment of this disclosure. This embodiment can be applied to a spatial encoder training device, such as... Figure 7 As shown, the spatial encoder training method may include the following steps: Step 701: Obtain the first image corresponding to the reference viewpoint and the second image corresponding to the target viewpoint.

[0131] Step 702: Process the first image and the second image to obtain the first spatial features corresponding to the reference viewpoint and the second spatial features corresponding to the target viewpoint.

[0132] Step 703: Using the spatial encoder to be trained, feature fusion processing is performed on the preset spatial query features and the first spatial features to obtain the global spatial features of the first image in three-dimensional space.

[0133] Step 704: Perform camera pose estimation on the second spatial features to obtain the camera pose corresponding to the target viewpoint.

[0134] Step 705: Process the global spatial features and camera pose to obtain the third spatial features corresponding to the target viewpoint.

[0135] Step 706: Based on the second and third spatial features, optimize the model parameters of the spatial encoder to be trained to obtain the trained spatial encoder.

[0136] In one embodiment, such as Figure 8 As shown above, in the above Figure 7 Based on the illustrated embodiment, step 702 may include the following steps: Step 801: Perform feature extraction processing on the first image to obtain the first visual features corresponding to the reference viewpoint.

[0137] Step 802: Perform feature extraction processing on the second image to obtain the second visual features corresponding to the target viewpoint.

[0138] Step 803: Based on a preset asymmetric mask, perform feature fusion processing on the first visual feature and the second visual feature to obtain the first spatial feature corresponding to the reference viewpoint and the second spatial feature corresponding to the target viewpoint.

[0139] In one embodiment, the asymmetric mask includes a first reserved mask corresponding to a first visual feature and a second reserved mask corresponding to a second visual feature, such as... Figure 9 As shown above, in the above Figure 8 Based on the illustrated embodiment, step 803 may include the following steps: Step 901: Based on the first preserved mask, construct query features and key-value features with the first visual features, perform feature fusion processing on the first visual features, and obtain the first spatial features corresponding to the reference viewpoint.

[0140] Step 902: Based on the first and second preserved masks, construct query features with second visual features, construct key-value features with the second and first visual features, and perform feature fusion processing on the second and first visual features to obtain the second spatial features corresponding to the target viewpoint.

[0141] In one embodiment, in the above Figure 7 Based on the illustrated embodiment, step 703 may include the following steps: The spatial encoder to be trained constructs query features with preset spatial query features and key features with first spatial features. The spatial query features and first spatial features are then fused to obtain the global spatial features of the first image in three-dimensional space.

[0142] In one embodiment, such as Figure 10 As shown above, in the above Figure 7 Based on the illustrated embodiment, step 705 may include the following steps: Step 1001: Based on the camera pose, generate the ray query features corresponding to the target viewpoint.

[0143] Step 1002: Construct query features using ray query features and key-value features using global spatial features. Perform feature fusion processing on ray query features and global spatial features to obtain the third spatial features corresponding to the target viewpoint.

[0144] It should be noted that the spatial encoder training method of this disclosure embodiment can be executed by the above-mentioned spatial encoder training device. The specific implementation method and technical effects correspond to the spatial encoder training device of this disclosure embodiment. For details, please refer to the spatial encoder training device section. In order to reduce redundancy, it will not be described in detail.

[0145] Exemplary electronic devices Figure 11 A structural diagram of an electronic device provided in an embodiment of this disclosure includes at least one processor 11 and a memory 12.

[0146] The processor 11 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 10 to perform desired functions.

[0147] The memory 12 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 11 may execute one or more computer program instructions to implement the methods and / or other desired functions of the various embodiments of this disclosure described above.

[0148] In one example, the electronic device 10 may also include an input device 13 and an output device 14, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).

[0149] The input device 13 may also include, for example, a keyboard, a mouse, etc.

[0150] The output device 14 can output various information to the outside, including, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0151] Of course, for the sake of simplicity, Figure 11 Only some of the components of the electronic device 10 relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device 10 may include any other suitable components depending on the specific application.

[0152] Exemplary computer program products and computer-readable storage media In addition to the methods and apparatus described above, embodiments of this disclosure may also provide a computer program product, including computer program instructions that, when executed by a processor, cause the processor to perform the steps of the methods in the various embodiments of this disclosure described in the "Exemplary Methods" section above.

[0153] Computer program products can be written in any combination of one or more programming languages ​​to perform the operations of embodiments of this disclosure. These programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0154] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps of the methods in the various embodiments of this disclosure described in the "Exemplary Methods" section above.

[0155] Computer-readable storage media may take the form of any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may include, but is not limited to, systems, apparatuses, or devices that are electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0156] The basic principles of this disclosure have been described above with reference to specific embodiments. However, the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.

[0157] Various modifications and variations can be made to this disclosure without departing from its spirit and scope. Therefore, this disclosure is also intended to include such modifications and variations if they fall within the scope of the claims of this disclosure and their equivalents.

Claims

1. A spatial reasoning apparatus, comprising one or more processors, said one or more processors being configured to: Acquire the task instructions and the corresponding first-person view image to be processed; where, The task instruction is used to obtain the global spatial description of the target object in the image to be processed; The image to be processed is processed to obtain the visual features and local spatial features corresponding to the first viewpoint; The local spatial features are processed by a pre-trained spatial encoder to obtain the global spatial features of the image to be processed in three-dimensional space. The visual features, the global spatial features, and the task instructions are processed to obtain a global spatial description of the target object in the three-dimensional space.

2. The apparatus according to claim 1, wherein, The step of processing the local spatial features using a pre-trained spatial encoder to obtain the global spatial features of the image to be processed in three-dimensional space includes: The pre-trained spatial encoder constructs query features with preset spatial query features, constructs key features with local spatial features, and performs feature fusion processing on the spatial query features and the local spatial features to obtain the global spatial features of the image to be processed in three-dimensional space.

3. The apparatus according to claim 1, wherein, The process of processing the visual features, the global spatial features, and the task instructions to obtain a global spatial description of the target object in the three-dimensional space includes: The visual features and the global spatial features are fused together to obtain fused features; The task instructions are subjected to feature extraction processing to obtain the semantic features corresponding to the task instructions; Based on the semantic features, spatial reasoning processing is performed on the fused features to obtain the global spatial description of the target object corresponding to the semantic features in the three-dimensional space.

4. The apparatus according to claim 3, wherein, The step of performing feature fusion processing on the visual features and the global spatial features to obtain fused features includes: The visual features are used to construct query features, the global spatial features are used to construct key-value features, and the visual features and the global spatial features are fused to obtain the fused features.

5. A spatial encoder training apparatus, comprising one or more processors, said one or more processors being configured to: Obtain the first image corresponding to the reference viewpoint and the second image corresponding to the target viewpoint; The first image and the second image are processed to obtain the first spatial feature corresponding to the reference viewpoint and the second spatial feature corresponding to the target viewpoint; The spatial encoder to be trained performs feature fusion processing on the preset spatial query features and the first spatial features to obtain the global spatial features of the first image in three-dimensional space. Camera pose estimation is performed on the second spatial features to obtain the camera pose corresponding to the target viewpoint; The global spatial features and the camera pose are processed to obtain the third spatial features corresponding to the target viewpoint; Based on the second spatial feature and the third spatial feature, the model parameters of the spatial encoder to be trained are optimized to obtain the trained spatial encoder.

6. The apparatus according to claim 5, wherein, The step of processing the first image and the second image to obtain the first spatial features corresponding to the reference viewpoint and the second spatial features corresponding to the target viewpoint includes: The first image is subjected to feature extraction processing to obtain the first visual features corresponding to the reference viewpoint; The second image is subjected to feature extraction processing to obtain the second visual features corresponding to the target viewpoint; Based on a preset asymmetric mask, feature fusion processing is performed on the first visual feature and the second visual feature to obtain the first spatial feature corresponding to the reference viewpoint and the second spatial feature corresponding to the target viewpoint.

7. The apparatus according to claim 6, wherein, The asymmetric mask includes a first reserved mask corresponding to the first visual feature and a second reserved mask corresponding to the second visual feature; The step of performing feature fusion processing on the first visual feature and the second visual feature based on a preset asymmetric mask to obtain the first spatial feature corresponding to the reference viewpoint and the second spatial feature corresponding to the target viewpoint includes: Based on the first preserved mask, query features and key features are constructed using the first visual features, and feature fusion processing is performed on the first visual features to obtain the first spatial features corresponding to the reference viewpoint. Based on the first and second preserved masks, query features are constructed using the second visual features, key-value features are constructed using the second and first visual features, and feature fusion processing is performed on the second and first visual features to obtain the second spatial features corresponding to the target viewpoint.

8. The apparatus according to claim 5, wherein, The step of using a spatial encoder to be trained to perform feature fusion processing on preset spatial query features and the first spatial features to obtain the global spatial features of the first image in three-dimensional space includes: The spatial encoder to be trained constructs query features with the preset spatial query features, constructs key features with the first spatial features, and performs feature fusion processing on the spatial query features and the first spatial features to obtain the global spatial features of the first image in three-dimensional space.

9. The apparatus according to claim 5, wherein, The process of processing the global spatial features and the camera pose to obtain the third spatial features corresponding to the target viewpoint includes: Based on the camera pose, generate the ray query features corresponding to the target viewpoint; A query feature is constructed using the ray query feature, and a key-value feature is constructed using the global spatial feature. The ray query feature and the global spatial feature are then fused to obtain the third spatial feature corresponding to the target viewpoint.

10. A spatial reasoning method, comprising: Obtain the task instruction and the image to be processed corresponding to the first-person perspective; wherein, the task instruction is used to obtain the global spatial description of the target object in the image to be processed; The image to be processed is processed to obtain the visual features and local spatial features corresponding to the first viewpoint; The local spatial features are processed by a pre-trained spatial encoder to obtain the global spatial features of the image to be processed in three-dimensional space. The visual features, the global spatial features, and the task instructions are processed to obtain a global spatial description of the target object in the three-dimensional space.

11. The method according to claim 10, wherein, The step of processing the local spatial features using a pre-trained spatial encoder to obtain the global spatial features of the image to be processed in three-dimensional space includes: The pre-trained spatial encoder constructs query features with preset spatial query features, constructs key features with local spatial features, and performs feature fusion processing on the spatial query features and the local spatial features to obtain the global spatial features of the image to be processed in three-dimensional space.

12. The method according to claim 10, wherein, The process of processing the visual features, the global spatial features, and the task instructions to obtain a global spatial description of the target object in the three-dimensional space includes: The visual features and the global spatial features are fused together to obtain fused features; The task instructions are subjected to feature extraction processing to obtain the semantic features corresponding to the task instructions; Based on the semantic features, spatial reasoning processing is performed on the fused features to obtain the global spatial description of the target object corresponding to the semantic features in the three-dimensional space.

13. The method according to claim 12, wherein, The step of performing feature fusion processing on the visual features and the global spatial features to obtain fused features includes: The visual features are used to construct query features, the global spatial features are used to construct key-value features, and the visual features and the global spatial features are fused to obtain the fused features.

14. A spatial encoder training method, comprising: Obtain the first image corresponding to the reference viewpoint and the second image corresponding to the target viewpoint; The first image and the second image are processed to obtain the first spatial feature corresponding to the reference viewpoint and the second spatial feature corresponding to the target viewpoint; The spatial encoder to be trained performs feature fusion processing on the preset spatial query features and the first spatial features to obtain the global spatial features of the first image in three-dimensional space. Camera pose estimation is performed on the second spatial features to obtain the camera pose corresponding to the target viewpoint; The global spatial features and the camera pose are processed to obtain the third spatial features corresponding to the target viewpoint; Based on the second spatial feature and the third spatial feature, the model parameters of the spatial encoder to be trained are optimized to obtain the trained spatial encoder.

15. An electronic device, the electronic device comprising: The apparatus as described in any one of claims 1 to 4, or 5 to 9; or, The electronic device includes a processor and a memory for storing instructions executable by the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the method of any one of claims 10 to 13 or 14.

16. A computer-readable storage medium storing a computer program that is executed by a processor to perform the method of any one of claims 10 to 13 or 14.