Spatial reasoning method and device based on image sequence, equipment and medium

By obtaining image sequences and camera pose information, screening key frames and constructing multimodal prompts, and using multimodal language models for spatial reasoning, the existing technology solves the dependence on complex three-dimensional data, realizes flexible and efficient spatial reasoning in the fields of financial technology and medical health, and improves the adaptability and practicality of the system.

CN120706573APending Publication Date: 2025-09-26PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510881713.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing three-dimensional spatial reasoning technology relies on complex three-dimensional structural data and specialized spatial modeling modules, making it difficult to achieve flexible and efficient spatial reasoning in dynamic, complex and resource-constrained environments, especially in the fields of financial technology and healthcare, where it is difficult to meet the requirements of rapid deployment and real-time performance.

Method used

By obtaining image sequences and camera pose information, filtering key frames and constructing multimodal prompts, the pre-trained multimodal language model is used to generate spatial reasoning results, reducing the dependence on complex three-dimensional data modeling. The key frames are filtered by combining visual language features and spatial features, multimodal prompts are constructed and input into the multimodal language model for spatial reasoning.

Benefits of technology

It achieves efficient and flexible spatial reasoning based on ordinary image data, reduces the system's dependence on hardware equipment and computing resources, improves the system's flexibility and practicality, and adapts to the rapid deployment needs of changing environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706573A_ABST
    Figure CN120706573A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a spatial reasoning method, device, equipment and medium based on an image sequence, and the method comprises the steps: obtaining the image sequence and camera pose information corresponding to each frame of image in the image sequence, and obtaining a spatial reasoning result based on the image sequence and the camera pose information; and screening key frames and camera pose information corresponding to the key frames in combination with the visual language features and the spatial features, constructing a multi-modal prompt based on the key frames, the camera pose information and the spatial reasoning request, and generating a spatial reasoning result by using a pre-trained multi-modal language model. According to the method, the key frames with richer information amount and wider space coverage are screened by combining the visual language features and the spatial features, and the multi-modal prompt is constructed based on the key frames and the camera pose information, so that the multi-modal language model can efficiently understand the spatial relationship in the scene, and the scene quality is improved. And flexible and low-cost spatial reasoning can be realized without depending on complex three-dimensional data modeling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a spatial reasoning method, device, equipment and storage medium based on image sequences. Background Art

[0002] Current 3D spatial reasoning technologies rely heavily on high-quality 3D structural data. Mainstream approaches typically utilize specialized depth sensors or multimodal data acquisition equipment to acquire point clouds, voxels, or other dense 3D information, then combine these with complex neural network models to model and understand spatial relationships. While these approaches offer some spatial reasoning capabilities, they still have significant shortcomings in practical applications.

[0003] On the one hand, existing technologies have high requirements for input data type and structure, resulting in high deployment costs and technical barriers to the overall system, making it difficult to adapt to the needs of dynamic, complex, and resource-constrained scenarios. In particular, in environments lacking complete three-dimensional information or with limited hardware equipment, existing methods struggle to achieve stable and reliable spatial reasoning results. On the other hand, existing spatial reasoning systems are mostly limited to specific closed scenes, lack flexible data acquisition and processing mechanisms, and are unable to efficiently complete spatial information expression and multimodal understanding through lightweight image data, which significantly limits the scope of application and practical value of the system.

[0004] In the field of financial technology, there is a demand for spatial understanding to assist business flow and service optimization in scenarios such as intelligent customer service, virtual counters, and remote risk control. However, due to the high cost and complex deployment of existing three-dimensional reasoning technology, it is difficult for financial systems to quickly obtain spatial information through flexible and lightweight means, and they are unable to meet the actual needs of efficient business development and rapid adaptation to the environment.

[0005] In the healthcare sector, scenarios such as ward management, equipment deployment, and personnel positioning place real-time and flexible demands on spatial reasoning. However, existing technologies rely on complex 3D modeling and structured data input, lacking universal and rapid spatial information extraction capabilities. This makes it difficult to adapt to the rapid deployment requirements of diverse medical scenarios, hindering the universality and efficiency of smart medical systems in changing environments.

[0006] Overall, existing three-dimensional spatial reasoning technology has obvious shortcomings in data acquisition form, system deployment flexibility, and cross-scenario adaptability. There is an urgent need to provide a technical path that does not rely on complex three-dimensional structures and can quickly complete spatial reasoning and multimodal understanding based on lightweight image data and relatively simple spatial information, so as to improve the system's versatility, practicality and cross-domain application capabilities. Summary of the Invention

[0007] The main purpose of the present invention is to provide a spatial reasoning method, device, equipment and storage medium based on image sequences, aiming to solve the technical problem that the existing technology generally relies on complex three-dimensional structural data and specialized spatial modeling modules, and it is difficult to achieve flexible and efficient spatial reasoning based only on ordinary images and simplified spatial information.

[0008] To achieve the above object, the present invention provides a spatial reasoning method based on image sequences, comprising:

[0009] Obtaining an image sequence and camera pose information corresponding to each frame of the image sequence;

[0010] According to the image sequence and the camera pose information, key frames and camera pose information corresponding to the key frames are screened based on visual language features and spatial features;

[0011] Constructing a multimodal prompt based on the keyframes, the camera pose information corresponding to the keyframes, and the spatial reasoning request;

[0012] The multimodal prompt is input into a pre-trained multimodal language model to generate a spatial reasoning result in response to the spatial reasoning request.

[0013] Furthermore, to achieve the above-mentioned object, the present invention provides a spatial reasoning device based on image sequences, comprising:

[0014] A data acquisition module is used to obtain an image sequence and camera pose information corresponding to each frame of the image sequence;

[0015] A key frame screening module is used to screen out key frames and camera pose information corresponding to the key frames based on visual language features and spatial features according to the image sequence and the camera pose information;

[0016] A prompt construction module, configured to construct a multimodal prompt based on the keyframes, the camera pose information corresponding to the keyframes, and the spatial reasoning request;

[0017] The spatial reasoning module is configured to input the multimodal prompt into a pre-trained multimodal language model and generate a spatial reasoning result in response to the spatial reasoning request.

[0018] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer device, which includes a memory, a processor, and an image sequence-based spatial reasoning program stored in the memory and runnable on the processor. When the image sequence-based spatial reasoning program is executed by the processor, the steps of the image sequence-based spatial reasoning method as described above are implemented.

[0019] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which a spatial reasoning program based on an image sequence is stored. When the spatial reasoning program based on an image sequence is executed by a processor, the steps of the spatial reasoning method based on an image sequence as described above are implemented.

[0020] Beneficial effects: The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as financial technology and medical health. A spatial reasoning method, apparatus, equipment and medium based on image sequences are disclosed, including: obtaining an image sequence and camera pose information corresponding to each frame in the image sequence, filtering key frames and camera pose information corresponding to the key frames based on the image sequence and the camera pose information, combining visual language features with spatial features, constructing multimodal prompts based on the key frames, camera pose information and spatial reasoning requests, and generating spatial reasoning results using a pre-trained multimodal language model. The present invention combines visual language features with spatial features to filter key frames with richer information and wider spatial coverage, and constructs multimodal prompts based on key frames and camera pose information, so that the multimodal language model can efficiently understand the spatial relationship in the scene, and can achieve flexible and low-cost spatial reasoning without relying on complex three-dimensional data modeling. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:

[0022] Figure 1 A schematic diagram of an application environment of a spatial reasoning method based on image sequences according to an embodiment of the present invention;

[0023] Figure 2 1. A schematic flow chart of an embodiment of a spatial reasoning method based on image sequences according to the present invention;

[0024] Figure 3 Schematic diagram of functional modules of a preferred embodiment of the spatial reasoning device based on image sequences of the present invention;

[0025] Figure 4 A schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0026] Figure 5 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0027] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0028] The spatial reasoning method based on image sequence provided by the embodiment of the present invention can be applied in the following fields: Figure 1In an application environment, the user terminal communicates with the server terminal through a network. The server terminal can obtain the image sequence and the camera pose information corresponding to each frame in the image sequence through the user terminal, and based on the image sequence and the camera pose information, combine the visual language features and the spatial features to screen the key frames and the camera pose information corresponding to the key frames, build a multimodal prompt based on the key frames, the camera pose information and the spatial reasoning request, and use the pre-trained multimodal language model to generate the spatial reasoning result. The present invention combines the visual language features and the spatial features to screen the key frames with richer information and wider spatial coverage, and builds the multimodal prompt based on the key frames and the camera pose information, so that the multimodal language model can efficiently understand the spatial relationship in the scene, and can achieve flexible and low-cost spatial reasoning without relying on complex three-dimensional data modeling. Among them, the user terminal can be but is not limited to various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server terminal can be implemented by an independent server or a server cluster composed of multiple servers. The present invention is described in detail below through specific embodiments.

[0029] See also Figure 2 , Figure 2 This is a flowchart of an embodiment of the spatial reasoning method based on image sequences provided by the present invention. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0030] like Figure 2 As shown, the spatial reasoning method based on image sequence proposed by the present invention includes the following steps:

[0031] S10, obtaining an image sequence and camera pose information corresponding to each frame of the image sequence;

[0032] In this embodiment, in the process of acquiring an image sequence and the camera pose information corresponding to each frame in the image sequence, an image sequence refers to a set of image data continuously acquired under specific temporal or spatial conditions. The image sequence can be acquired from a variety of visual acquisition devices, including monocular cameras, binocular cameras, panoramic cameras, or mobile terminal devices with image recording capabilities. Each frame in the image sequence refers to a single static image that constitutes the sequence, typically arranged in chronological order or spatial position, and each frame carries independent visual information. Camera pose information refers to the spatial position and orientation parameters of the camera corresponding to each frame in three-dimensional space. The spatial position describes the specific position of the camera in a three-dimensional coordinate system and is typically expressed as numerical values ​​in the X, Y, and Z dimensions using a Cartesian coordinate system. The orientation parameters describe the posture or rotation state of the camera and can be expressed in the form of Euler angles, quaternions, direction cosine matrices, etc. The Euler angle form includes rotation angles about the X, Y, and Z axes. The process of acquiring image sequences and camera pose information is typically accomplished using data acquisition equipment for multi-view scenarios. Multi-view scenarios encompass images captured from multiple angles, spatial locations, or timeframes. This data acquisition equipment can be a fixed-mount camera system, an imaging system mounted on a drone, a robot-integrated vision module, or a handheld terminal, capable of simultaneously acquiring image data and spatial positioning information. In practice, the data acquisition equipment continuously captures image sequences and works with a spatial positioning module to record the spatial position and orientation parameters of each frame in real time. The spatial positioning module can be based on a GNSS global navigation system, an IMU inertial measurement unit, SLAM simultaneous localization and mapping technology, visual odometry, or a combined navigation system that integrates multiple sensor information. Each image frame is synchronized with its corresponding spatial position and orientation parameters using a timestamp or data index, forming a complete image and pose data pair. This data synchronization mechanism ensures a one-to-one correspondence between image and pose information in both temporal and spatial dimensions, preventing data misalignment or loss and further improving the accuracy of subsequent spatial reasoning. In terms of data structure, image sequences can be stored in standard image file formats or compressed video streams, and camera pose information can be stored in independent data files or image metadata in the form of structured text, arrays, or matrices, enabling efficient data management and call.

[0033] In specific implementations, a fixed monitoring system can be deployed in indoor or outdoor scenarios, using a camera array to capture image sequences from different angles. Combined with a lidar or structured light sensor, the system can obtain the spatial position and orientation information of each image frame in real time, enabling large-scale, static image and pose data acquisition. Mobile robots can also integrate vision modules and IMU units to capture continuous image sequences and spatial positioning information in real time during tasks such as spatial inspection, navigation, and operations, making them suitable for data acquisition in dynamic environments. Unmanned aerial vehicles (UAVs) equipped with high-definition cameras, satellite positioning systems, and attitude sensors can be used for low-altitude aerial photography, acquiring large-scale, multi-angle image sequences and precise camera pose information. This is suitable for scenarios such as geographic surveying and mapping, environmental monitoring, and infrastructure inspections. Handheld devices can also combine visual positioning technology with inertial navigation systems to capture image sequences and high-precision camera pose information in any scenario while personnel are freely moving, improving the flexibility and applicability of data acquisition.

[0034] Example: In the healthcare business, by deploying indoor positioning systems and camera networks, you can obtain real-time image sequences and corresponding camera pose information in environments such as wards, operating rooms, and emergency areas. This assists medical systems in performing spatial layout analysis, patient behavior monitoring, and equipment management, thereby improving the spatial perception capabilities of the medical environment.

[0035] In the field of financial technology business, the security monitoring system of business outlets can be combined with spatial positioning technology to obtain image sequences and camera posture information of business halls, counters, self-service areas and other areas, supporting financial institutions to conduct customer behavior analysis, risk monitoring and space optimization management, and improve the intelligence level and security management capabilities of financial service spaces.

[0036] This embodiment obtains the image sequence and the camera pose information corresponding to each frame of the image sequence, and can establish a correspondence between data and space based on ordinary image data and spatial position parameters without relying on complex hardware structures. It avoids the high cost and high computing power burden brought by the existing reliance on complex data forms such as point clouds, voxels or structured atlases, reduces the system's dependence on hardware equipment and computing resources, and improves the system's flexibility and practicality.

[0037] S20, filtering out key frames and camera pose information corresponding to the key frames based on visual language features and spatial features according to the image sequence and the camera pose information;

[0038] In this embodiment, an image sequence refers to a collection of multiple consecutive images captured by an image acquisition device. The image sequence can originate from fixed monitoring equipment, a camera module mounted on a mobile terminal, an unmanned aerial vehicle (UAV) vision system, or other devices with image capture capabilities. Each frame in the image sequence is independent image data, reflecting visual information at different times or spatial locations. Camera pose information refers to the spatial position and orientation parameters corresponding to each frame in the image sequence. The spatial position can be expressed using position coordinates in a three-dimensional coordinate system, and the orientation parameters can be expressed based on Euler angles, direction vectors, or quaternions to accurately reflect the specific posture state of the image acquisition device in three-dimensional space.

[0039] Visual language features refer to high-dimensional feature expressions with semantic information extracted from image content. They not only include basic information such as the image's texture structure, color distribution, and shape boundaries, but also integrate scene understanding, object category recognition, and spatial semantic reasoning capabilities acquired through training visual models. Visual language features can be expressed in the form of feature vectors, structured labels, or other multidimensional forms. Spatial features refer to a set of parameters generated by combining camera pose information to reflect the position, distribution, and orientation of an image in the overall space. Spatial features can include the spatial position of a single frame, shooting angle, field of view coverage, spatial proximity between images, and position difference indicators.

[0040] Keyframe screening involves removing images with high information redundancy and repetitive spatial distribution from an image sequence, retaining a collection of image frames with rich information expression, reasonable spatial layout, and strong semantic complementarity. This process comprehensively considers both visual language and spatial features to ensure that the screening results are representative and diverse in both semantic content and spatial information. Keyframes are subsets of images selected from an image sequence that possess valuable spatial structural information and visual semantic expression. The camera pose information corresponding to keyframes refers to the spatial position and orientation data that corresponds one-to-one with the selected keyframes, and is used for subsequent spatial relationship analysis, scene modeling, or multimodal reasoning tasks.

[0041] The screening process involves extracting visual language features for each image frame, generating spatial features based on camera pose information, and calculating the semantic similarity and spatial distance between any two frames in the image sequence. Semantic similarity reflects the degree of redundancy in the semantic information dimension, while spatial distance reflects the positional difference between the images in three-dimensional space. A joint distance metric is generated by fusing semantic similarity and spatial distance. This metric comprehensively measures the degree of redundancy or information complementarity between the two frames in both the semantic and spatial dimensions. The image sequence is screened using the joint distance metric, prioritizing the retention of frames with larger joint distance metric values ​​and eliminating frames with smaller joint distance metric values ​​and higher information redundancy to form a set of key frames. The selected key frames are combined with the corresponding camera pose information to ensure that the spatial reasoning process is analyzed and processed only based on semantically rich and spatially well-organized image and spatial data.

[0042] Visual encoding networks can be used to extract visual language features from each frame in an image sequence. These networks can be implemented using convolutional neural networks, visual transformer architectures, or other models with semantic extraction capabilities. Visual language features can be expressed using multidimensional feature vectors, semantic label sets, or visual concept structures. Camera pose information corresponding to each frame can be generated using a high-precision positioning system or the spatial positioning module built into the image acquisition device. Position data in the camera pose information can be expressed in terms of Global Positioning System coordinates, local space coordinates, or a custom reference coordinate system, and orientation data can be expressed in terms of Euler angles, orientation matrices, or rotation vectors. Spatial features can be calculated for each image frame using a spatial analysis module. These features may include spatial position parameters, shooting angle, field of view coverage, or spatial proximity between images. Semantic similarity between image pairs can be calculated using a similarity calculation module based on the visual language features. Semantic similarity can be quantified using cosine similarity, Euclidean distance, contrastive learning results, or other metrics. Spatial distance between image pairs can be calculated based on these spatial features. Spatial distance can be expressed using Euclidean distance, Mahalanobis distance, or a weighted distance model that incorporates orientation. A weighted fusion strategy can be used to generate a joint distance metric. This metric can adjust the weights of semantic similarity and spatial distance based on actual application requirements, improving screening flexibility and applicability. Image sequences can be screened based on this joint distance metric, removing frames with high information redundancy and repetitive spatial distribution. Frames with strong information expression and complementary spatial layouts are retained as a set of keyframes. The camera pose information corresponding to these keyframes is then extracted to form the data input foundation for subsequent analysis.

[0043] Example: In the healthcare business, image acquisition devices can be deployed in hospital wards, operating rooms, or emergency environments to obtain environmental image sequences and camera pose information. Visual language features and spatial features are combined to select key frames and corresponding camera pose information to form an image collection covering key areas of the environment and key equipment. This assists the medical system in real-time understanding the environmental layout, personnel location, and equipment status, thereby improving the efficiency of medical environment monitoring and emergency response.

[0044] In the field of financial technology business, for bank branches, smart teller areas or financial self-service terminal areas, business scene image sequences and camera pose information can be obtained, key frames and corresponding camera pose information can be selected based on visual language features and spatial features, and scene data sets with spatial structure expression capabilities can be constructed to improve the spatial perception ability and intelligent analysis level of financial scenes, optimize business processing procedures, and strengthen risk monitoring and security management capabilities.

[0045] This embodiment combines visual language features with spatial features to filter image sequences. This can reduce the size of image data while retaining key spatial layout information and semantic expression content, improve the data efficiency and accuracy of spatial analysis and multimodal reasoning tasks, reduce the system's dependence on hardware equipment and computing resources, and enhance the flexibility, real-time performance, and environmental adaptability of the spatial perception system.

[0046] S30, constructing a multimodal prompt based on the key frame, the camera pose information corresponding to the key frame, and the spatial reasoning request;

[0047] In this embodiment, a keyframe refers to a representative subset of images selected from an image sequence. Keyframes are obtained by comprehensively considering the image's visual information expression capabilities and spatial layout characteristics. They exhibit high information density, rich semantic expression, and a reasonable spatial distribution. Camera pose information is the spatial position and orientation data corresponding to the keyframe. This information reflects the specific position coordinates and orientation of the image acquisition device in three-dimensional space. Position data can be expressed based on a spatial coordinate system, while orientation data can be expressed using Euler angles, direction vectors, rotation matrices, or quaternions, ensuring complete and accurate spatial reference information during spatial reasoning.

[0048] A spatial reasoning request refers to a natural language expression made by a user or system regarding spatial relationships, location layout, or object orientation. A spatial reasoning request can involve spatial structure understanding, relative position judgment, spatial layout inference, or spatial semantic analysis. The expression can be in the form of natural language questions, instructions, declarative sentences, or other structured or unstructured language content. The source of a spatial reasoning request can be user input, system autonomous generation, or synchronous acquisition of external information.

[0049] Multimodal prompts refer to the fusion expression of image information, spatial information and language information, which are used to guide multimodal language models to perform spatial reasoning tasks. Multimodal prompts contain multidimensional information such as images, text, spatial parameters, etc., ensuring that the model can comprehensively understand the relationship between visual, spatial and language information during the reasoning process, and generate reasoning results that conform to spatial reality and language habits.

[0050] The process of constructing multimodal prompts involves extracting keyframe images and camera pose information, generating scene description text, image tiles, and spatial reference information, and combining them with spatial reasoning requests to form a structured input data set. Based on the camera pose information, position coordinates and orientation parameters can be extracted. Position coordinates represent the positional relationship of the image in three-dimensional space, and the orientation parameters can be converted into Euler angles for subsequent language model understanding and processing. Spatial position description text can be generated based on the position coordinates and Euler angles. This spatial position description text expresses the spatial position and viewing direction of the image in natural language, improving the language model's ability to understand spatial information. Keyframe images and spatial position description text can be combined into image tiles. These tiles unify visual and spatial information in a structured data format, facilitating multimodal language model processing and fusion. Multimodal prompts can be combined with spatial reasoning requests to integrate image tiles, spatial position description text, and spatial reasoning requests. This ensures that the model input contains complete semantic information, spatial structure information, and user requirement information, improving the model's understanding and generation accuracy in spatial reasoning tasks.

[0051] The spatial position of keyframes can be extracted based on the 3D coordinates in the camera pose information. Spatial positions can be expressed in a global coordinate system, a local reference system, or a custom coordinate system to ensure uniformity and accuracy. Euler angles can be generated based on the orientation parameters in the camera pose information. Euler angles include pitch, yaw, and roll, clearly expressing the camera's viewing direction and attitude, facilitating spatial relationship analysis and language description. Spatial position description text can be generated based on the position coordinates and Euler angles. This description text can include information such as image acquisition location, orientation angle, and field of view. It uses a natural language structure to enhance the language model's ability to perceive and understand spatial information. Keyframe images and spatial position description text can be combined to form a graphic block. This graphic block can be structured data and contain image data, spatial parameters, and auxiliary text information, ensuring the coordinated expression of visual and spatial information. Spatial reasoning requests can be received. Spatial reasoning requests can express specific spatial relationship judgment, layout analysis, or position inference requirements in natural language, and can be expressed as questions, instructions, or statements. It can integrate graphic blocks, spatial location description text and spatial reasoning requests to generate multimodal prompts. Multimodal prompts have the characteristics of structured expression, comprehensive information and clear semantics, which facilitates the understanding and reasoning of multimodal language models.

[0052] Example: In the healthcare business, based on keyframe images and camera pose information of hospital wards, emergency environments, or operating rooms, combined with spatial reasoning requests from doctors or nurses, multimodal prompts can be constructed that include scene images, spatial location descriptions, and spatial reasoning requirements. This assists intelligent systems in understanding environmental layout, equipment location, and personnel orientation, thereby improving medical space perception and emergency response capabilities.

[0053] In the field of financial technology business, based on the key frame images and camera pose information of bank branches, smart teller areas or financial terminal areas, combined with spatial reasoning requests generated by users or systems, multimodal prompts containing scene images, spatial structure information and spatial reasoning requirements can be constructed to improve spatial perception, location analysis and risk monitoring capabilities in financial scenarios, and optimize business processing procedures and security levels.

[0054] By constructing multimodal prompts that include visual, spatial, and language information, this embodiment can effectively improve the understanding ability and generation accuracy of multimodal language models in spatial reasoning tasks, reduce dependence on complex three-dimensional structure modeling and specialized encoder design, enhance the system's adaptability and generalization capabilities in changing scenarios and complex tasks, and improve the practicality, flexibility, and deployment efficiency of the spatial reasoning system.

[0055] S40: Input the multimodal prompt into a pre-trained multimodal language model to generate a spatial reasoning result in response to the spatial reasoning request.

[0056] In this embodiment, a multimodal prompt is a structured data set containing visual, spatial, and linguistic information. Constructed by integrating keyframe images, corresponding camera pose information, and spatial reasoning requests, the multimodal prompt has the ability to fully express both semantic and spatial structural information. A multimodal language model is a parameterized neural network model capable of simultaneously understanding both visual and linguistic information. It is typically trained on large-scale data to achieve generalized expression and reasoning capabilities across multiple domains and tasks. The multimodal language model is capable of processing images, text, and spatial data, and can generate reasoning results that align with semantic logic and spatial facts based on the input multimodal prompt.

[0057] A spatial reasoning request is a linguistic expression from a user or system regarding spatial relationships, object locations, or environmental layout. It can be a question, instruction, or statement in natural language. It serves to clarify the goal and direction of the reasoning task. Spatial reasoning results are language outputs generated based on multimodal prompts and spatial reasoning requests. These results can be descriptions of spatial relationships, position determinations, layout analyses, or other spatial information expressions that conform to language logic. These results are expressed in natural language, making them easier for users to understand and for subsequent system applications.

[0058] Inputting multimodal prompts into a multimodal language model allows for unified encoding of image data, text data, and spatial information into a model-recognizable input format through interface calls, data structure encapsulation, or serialization. Based on the input multimodal prompts, the multimodal language model combines its internally trained multimodal expression and spatial reasoning capabilities to automatically perform information fusion, semantic understanding, and spatial reasoning, generating language output that meets the requirements of spatial reasoning requests. The spatial reasoning results are output through the model decoder, featuring clear language expression, accurate spatial logic, and complete semantic structure, ensuring the effective completion of reasoning tasks.

[0059] Based on an API interface or local call method, the image information, spatial information, and language information in the multimodal prompt can be formatted and encoded into the data input structure required by the multimodal language model. The data structure can use sequence encoding, embedded representation, or other standard formats to ensure the accuracy and completeness of information transmission. Based on a large-scale multimodal language model that has been trained, the model reasoning interface can be called to input multimodal prompts. The model automatically completes the visual information parsing, spatial information mapping, and language information understanding processes based on the image and text information fusion module and spatial reasoning mechanism. Based on the semantic content of the spatial reasoning request, the multimodal language model can be guided to focus on information such as spatial relationships, location layout, or object orientation. The model generates language output results that conform to semantic logic and spatial facts based on training parameters and reasoning mechanisms. The final spatial reasoning results can be parsed based on the sequence data or text data output by the model. The spatial reasoning results have a natural language expression form, which is convenient for user understanding, system feedback, or subsequent intelligent processing.

[0060] Example description: In the healthcare business field, based on the key frame images and spatial location data of the hospital emergency room, intensive care unit or operating room, combined with the spatial reasoning requests raised by medical staff regarding equipment layout, personnel orientation or environmental accessibility, the constructed multimodal prompts can be input into the multimodal language model to automatically generate spatial reasoning results about the position relationship of objects, the relative orientation of personnel or environmental layout analysis, to assist in medical space management and emergency response.

[0061] In the field of financial technology business, based on the key frame images and camera pose information of the bank business area, self-service terminal area or financial service area, combined with the spatial reasoning requests raised by the system or user regarding the physical layout, route or spatial safety, multimodal prompts can be input into the multimodal language model to automatically generate spatial reasoning results about location distribution, spatial relationships or safety hazard analysis, thereby improving the level of spatial perception and intelligent decision-making in the financial business environment.

[0062] By inputting multimodal cues containing visual, spatial, and language information into a multimodal language model, this embodiment can fully leverage the comprehensive capabilities of large-scale models in multimodal fusion, language generation, and spatial reasoning, effectively improving the system's information understanding ability, reasoning accuracy, and output expression effects in spatial reasoning tasks, reducing its reliance on complex three-dimensional structure modeling and specialized spatial reasoning module design, enhancing the system's adaptability and generalization capabilities in unknown environments and changing tasks, and improving the practicality, flexibility, and deployment efficiency of the overall spatial reasoning system.

[0063] The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as financial technology and medical health. A spatial reasoning method, device, equipment and medium based on image sequences are disclosed, including: obtaining an image sequence and camera pose information corresponding to each frame in the image sequence, based on the image sequence and the camera pose information, combining visual language features with spatial features to screen key frames and camera pose information corresponding to the key frames, constructing multimodal prompts based on the key frames, camera pose information and spatial reasoning requests, and generating spatial reasoning results using a pre-trained multimodal language model. The present invention combines visual language features with spatial features to screen key frames with richer information and wider spatial coverage, and constructs multimodal prompts based on the key frames and camera pose information, so that the multimodal language model can efficiently understand the spatial relationship in the scene, and can achieve flexible and low-cost spatial reasoning without relying on complex three-dimensional data modeling.

[0064] In one embodiment, the above step S10 includes:

[0065] S101, collecting image sequences of multi-view scenes;

[0066] S102, obtaining camera position information of each frame of image in the image sequence;

[0067] S103, obtaining camera orientation information of each frame of image in the image sequence;

[0068] S104: Generate camera pose information corresponding to each frame of image based on the camera position information and the camera orientation information.

[0069] In this embodiment, an image sequence is a collection of multiple frames of images obtained by continuously capturing the same scene from different positions or angles in space using an image acquisition device. This image sequence can include image information from different perspectives, orientations, and time points. The acquisition of this image sequence provides a rich source of visual information for subsequent spatial structure understanding and positional relationship reasoning. Image sequences can be acquired using cameras installed at different locations or by continuously capturing images from different locations within a scene using a single mobile device. Image sequences possess diversity, continuity, and spatial distribution information.

[0070] A multi-perspective scene refers to the spatial information displayed by the photographed object or environment at different shooting positions or angles. The acquisition of image sequences of multi-perspective scenes can be achieved by arranging multiple fixed cameras, deploying mobile acquisition equipment, or using carriers such as drones. Obtaining multi-perspective image sequences can enhance the comprehensiveness and completeness of spatial information expression, which is helpful for subsequent spatial feature extraction and relative position relationship analysis.

[0071] Camera position information refers to the position coordinate information of the camera in three-dimensional space corresponding to each frame of the image. It is usually expressed as X, Y, and Z values ​​in a Cartesian coordinate system, reflecting the absolute position of the camera in the scene. Camera position information can be obtained through GPS positioning, inertial navigation, visual odometry, or other spatial positioning technologies. Camera orientation information refers to the direction or orientation angle of the camera when capturing the image. It is usually expressed in the form of Euler angles, quaternions, or rotation matrices. Camera orientation information reflects the directional characteristics of the image capture process and plays an important role in accurately restoring spatial structure and positional relationships.

[0072] Camera pose information is a joint expression based on camera position information and camera orientation information. It fully describes the position and orientation of the shooting device in space. It is usually expressed in six-degree-of-freedom parameters, pose matrix or other standard expression structures. Camera pose information is used to accurately locate the shooting position and direction of each frame in three-dimensional space, providing a data basis for subsequent spatial reasoning and scene reconstruction.

[0073] This embodiment captures image sequences of multi-view scenes, combines the camera position information and camera orientation information corresponding to each frame of image, and generates camera pose information corresponding to each frame of image. This can comprehensively and accurately express the spatial structure and image acquisition position status in the scene, enhance the system's ability to understand and express spatial information, and provide reliable data support for subsequent spatial feature extraction, key frame screening, and spatial reasoning tasks, effectively improving the integrity of spatial perception, the accuracy of position relationship expression, and the flexibility of system processing.

[0074] In one embodiment, the above step S20 includes:

[0075] S201, extracting visual language features of each frame of the image in the image sequence;

[0076] S202, generating a spatial feature of each frame of image based on the camera pose information;

[0077] S203, determining the semantic similarity between every two frames of the image sequence based on the visual language features;

[0078] S204, determining a spatial distance between every two frames of image in the image sequence based on the spatial feature;

[0079] S205, fusing the semantic similarity and spatial distance to generate a joint distance measurement value between every two frames of images;

[0080] S206 , screening a key frame subset from the image sequence according to the joint distance metric value, and extracting camera pose information corresponding to the key frame subset from the camera pose information.

[0081] In this embodiment, an image sequence refers to a collection of multiple images captured from the same spatial scene. Each frame in the image sequence exhibits spatial continuity and positional correlation. Image sequences can be acquired using fixed cameras, mobile terminals, or unmanned devices, covering different locations and angles of the scene, providing a multi-perspective, multi-directional source of visual information. Camera pose information describes the spatial position and orientation of each frame in the image sequence. This information, combined with the image sequence, expresses the distribution characteristics and viewing angles of each frame in space.

[0082] Visual language features refer to high-dimensional feature expressions formed by encoding the semantic information and visual expression information of each frame in an image sequence through an image analysis model. Visual language features may include image content, scene attributes, object information and semantic labels. The extraction process of visual language features is usually based on a visual encoding network, a visual language pre-training model or an image feature extraction algorithm, and is used to express the semantic connotation and visual information structure of the image.

[0083] Spatial features are expression information that reflects the position and orientation characteristics of each frame image in three-dimensional space, which is calculated based on the camera pose information. Spatial features can be generated by coordinate mapping, spatial projection or spatial encoding to express the spatial position, orientation distribution and relative relationship of each frame image. Spatial features are used to quantify the spatial structure information and position relationship of each frame image in the image sequence.

[0084] Semantic similarity is a measure of the degree of similarity between two frames in an image sequence at the semantic level, calculated based on visual language features. Semantic similarity can be obtained through cosine similarity, Euclidean distance, or other similarity calculation methods, and is used to analyze the relevance of different frames in terms of content, structure, or semantic expression.

[0085] Spatial distance is a metric value calculated based on spatial features that reflects the relative distance between every two frames in an image sequence in three-dimensional space. Spatial distance is usually calculated based on spatial coordinate differences, Euclidean distance, Mahalanobis distance, or spatial position relationship analysis methods. Spatial distance is used to quantify the degree of proximity and position overlap between different frames in spatial distribution.

[0086] The joint distance metric is a comprehensive metric generated by fusing semantic similarity and spatial distance. The joint distance metric combines the semantic information and spatial information of the image to comprehensively evaluate the overall redundancy and complementarity between each two frames. The generation process of the joint distance metric can be implemented based on weighted fusion, distance mapping or comprehensive evaluation model. The joint distance metric is used to guide the key frame screening process.

[0087] A keyframe subset is a collection of image frames selected from an image sequence based on a joint distance metric, characterized by rich information, spatial complementarity, and low redundancy. This subset preserves the spatial structure and core semantic information in the image sequence, helping to reduce data redundancy and computational complexity. The camera pose information corresponding to the keyframe subset refers to the spatial position and orientation information extracted from the camera pose information, corresponding to the keyframe subset. The keyframe subset and the corresponding camera pose information are used together for subsequent spatial reasoning and multimodal information fusion.

[0088] In the specific screening process, the threshold of the joint distance metric value can be set to 0.75, and the value range of the joint distance metric value is 0 to 1. The smaller the value, the higher the similarity between the two frames of images in terms of semantic information and spatial information, and the greater the redundancy. For any two frames of images in the image sequence, if the calculated joint distance metric value is less than 0.75, it is determined that the content of the two frames of images is redundant, and the latter can be eliminated. If the joint distance metric value is greater than or equal to 0.75, it indicates that the information between the two frames of images is complementary and both can be retained. For example, for an image sequence containing 50 frames of images, the above threshold screening is used, and eventually 10 to 15 frames of key frames with rich information, dispersed spatial positions, and complementary semantic expressions can be retained, and the corresponding camera pose information can be extracted at the same time to achieve optimization of spatial structure expression and effective compression of data scale.

[0089] This embodiment extracts visual language features from each frame in an image sequence and generates spatial features based on camera pose information, jointly analyzes the semantic information and spatial structure of the image, calculates the semantic similarity and spatial distance between every two frames, and fuses them to generate a joint distance metric. This can comprehensively evaluate the content redundancy and spatial distribution characteristics of the image sequence, and uses the joint distance metric to filter a subset of key frames and the corresponding camera pose information, effectively preserving the spatial information expression and semantic expression of the image sequence, reducing data redundancy, improving the efficiency and accuracy of spatial structure reconstruction and reasoning, and enhancing the system's applicability and spatial perception capabilities in different scenarios.

[0090] In one embodiment, the above step S206 includes:

[0091] S2061, determining an edge sharpness value of each frame of the image in the image sequence as a clarity index;

[0092] S2062, determining the field of view angle coverage of each frame of image according to the camera pose information;

[0093] S2063: Adding a clarity weight coefficient to the joint distance metric based on the clarity index to generate a joint distance metric with clarity weight;

[0094] S2064: Based on the field of view angle coverage, add a field of view angle priority weight to the frames with complementary perspectives to generate a joint distance metric value with the field of view angle weight;

[0095] S2065, fusing the joint distance metric value with the clarity weight and the joint distance metric value with the field of view angle weight to obtain a weighted joint distance metric value;

[0096] S2066 , filtering a key frame subset from the image sequence according to the weighted joint distance metric value;

[0097] S2067: Extract camera pose information corresponding to the key frame subset from the camera pose information.

[0098] In this embodiment, an image sequence refers to a set of image data captured by a single camera or multiple cameras, arranged in chronological or logical order. This image sequence can originate from video frame capture, fixed-point timer photography, or continuous filming by drones, robots, or surveillance systems. Each frame in the image sequence corresponds to an independent spatial perspective, reflecting scene information from different positions and orientations, helping to capture the complete spatial layout and visual details.

[0099] Camera pose information is used to describe the spatial position and orientation of each frame of image. The spatial position is usually expressed by the X, Y, and Z coordinates in a three-dimensional coordinate system, and the spatial orientation can be expressed in the form of Euler angles, rotation matrices, or quaternions. Combining the two can accurately locate the specific position and posture of the camera in space, providing a data basis for subsequent spatial relationship reasoning and image position association.

[0100] Edge sharpness is an important indicator for measuring image clarity. It is usually calculated through methods such as Laplace transform, Sobel operator, Canny edge detection, or gradient amplitude analysis. A larger value indicates a clearer image boundary contour and richer detail information, which helps to select image frames with high visual quality and complete information expression.

[0101] The field of view angle coverage is calculated based on the camera pose information. Combined with the camera's optical parameters, focal length, imaging size and orientation information, spatial geometric transformation is used to determine the coverage area of ​​each frame image in three-dimensional space. The larger the field of view angle coverage, the wider the spatial information provided by the image and the stronger the spatial expression ability.

[0102] The clarity weight coefficient is dynamically adjusted according to the edge sharpness value of each frame image. Normalization or interval mapping is usually used to keep the weight value within a reasonable range, such as the range of 0 to 1. Images with high clarity are given greater weights to increase the image's participation in the joint distance metric value, and image frames with high clarity and clear information expression are retained first.

[0103] The joint distance metric with clarity weight is calculated based on the original joint distance metric and the clarity weight coefficient. It retains the basic information of image semantic similarity and spatial distance, while introducing the clarity factor to enhance the role of clarity in the screening process and improve the visual information quality of the final keyframe.

[0104] The field of view priority weight is calculated based on the field of view coverage of each frame image. The weight value is dynamically allocated with the spatial coverage area and spatial complementarity. Images with large spatial coverage, unique perspective information or complementarity are given higher weights, which enhances the participation of the image in the joint distance metric and optimizes the spatial layout expression effect.

[0105] The joint distance metric with field of view angle weight is generated based on the original joint distance metric and combined with the field of view angle priority weight, ensuring that the spatial information coverage factor is reflected in the screening process and improving the overall spatial information contribution and expression integrity of the retained image.

[0106] The weighted joint distance metric is obtained by fusing the joint distance metric with clarity weight and the joint distance metric with field of view weight. The fusion process can be implemented based on weighted average, linear combination or other mathematical methods, taking into account the image clarity, spatial coverage, semantics and spatial differences to form a screening reference indicator with strong expression ability and high information value.

[0107] The keyframe subset is filtered based on the weighted joint distance metric. A reasonable filtering threshold is set, such as 0.75. The joint distance metric value range is usually between 0 and 1. The smaller the value, the more similar the semantic and spatial information of the two frames of image are, and the higher the redundancy. The larger the value, the stronger the complementarity of the information of the two frames of image is, and the higher the retention value. If the weighted joint distance metric of the two frames of image is less than or equal to the set threshold, it is determined that there is a high degree of redundancy in the content of the two frames of image, and subsequent images can be eliminated; if the weighted joint distance metric is greater than the set threshold, it indicates that the image information is clearly complementary, and the frame of image is retained as a keyframe. After the screening is completed, a subset of keyframes with rich information, reasonable spatial distribution, and complementary semantic expression is retained, while effectively controlling data redundancy, reducing the computational burden, and improving the efficiency of spatial expression.

[0108] Extracting the camera pose information corresponding to the keyframe subset means retrieving the corresponding camera pose information in the image sequence based on the filtered keyframe subset, extracting the spatial position and orientation information corresponding to each keyframe, and ensuring a one-to-one correspondence between the keyframe and the spatial position and orientation data, which facilitates subsequent spatial reasoning, three-dimensional relationship modeling, and multimodal information fusion, and ensures the data integrity and spatial expression accuracy of the reasoning process.

[0109] In the actual screening process, the edge sharpness value is first calculated for each frame in the image sequence. Assume that after the second-order derivative operation of the Laplace operator is performed on the image, the edge sharpness value of image A is 250, the edge sharpness value of image B is 180, and the edge sharpness value of image C is 90. The sharpness values ​​are normalized to the range of 0 to 1, and the corresponding clarity weight coefficients for images A, B, and C are 1.0, 0.72, and 0.36, respectively. The higher the clarity weight, the better the image quality.

[0110] Combined with the camera pose information, the field of view coverage of each frame is calculated. Assume that the spatial coverage area of ​​image A is 20 square meters, that of image B is 15 square meters, and that of image C is 8 square meters. Assign field of view priority weights based on the spatial coverage area, setting the maximum spatial coverage weight to 1.0. The field of view priority weights for images A, B, and C are 1.0, 0.75, and 0.4, respectively. A larger coverage area indicates a higher contribution to spatial information.

[0111] During the fusion phase, the original joint distance metric is calculated based on the assumption that the distances between images A and B are 0.15, between B and C is 0.08, and between A and C is 0.20. After adding clarity weighting and field of view priority weighting to these values, a clarity-weighted joint distance metric and a field of view-weighted joint distance metric are generated. For images A and B, the clarity-weighted joint distance metric is 0.15 × (1.0 + 0.72) / 2 = 0.129, and the field of view-weighted joint distance metric is 0.15 × (1.0 + 0.75) / 2 = 0.13125.

[0112] Finally, the two types of weighted results are fused and a weighted joint distance metric value is generated by weighted averaging. The weighted joint distance metric value of images A and B is (0.129+0.13125) / 2=0.130125.

[0113] The weighted joint distance metric screening threshold is set to 0.12. If the weighted joint distance metric between two frames is less than or equal to the threshold, it is determined that the image information is redundant and the frame with relatively poor information quality needs to be eliminated. If it is greater than the threshold, the two frames of image are retained.

[0114] Combining the above data, the weighted joint distance metric between images A and B is 0.130125, which is greater than the threshold of 0.12, indicating that the information in the two frames is complementary and both are retained. The weighted joint distance metric between images B and C is assumed to be 0.09, which is lower than the threshold of 0.12, indicating that there is information redundancy. Comparing the clarity and spatial coverage indicators, image B has a clarity weight of 0.72 and a field of view weight of 0.75. Image C has a clarity weight of 0.36 and a field of view weight of 0.4. After comprehensive consideration, image B has higher information quality, so image B is retained and image C is eliminated.

[0115] After the screening is completed, images A and B are retained to form a keyframe subset, and the corresponding camera pose information is further extracted to ensure that each keyframe matches the complete spatial position and orientation data, providing a data foundation with clear information and complete spatial expression for subsequent spatial reasoning.

[0116] This embodiment combines the image edge sharpness value and the camera field of view angle coverage range and introduces a dynamic weight adjustment mechanism for the joint distance metric value. This can effectively improve the comprehensive consideration of image quality and spatial information coverage during the key frame screening process, avoid retaining image frames that are blurred, have low information contribution, or have severe spatial overlap, and thus screen out a subset of key frames with high clarity, complementary spatial positions, and rich information expression while controlling the data scale. The corresponding camera pose information is also extracted simultaneously, further improving the spatial information integrity and semantic expression accuracy in the multimodal spatial reasoning process, and enhancing the reliability and practicality of spatial structure reconstruction and reasoning results.

[0117] In one embodiment, the above step S30 includes:

[0118] S301, receiving a spatial reasoning request;

[0119] S302, generating a scene guide for describing an image acquisition method and spatial characteristics for the key frame;

[0120] S303, combining each frame image in the key frame and the corresponding camera pose information into a graphic block;

[0121] S304, generating an annotation instruction for restricting the model from using image sequence number references;

[0122] S305: Integrate the spatial reasoning request, scene guide, graphic block, and annotation instruction to construct a multimodal prompt.

[0123] In this embodiment, a keyframe refers to a collection of information-rich, spatially distributed, and representative images retained through pre-filtering. Each keyframe corresponds to unique spatial position and orientation parameters, used to express scene structure and visual information. Camera pose information includes spatial position and orientation. Spatial position can be expressed as three-dimensional coordinates, while spatial orientation can be expressed as Euler angles, rotation matrices, or quaternions. Together, they reflect the specific shooting position and orientation of the keyframe in three-dimensional space.

[0124] Spatial reasoning requests are spatial understanding instructions input by external users or upper-level systems. The content may involve requirements such as spatial position determination, object relative orientation reasoning, and scene structure analysis. Spatial reasoning requests are expressed in the form of natural language text, structured commands, or parameter sets, and have clear reasoning goals and information requirements.

[0125] The process of receiving spatial reasoning requests refers to the system receiving spatial reasoning request information from the outside through input interfaces, message queues, interactive instructions, or program calls, and performing format recognition and structural analysis on the request content to ensure that the reasoning request can be correctly understood and used by subsequent modules.

[0126] The process of generating scene guides for keyframes, describing the image acquisition method and spatial characteristics, involves combining the keyframe's camera pose information with the shooting context to generate auxiliary prompts. This content can include information such as the image acquisition method (e.g., mobile device capture, drone aerial photography, fixed camera surveillance), capture time, spatial location distribution, scene layout characteristics, and ambient lighting conditions. Scene guides are typically output as natural language text or structured information to provide background knowledge and scene references for downstream reasoning models, improving spatial reasoning accuracy and semantic expression capabilities.

[0127] The operation of combining each image frame in the key frame with the corresponding camera pose information into a graphic block means extracting the image data and spatial position and orientation information for each key frame, and packaging the two in the form of a structured unit. The graphic block may include image files, image data identifiers, three-dimensional coordinate parameters, spatial orientation parameters, etc. The graphic block serves as a multimodal information input structure to ensure the synchronous expression of visual information and spatial information, which is convenient for direct processing by the multimodal model.

[0128] The process of generating annotation instructions that restrict the model's use of image sequence numbers involves designing constraints for image blocks and overall prompt content to prevent multimodal models from inferring scene structure based solely on image sequence numbers or sequential position, thereby avoiding reasoning bias that is not based on real spatial information. Annotation instructions can take the form of explicit prohibitions, regular constraint expressions, or precautions to ensure that the reasoning process relies on real image content and spatial information, rather than logical position or image order.

[0129] The process of integrating spatial reasoning requests, scene guides, graphic blocks, and annotation instructions to construct multimodal prompts refers to combining the aforementioned information into a complete multimodal prompt input according to the preset format, logical structure, and data protocol. The prompt structure can adopt a mixed format of text and images, a structured data set, a sequence encoding data, etc., to ensure that the multimodal model can fully receive multi-source information such as images, space, and text, and support high-precision spatial reasoning and multimodal fusion tasks.

[0130] Through the above operations, the system of this embodiment effectively integrates key frame images, camera pose information and external spatial reasoning requests to form a unified prompt structure containing semantic, visual and spatial multi-dimensional information. The combination of graphic blocks and scene guides significantly reduces the cost of acquiring three-dimensional data and the complexity of the system, avoiding the hardware burden of relying on high-precision depth cameras or point cloud sensors. At the same time, with the help of annotation instructions, the model's incorrect dependence on sequence positions is limited, thereby enhancing the authenticity and robustness of the spatial reasoning results. The overall process does not require the construction of a complex spatial structure encoder, and can use the multimodal language model to have zero-sample spatial reasoning capabilities, thereby improving the adaptability, scalability and practicality of the spatial understanding system.

[0131] In one embodiment, the above step S303 includes:

[0132] S3031, extracting camera position information and camera orientation information from the camera pose information;

[0133] S3032, converting the camera orientation information into Euler angle representation;

[0134] S3033, combining the camera position information and Euler angle representation to generate a posture description text;

[0135] S3034: Combine the posture description text with the image in the corresponding key frame to generate an image-text block.

[0136] In this embodiment, a keyframe refers to a collection of representative, information-rich, and spatially complementary images selected from an image sequence. Each keyframe corresponds to unique spatial position and orientation information, facilitating the expression of the spatial structure of the overall scene. Camera pose information includes the camera's position and orientation in three-dimensional space. Spatial position is generally represented by X, Y, and Z coordinates in a three-dimensional coordinate system, reflecting the camera's shooting point. Spatial orientation can be expressed using Euler angles, rotation matrices, or quaternions to describe the camera lens's orientation, tilt angle, and rotation state.

[0137] The process of extracting camera position information and camera orientation information from camera pose information refers to performing structured analysis on the camera pose information corresponding to each keyframe, extracting and separating the coordinate parameters expressing the spatial position and the direction parameters expressing the spatial orientation, respectively, to ensure that the data source is clear and the content is independent in the subsequent information combination process.

[0138] The process of converting camera orientation information into Euler angles involves applying mathematical transformations or spatial geometric inference to the extracted spatial orientation parameters, converting them into Euler angles. Euler angles typically include pitch, yaw, and roll, which visually represent the specific rotational state of the camera lens along the three-dimensional axis. Euler angles are simple and easy to understand, helping to improve the ability of multimodal models to analyze spatial orientation information and avoid misunderstandings caused by complex representations.

[0139] The process of combining camera position information and Euler angle representation to generate pose description text refers to integrating the camera's spatial position parameters and orientation angle information into a piece of text content through structured templates or natural language expression rules. The pose description text clearly expresses the shooting position and direction information of each key frame in three-dimensional space, improves the readability and interactivity of spatial information, and facilitates the multimodal language model to directly obtain spatial structural features.

[0140] The process of combining the pose description text with the image in the corresponding key frame to generate a graph-text block refers to constructing a structured input unit containing visual information and spatial information for each key frame based on the image data and the generated pose description text. The graph-text block usually contains image content, spatial position data, orientation parameters and related text information. As the basic data structure of the multimodal system input, the graph-text block ensures the synchronous expression of image information and spatial features, which helps the downstream multimodal model to efficiently and accurately understand the overall scene and spatial relationship.

[0141] Through the above operations, this embodiment can effectively deeply fuse the key frame image with the corresponding camera pose information to form a graph block containing image, spatial position and orientation data, avoiding the problems of traditional spatial reasoning methods that require separate input of multiple types of information, complex structure, and synchronization difficulties. The use of Euler angles to express spatial orientation simplifies the representation of spatial information and improves the accuracy of the multimodal language model in analyzing spatial orientation features. The introduction of pose description text enhances the interpretability and text expression ability of spatial information. The establishment of the graph block structure enables image information and spatial information to have a unified expression channel, reduces the difficulty of integrating multimodal input, and improves the flexibility and practicality of the spatial reasoning system.

[0142] In one embodiment, the above step S40 includes:

[0143] S401, parsing camera pose information in the multimodal prompt using a multimodal language model;

[0144] S402, establishing a three-dimensional space coordinate system based on the camera pose information;

[0145] S403, mapping the key frame content in the multimodal prompt to the three-dimensional space coordinate system through a multimodal language model to generate spatial mapping content;

[0146] S404: performing spatial relationship recognition processing based on the spatial reasoning request and the spatial mapping content using the multimodal language model to generate a spatial relationship description;

[0147] S405: Generate a spatial reasoning result in natural language based on the spatial relationship description.

[0148] In this embodiment, multimodal prompts include data content in various forms, including images, text, and spatial information. The overall structure combines visual expression, spatial structure, and semantic information, aiming to provide a comprehensive source of input information for multimodal language models. A pre-trained multimodal language model is a generative model trained on large-scale cross-modal data, capable of language understanding, visual parsing, and spatial reasoning. It typically possesses the ability to jointly understand and reason about images, text, and spatial information, and can be directly applied to complex spatial reasoning tasks without the need for additional task-specific training.

[0149] Parsing the camera pose information in multimodal prompts refers to performing structured analysis of the spatial position and orientation information in the input data through the spatial information understanding mechanism built into the multimodal language model, accurately extracting the spatial position coordinates and orientation parameters associated with the keyframe image, and converting them into a unified expression format that is convenient for processing by the spatial reasoning module within the model, ensuring the effective use of spatial structure information in downstream steps.

[0150] The process of establishing a 3D spatial coordinate system based on camera pose information involves building a unified 3D spatial reference frame based on the parsed spatial position and orientation data, combined with a fixed world coordinate standard or relative coordinate system. This spatial coordinate system provides the foundation for spatial mapping and positional relationship recognition of all keyframe content, avoiding spatial structure confusion or misinterpretation of positional relationships caused by the lack of a standard spatial framework.

[0151] Mapping the keyframe content in the multimodal prompt to the three-dimensional spatial coordinate system means that the multimodal language model combines the established spatial coordinate system to perform spatial position positioning and direction matching operations on the input image information, accurately associates the spatial position and orientation information corresponding to the keyframe image with the three-dimensional spatial coordinate system, and generates complete spatial mapping data. The spatial mapping content expresses the specific position relationship of each keyframe in the spatial structure, providing a reliable structural foundation for subsequent spatial reasoning.

[0152] The process of performing spatial relationship recognition processing based on spatial reasoning requests and spatial mapping content refers to the multimodal language model combining the task requirements and question types contained in the input spatial reasoning request, relying on the spatial structure information provided by the spatial mapping content, and using the built-in spatial relationship understanding mechanism to identify the spatial position relationship, directional relativity and spatial layout characteristics between different key frames, forming structured spatial relationship description data. The spatial relationship description accurately expresses the spatial structure characteristics and object position relationship in the input scene, and has clear spatial logical expression capabilities.

[0153] The process of generating spatial reasoning results in natural language based on spatial relationship descriptions refers to the multimodal language model converting structured spatial relationship description information into text content that conforms to natural language expression habits, has clear semantics, and complete structure, and finally outputs human-readable reasoning results. The reasoning results not only accurately express the spatial layout and object relationships, but also have good language fluency and expression logic, which is convenient for users to understand and downstream system calls.

[0154] Example: In a financial business scenario, multimodal spatial reasoning methods can be used to perform spatial understanding and semantic reasoning of equipment layout, security structure, and personnel distribution in complex spatial environments, such as bank branches, intelligent teller machines, automated risk control terminals, or financial transaction venues. The specific process includes the following steps:

[0155] First, multiple fixed surveillance cameras, mobile inspection equipment, or intelligent robotic devices deployed at financial outlets or trading centers capture image sequences covering the location. These sequences present spatial environmental information from different perspectives and locations. Each frame in the sequence captures visual information such as the distribution of equipment within the outlet, personnel activities, counter locations, and blind spots, ensuring diverse data sources and a complete representation of the spatial structure.

[0156] Combined with the image acquisition process, the system synchronously obtains the camera pose information corresponding to each frame of image. The spatial position expresses the specific position of the camera in the grid through three-dimensional coordinates, and the spatial orientation uses Euler angles or quaternions to represent the camera orientation direction, which fully reflects the shooting position and angle of each frame of image, facilitating subsequent spatial structure reconstruction.

[0157] Based on the acquired image sequences, the system leverages a built-in visual analysis model to extract visual language features from each frame. These features encompass information such as business windows, ticket collection devices, personnel flow paths, security monitoring areas, and risk monitoring signs within the financial branch. Simultaneously, the system calculates the spatial features of each frame based on the camera's position and quantifies the specific position and spatial distribution of each image in three-dimensional space.

[0158] Furthermore, the system calculates the semantic similarity between any two images in the image sequence based on visual language features. By comparing image content, identifying financial device labels, and analyzing business area distribution, the system determines the similarity in the two images' expressive content. Incorporating spatial features, the system calculates the spatial distance between any two images, reflecting their physical location relationship and spatial coverage.

[0159] The system fuses semantic similarity and spatial distance to generate a joint distance metric. This metric is used to evaluate the information redundancy and complementarity of each frame in the image sequence, selecting a subset of keyframes that are rich in information, broad in coverage, and complete in expression. Simultaneously, the system extracts the camera pose information corresponding to each keyframe subset, ensuring that each keyframe has accurate spatial position and orientation data.

[0160] During the keyframe screening process, the system calculates the edge sharpness value of each frame, evaluates the image clarity, and generates a clarity weight coefficient based on the clarity index, prioritizing frames with rich image details and clear information expression. Combined with the camera pose information, the field of view coverage of each frame is calculated, and field of view priority weights are assigned based on the coverage area, increasing the weight of frames with strong spatial complementarity and unique perspectives in the screening process. The joint distance metric with clarity weight and field of view weight is combined to generate a weighted joint distance metric. The system uses this metric to screen a subset of keyframes and extract the corresponding camera pose information.

[0161] The system receives spatial reasoning requests from financial business scenarios. These requests can include branch layout optimization, blind spot identification, ATM location relationship analysis, pedestrian path detection, or surveillance blind spot inspection. For key frames, the system generates a scene guide describing the image acquisition method and spatial characteristics. This guide combines the camera position information with the orientation information converted from Euler angles to form a pose description. The image and pose description are then combined to form a graphic block.

[0162] The system generates annotation instructions that restrict the model's use of image sequence numbers to reference, preventing the model from relying on image sequence numbers during reasoning, which can affect inference accuracy. By integrating spatial reasoning requests, scene guidance, image blocks, and annotation instructions, the system constructs a complete multimodal prompt data set.

[0163] Multimodal prompts are fed into a pre-trained multimodal language model. The model interprets the camera pose information contained in the prompts and establishes a three-dimensional spatial coordinate system based on the results, accurately recreating the spatial structure of the venue. The model maps the keyframe content into three-dimensional space using this spatial coordinate system, generating a spatial mapping that clearly depicts the layout of the financial business area, the relative positions of ATMs, the movement paths of personnel, and the distribution of equipment.

[0164] Based on the spatial reasoning request and spatial mapping content, the model performs spatial relationship recognition, determining the spatial location relationships, relative orientations, and layout characteristics between different financial business facilities, and generating a spatial relationship description. Ultimately, based on this spatial relationship description, the model outputs spatial reasoning results in natural language. These results may include: specifying the relative spatial position of a self-service terminal device and the main business window within the branch; describing the spatial relationship between high-frequency personnel activity paths and blind spots; providing feedback on the rationality of the layout of ATM deployment locations and customer waiting areas; and providing spatial layout optimization suggestions and risk area warnings.

[0165] In healthcare scenarios, multimodal spatial reasoning methods can be used to understand the spatial structure, analyze equipment layout, and reason about the relationships between personnel activities within medical facilities, such as hospital wards, operating rooms, emergency rooms, or smart medical terminals. The specific process includes the following operations:

[0166] First, multiple fixed cameras, bedside monitoring devices, mobile nursing robots, or medical inspection terminals deployed within the medical environment capture image sequences covering the hospital interior. These image sequences record multi-dimensional spatial information, including ward bed layout, medical equipment distribution, personnel movement paths, and patient locations. Each image frame reflects the medical scene from a specific spatial perspective, allowing the system to obtain complete spatial visual information.

[0167] In conjunction with the image acquisition process, the system synchronously obtains the camera pose information of each frame of image. The spatial position expresses the specific position of the camera in the ward or medical area through three-dimensional coordinates, and the spatial orientation uses Euler angles or other angle parameters to represent the camera shooting direction, comprehensively reflecting the spatial position and angle characteristics of image acquisition.

[0168] For image sequences, the system applies a built-in visual analysis model to extract visual language features from each frame. These features include medical device labels, bed numbers, nursing workspaces, monitoring system status, and personnel identities, reflecting the semantic information and visual structure of the hospital space. Simultaneously, combined with camera pose information, the system calculates the spatial features of each frame, clarifying the image's positional relationship and spatial distribution within the overall medical facility.

[0169] Furthermore, the system calculates the semantic similarity between any two images based on visual language features. By comparing the medical device information, patient position changes, and scene structure within the images, it identifies similarities in content expression between different images. Incorporating spatial features, the system calculates the spatial distance between each frame in the image sequence, reflecting the actual spatial position differences and spatial layout information of the corresponding images.

[0170] The system integrates semantic similarity and spatial distance to generate a joint distance metric, which is used to evaluate image information redundancy and spatial complementarity, screen out a subset of key frames with clear information expression and reasonable spatial distribution, and simultaneously extract the camera pose information corresponding to the key frames to ensure that the spatial reasoning process has accurate position and orientation data.

[0171] During the keyframe screening process, the system calculates the edge sharpness of each frame, reflecting the detail of the medical scene through image clarity. Combined with this clarity index, it generates a clarity weight coefficient, prioritizing high-quality images in the screening process. Combined with camera pose information, the system calculates the field of view coverage of each frame and assigns a field of view priority weight based on the coverage area, prioritizing frames with rich spatial information and minimal blind spots, thereby enhancing the comprehensiveness and diversity of spatial expression.

[0172] The system fuses the joint distance metric with clarity weight and the joint distance metric with field of view weight to generate a weighted joint distance metric. Based on this metric, it filters the key frame subset and extracts the corresponding camera pose information.

[0173] In medical scenarios, the system receives spatial reasoning requests, which can include analyzing the relative layout of medical equipment, determining the relationship between beds and nursing corridors, planning emergency cart paths, identifying blind spots in medical monitoring, or evaluating ward space utilization. For selected keyframes, the system generates scene guidance describing the image acquisition method and spatial characteristics. This guidance, combined with camera position and orientation information, creates a pose description text, which is then combined with the image to create a graphic block.

[0174] The system generates annotation instructions that restrict the model's reference image numbers, integrates spatial reasoning requests, scene guides, image blocks and annotation instructions, and constructs a complete multimodal prompt.

[0175] Multimodal prompts are fed into a pre-trained multimodal language model, which interprets the camera's pose information and establishes a three-dimensional spatial coordinate system for the medical facility, accurately recreating the spatial structure of the ward, nursing station, or surgical area. The model, combined with the spatial coordinate system, maps keyframe content into three-dimensional space, generating a spatial map that intuitively displays the distribution of medical equipment, bed arrangement, personnel movement paths, and monitoring range.

[0176] Based on spatial reasoning requests and spatial mapping content, the model performs spatial relationship recognition, analyzing the spatial positional relationships between equipment, the relative orientation of beds, personnel flow paths, and the layout of safety corridors, generating a spatial relationship description. Ultimately, the model outputs spatial reasoning results in natural language, including: describing the spatial positional relationship between specific medical equipment and nursing corridors; specifying the layout relationship between patient beds and vital signs monitoring systems; identifying spatial blind spots and blind spots, providing optimization suggestions; and providing feedback on the spatial conflict risk between equipment storage areas and emergency evacuation routes.

[0177] This embodiment avoids the problem of relying on external three-dimensional modeling, structural diagram construction or independent spatial reasoning modules in traditional methods through the spatial information parsing capability and spatial reasoning mechanism of the multimodal language model, and realizes an integrated reasoning process of image, spatial information and language expression. Combining a unified three-dimensional spatial coordinate system and spatial mapping content, it ensures the accurate transmission of spatial structural information and the clear expression of positional relationships, and improves the structural integrity and expression accuracy of spatial reasoning results. Through multimodal prompt input, the difficulty of data integration is reduced, and complex data conversion and format compatibility issues are avoided. The final output of natural language formal reasoning results has good readability and user-friendliness, which is convenient for flexible application in different business scenarios, and overall improves the accuracy, generalization ability and practicality of the spatial reasoning system.

[0178] In one embodiment, a spatial reasoning device based on an image sequence is provided, and the spatial reasoning device based on an image sequence corresponds one-to-one to the spatial reasoning method based on an image sequence in the above embodiment. Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the image sequence-based spatial reasoning device of the present invention. It includes a data acquisition module 10, a key frame screening module 20, a prompt construction module 30, and a spatial reasoning module 40. Each functional module is described in detail below:

[0179] The data acquisition module 10 is used to obtain the image sequence and the camera pose information corresponding to each frame of the image sequence;

[0180] A key frame screening module 20 is configured to screen out key frames and camera pose information corresponding to the key frames based on visual language features and spatial features according to the image sequence and the camera pose information;

[0181] a prompt construction module 30 for constructing a multimodal prompt based on the keyframes, the camera pose information corresponding to the keyframes, and the spatial reasoning request;

[0182] The spatial reasoning module 40 is configured to input the multimodal prompt into a pre-trained multimodal language model to generate a spatial reasoning result in response to the spatial reasoning request.

[0183] In one embodiment, the data acquisition module 10 is specifically configured to:

[0184] Acquire image sequences of multi-view scenes;

[0185] Obtaining camera position information of each frame of image in the image sequence;

[0186] Obtaining camera orientation information of each frame of the image sequence;

[0187] Based on the camera position information and the camera orientation information, camera pose information corresponding to each frame of image is generated.

[0188] In one embodiment, the key frame screening module 20 is specifically configured to:

[0189] Extracting visual language features of each frame of the image sequence;

[0190] Generate spatial features of each frame of image based on the camera pose information;

[0191] Determining the semantic similarity between every two frames of the image sequence based on the visual language features;

[0192] Determining a spatial distance between every two frames of image in the image sequence based on the spatial feature;

[0193] fusing the semantic similarity and spatial distance to generate a joint distance metric between every two frames of images;

[0194] A key frame subset is filtered from the image sequence according to the joint distance metric value, and camera pose information corresponding to the key frame subset is extracted from the camera pose information.

[0195] In one embodiment, the key frame screening module 20 is specifically configured to:

[0196] Determining an edge sharpness value of each frame of the image in the image sequence as a clarity index;

[0197] Determine the field of view angle coverage of each frame image according to the camera pose information;

[0198] Based on the clarity index, adding a clarity weight coefficient to the joint distance metric value to generate a joint distance metric value with clarity weight;

[0199] Based on the field of view angle coverage, adding a field of view angle priority weight to frames with complementary perspectives to generate a joint distance metric value with the field of view angle weight;

[0200] fusing the joint distance metric value with the clarity weight and the joint distance metric value with the field of view angle weight to obtain a weighted joint distance metric value;

[0201] Filtering a subset of key frames from the image sequence according to the weighted joint distance metric;

[0202] The camera pose information corresponding to the key frame subset is extracted from the camera pose information.

[0203] In one embodiment, the prompt building module 30 is specifically configured to:

[0204] receiving spatial reasoning requests;

[0205] Generating scene guides for describing image acquisition methods and spatial characteristics for the key frames;

[0206] Combining each frame image in the key frame and the corresponding camera pose information into a picture block;

[0207] Generate annotation instructions for restricting the model to use image number references;

[0208] The spatial reasoning request, scene guide, graphic blocks and annotation instructions are integrated to construct a multimodal prompt.

[0209] In one embodiment, the prompt building module 30 is specifically configured to:

[0210] Extracting camera position information and camera orientation information from the camera pose information;

[0211] Convert the camera orientation information into Euler angle representation;

[0212] Combining the camera position information and Euler angle representation to generate a posture description text;

[0213] The posture description text is combined with the image in the corresponding key frame to generate an image-text block.

[0214] In one embodiment, the spatial reasoning module 40 is specifically configured to:

[0215] parsing camera pose information in the multimodal prompt using a multimodal language model;

[0216] Establishing a three-dimensional space coordinate system based on the camera pose information;

[0217] Mapping key frame content in the multimodal prompt to the three-dimensional space coordinate system through a multimodal language model to generate spatial mapping content;

[0218] Performing spatial relationship recognition processing based on the spatial reasoning request and spatial mapping content using the multimodal language model to generate a spatial relationship description;

[0219] A spatial reasoning result in natural language is generated based on the spatial relationship description.

[0220] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide determination and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external user terminal via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a spatial reasoning method based on an image sequence.

[0221] In one embodiment, a computer device is provided. The computer device may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide determination and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the user side of a spatial reasoning method based on an image sequence.

[0222] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0223] Obtaining an image sequence and camera pose information corresponding to each frame of the image sequence;

[0224] According to the image sequence and the camera pose information, key frames and camera pose information corresponding to the key frames are screened based on visual language features and spatial features;

[0225] Constructing a multimodal prompt based on the keyframes, the camera pose information corresponding to the keyframes, and the spatial reasoning request;

[0226] The multimodal prompt is input into a pre-trained multimodal language model to generate a spatial reasoning result in response to the spatial reasoning request.

[0227] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0228] Obtaining an image sequence and camera pose information corresponding to each frame of the image sequence;

[0229] According to the image sequence and the camera pose information, key frames and camera pose information corresponding to the key frames are screened based on visual language features and spatial features;

[0230] Constructing a multimodal prompt based on the keyframes, the camera pose information corresponding to the keyframes, and the spatial reasoning request;

[0231] The multimodal prompt is input into a pre-trained multimodal language model to generate a spatial reasoning result in response to the spatial reasoning request.

[0232] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the user side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0233] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0234] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0235] It should be noted that if any software tools or components other than those of the Company appear in the embodiments of this application, they are merely for illustration and do not represent actual use. The above embodiments are intended only to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some of the technical features therein with equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A spatial reasoning method based on image sequences, characterized in that: The following steps are involved: Obtaining an image sequence and camera pose information corresponding to each frame of the image sequence; According to the image sequence and the camera pose information, key frames and camera pose information corresponding to the key frames are screened based on visual language features and spatial features; Constructing a multimodal prompt based on the keyframes, the camera pose information corresponding to the keyframes, and the spatial reasoning request; The multimodal prompt is input into a pre-trained multimodal language model to generate a spatial reasoning result in response to the spatial reasoning request.

2. The spatial reasoning method based on image sequences according to claim 1, characterized in that: Obtaining an image sequence and camera pose information corresponding to each frame in the image sequence, including: Acquire image sequences of multi-view scenes; Obtaining camera position information of each frame of image in the image sequence; Obtaining camera orientation information of each frame of the image sequence; Based on the camera position information and the camera orientation information, camera pose information corresponding to each frame of image is generated.

3. The spatial reasoning method based on image sequences according to claim 1, characterized in that: According to the image sequence and the camera pose information, key frames and camera pose information corresponding to the key frames are screened based on visual language features and spatial features, including: Extracting visual language features of each frame of the image sequence; Generate spatial features of each frame of image based on the camera pose information; Determining the semantic similarity between every two frames of the image sequence based on the visual language features; Determining a spatial distance between every two frames of image in the image sequence based on the spatial feature; fusing the semantic similarity and spatial distance to generate a joint distance metric between every two frames of images; A key frame subset is filtered from the image sequence according to the joint distance metric value, and camera pose information corresponding to the key frame subset is extracted from the camera pose information.

4. The spatial reasoning method based on image sequences according to claim 3, characterized in that: According to the joint distance metric value, a key frame subset is selected from the image sequence, and camera pose information corresponding to the key frame subset is extracted from the camera pose information, including: Determining an edge sharpness value of each frame of the image in the image sequence as a clarity index; Determine the field of view angle coverage of each frame image according to the camera pose information; Based on the clarity index, adding a clarity weight coefficient to the joint distance metric value to generate a joint distance metric value with clarity weight; Based on the field of view angle coverage, adding a field of view angle priority weight to frames with complementary perspectives to generate a joint distance metric value with the field of view angle weight; fusing the joint distance metric value with the clarity weight and the joint distance metric value with the field of view angle weight to obtain a weighted joint distance metric value; Filtering a subset of key frames from the image sequence according to the weighted joint distance metric; The camera pose information corresponding to the key frame subset is extracted from the camera pose information.

5. The spatial reasoning method based on image sequences according to claim 1, characterized in that: Constructing a multimodal prompt based on the keyframe, the camera pose information corresponding to the keyframe, and the spatial reasoning request, including: receiving spatial reasoning requests; Generating scene guides for describing image acquisition methods and spatial characteristics for the key frames; Combining each frame image in the key frame and the corresponding camera pose information into a picture block; Generate annotation instructions for restricting the model to use image number references; The spatial reasoning request, scene guide, graphic blocks and annotation instructions are integrated to construct a multimodal prompt.

6. The spatial reasoning method based on image sequences according to claim 5, characterized in that: Each frame image in the key frame and the corresponding camera pose information are combined into a graphic block, including: Extracting camera position information and camera orientation information from the camera pose information; Convert the camera orientation information into Euler angle representation; Combining the camera position information and Euler angle representation to generate a posture description text; The posture description text is combined with the image in the corresponding key frame to generate an image-text block.

7. The spatial reasoning method based on image sequences according to claim 1, characterized in that: Inputting the multimodal prompt into a pre-trained multimodal language model to generate a spatial reasoning result in response to the spatial reasoning request, including: parsing camera pose information in the multimodal prompt using a multimodal language model; Establishing a three-dimensional space coordinate system based on the camera pose information; Mapping key frame content in the multimodal prompt to the three-dimensional space coordinate system through a multimodal language model to generate spatial mapping content; Performing spatial relationship recognition processing based on the spatial reasoning request and spatial mapping content using the multimodal language model to generate a spatial relationship description; A spatial reasoning result in natural language is generated based on the spatial relationship description.

8. A spatial reasoning device based on image sequences, characterized in that: The spatial reasoning device based on image sequences comprises: A data acquisition module is used to obtain an image sequence and camera pose information corresponding to each frame of the image sequence; A key frame screening module is used to screen out key frames and camera pose information corresponding to the key frames based on visual language features and spatial features according to the image sequence and the camera pose information; A prompt construction module, configured to construct a multimodal prompt based on the keyframes, the camera pose information corresponding to the keyframes, and the spatial reasoning request; The spatial reasoning module is configured to input the multimodal prompt into a pre-trained multimodal language model and generate a spatial reasoning result in response to the spatial reasoning request.

9. A computer device, characterized in that: The computer device includes a memory, a processor, and an image sequence-based spatial reasoning program stored in the memory and capable of running on the processor. When the image sequence-based spatial reasoning program is executed by the processor, the steps of the image sequence-based spatial reasoning method as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The storage medium stores a spatial reasoning program based on an image sequence, and when the spatial reasoning program based on an image sequence is executed by a processor, the steps of the spatial reasoning method based on an image sequence as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • 3D form and attitude estimation method and device based on space-time correlation image

    CN113298047A

  • Camera pose estimation method and system based on semantics

    CN114708321A