Method for constructing large model for artificial general spatiotemporal intelligence

The AGSTI model addresses LLMs' limitations by employing multidimensional object abstraction and spatiotemporal encoding to process and train models, enabling accurate spatiotemporal data handling and improving decision-making in various fields.

GB2636309APending Publication Date: 2025-06-11BEIJING LONGRUAN TECHNOLOGIES INC +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
GB2025002709
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-10-15
Filing Date
2025-02-25
Publication Date
2025-06-11

AI Technical Summary

Technical Problem

Large language models (LLMs) face limitations in processing multidimensional spatiotemporal information due to their one-dimensional data structure, lacking spatial and temporal interaction capabilities, which hinders their application in fields requiring accurate spatiotemporal data processing, such as autonomous driving and robotics.

Method used

A method for constructing a large model for artificial general spatiotemporal intelligence (AGSTI) that processes objects and spatial relationships using multidimensional object abstraction, employs a multimodal spatiotemporal encoding algorithm, and trains the model with a three-dimensional representation-based multimodal spatiotemporal training method to handle spatiotemporal data effectively.

Benefits of technology

The AGSTI model can uniformly represent and process spatiotemporal data, enhancing decision-making accuracy and adaptability in complex scenarios, and is applicable in fields like multimedia information processing, autonomous driving, intelligent robotics, and industrial engineering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Disclosed is a method for constructing a large model for artificial general spatiotemporal intelligence (AGSTI) that includes establishing a unified and structured topological representation and stora
Need to check novelty before this filing date? Find Prior Art

Description

FIELD

[0001] The present disclosure relates to the technical field of artificial intelligence, and particularly, to a method for constructing a large model for artificial general spatiotemporal intelligence (AGSTI). BACKGROUND

[0002] In recent years, artificial intelligence (AI) technology has developed rapidly. Large language models (LLMs) have achieved significant breakthroughs in fields of natural language processing, speech recognition, computer vision, and the like by virtue of their remarkable representation learning and generalization capabilities. LLMs have become a prominent focus of AI research and development, finding widespread applications in industries such as literature, art, education, and finance, and exerting a profound impact on the development of human society.

[0003] Currently, the number of parameters of LLMs typically reaches billions or even hundreds of billions, placing extremely high demands on computing resources and training technology. Training such a large-scale model requires large volumes of data. The process of acquiring, cleaning, and annotating these data is expensive and time-consuming. This creates a bottleneck in computing and data resources.

[0004] Moreover, LLMs are inherently designed to process language-based text data. While LLMs can encode multimedia data such as images and videos, their fundamental data structure remains onedimensional, represented by tokenized text sequences. LLMs have relatively simple structures and are lack of capability to represent spatial relationships. Additionally, as LLMs are primarily trained based on text data, they lack direct interaction experience with the physical world. Unlike humans, LLMs cannot leam spatial, temporal, and physical interactions through personal experience. Consequently, LLMs struggle to effectively integrate multidimensional spatiotemporal information from the real world, and cannot support the perception, understanding, and operation of three-dimensional spaces. This limitation prevents their direct application in fields requiring multidimensional spatiotemporal information interaction, such as industrial production and engineering construction. For example, in autonomous driving and robotics, accurate modeling and prediction for spatiotemporal changes of each object in the environment are essential. However, a processing method of LLMs, based on one-dimensional data structures, requires extremely high computational power and lacks a unified analytical framework for temporal and spatial dimensions. As a result, LLMs fail to provide accurate and consistent spatiotemporal association analyses.

[0005] Currently, LLMs are primarily suited for intelligent processing in the humanities which are referred to as "humanities intelligence" and cover text, images, and videos, etc. LLMs are not well-suited for processing scientific and engineering data, particularly not suited for processing, interpreting or analyzing spatiotemporal data in human social engineering fields, such as the intelligent processing of multidimensional engineering design graphics. Graphic design and application in the field of engineering are critical issues that must be addressed in an industrial and intelligent society. Otherwise, the application scenarios of artificial intelligence will remain localized and incomplete.

[0006] The human real world consists of three-dimensional physical objects or multimedia information objects, both of which have temporal and spatial features. The three-dimensional world follows physical laws and possesses inherent structures, attributes, and spatiotemporal relationships. Due to the innate limitations of LLMs with their one-dimensional data structure and data modeling methods, there is a need to find a unified and universal method for expressing and processing spatiotemporal data, and to develop a spatiotemporal large model suitable for representation, perception, reasoning, and interactive decision-making in the real spatial world. Such large models are capable of comprehensively processing multimedia data faced by humanity, such as "text, graphics, images, audio, and video", and addressing problems across all professional domains, including the "humanities, sciences, and engineering". SUMMARY

[0007] In view of the above problems, the present disclosure proposes a method for constructing a large model for artificial general spatiotemporal intelligence (AGSTI)

[0008] The present disclosure discloses a method for constructing a large model for artificial general spatiotemporal intelligence (AGSTI), including:

[0009] processing each object and a spatial relationship of the object in the real world using a multidimensional object abstraction representation manner, and establishing a structured spatiotemporal object representation that can be uniformly processed by a spatiotemporal intelligence and data storage, wherein multidimensional object abstraction refers to zero-dimensional objects, one-dimensional objects, two-dimensional objects, and three-dimensional objects;

[0010] processing a spatiotemporal object by a multimodal spatiotemporal encoding algorithm, to transform spatiotemporal data corresponding to the spatiotemporal object into multimodal spatiotemporal encoded data for subsequent training and application of the spatiotemporal intelligence large model;

[0011] constructing a multimodal spatiotemporal training dataset based on the multimodal spatiotemporal encoded data, and training the spatiotemporal intelligence large model by adopting a three-dimensional representation-based multimodal spatiotemporal training method; and

[0012] deploying the trained spatiotemporal intelligence as a spatiotemporal intelligence inference service to provide spatiotemporal intelligence capability, the spatiotemporal intelligence capability includes: spatiotemporal relationship representation, semantic understanding, environmental perception, spatiotemporal reasoning, and spatial decision-making.

[0013] Optionally, the processing each object and a spatial relationship of the object in the real world using a multidimensional object abstraction representation manner includes:

[0014] collecting data of various objects in the real world, wherein the data includes, but is not limited to, multimedia information formats, identifying and analyzing spatiotemporal information features of various objects within the multimedia information formats, and combining the environment, scene, or context in which the objects are located, their relative spatial positions, as well as knowledge graph associations and spatial relationships of the objects, to generate multimodal and structured spatiotemporal object representations in the form of zero-dimensional objects, onedimensional objects, two-dimensional objects, and three-dimensional volumes. The multimedia information formats include: text format, graphic format, image format, audio format, and video format, and the spatiotemporal information features include: location, form, state, and temporality; and

[0015] the spatiotemporal object representation includes features of location, temporality, and spatial relationships to ensure consistency and association of objects in temporal and spatial dimensions, the relationships between the spatiotemporal objects are represented through topological relationships, directional or sequential relationships, and metric relationships, the topological relationships include: association, adjacency, and containment, the directional or sequential relationships include: up, down, left, right, and north, south, east, west, and the metric relationships include: distance, angle, or proximity.

[0016] Optionally, the objects in the real world include: real-world physical geographic objects, virtual geographic objects, and multimedia objects in inherent relative space of multimedia contexts, and the multimedia objects include: text, graphics, images, audio, and videos;

[0017] the zero-dimensional object refers to the object-based modeling representation of fundamental units in the real world; the fundamental units refer to various physical objects with spatial locations in the real world, as well as multimedia information objects which include: text, graphics, images, audio, and videos; the representation of the physical objects is based on spatial locations, environment, and spatial relationships of the objects, the representation of the multimedia information objects is based on digital features, scene or context, and relative space of the objects, including but not limited to a sequential location of text in a paragraph and relevance to other text, graphic vector nodes, as well as sequential locations of image frames in a video and relevance to other frames;

[0018] the one-dimensional object refers to the modeling representation of changes and evolution processes of real-world objects in the temporal dimension, including: representing a continuous change process or trajectory of an object in temporal and spatial dimensions as a one-dimensional object, or representing data sequences of zero-dimensional objects or multiple similar zerodimensional objects as a one-dimensional object, and recording a flow process, a logical relationship, and a topological adjacency relationship of the object in temporal and spatial dimensions, wherein the data sequences include but are not limited to: coordinate sequences, character sequences, text paragraphs, mathematical formulas, graphic vector segments, image sets, and video frames;

[0019] the two-dimensional object refers to the modeling representation of adjacency and aggregation of real-world objects, including: representing similar one-dimensional objects as a two-dimensional object, and recording adjacency relationships of the objects in temporal and spatial dimensions for larger-scale data analysis and association modeling; and when processing adjacency and aggregation of real-world objects, fully recording intrinsic features of the fundamental units, as well as complex derivation processes or combined association features of multiple data based on the locations, types, or changing processes of the objects in temporal and spatial dimensions; and

[0020] the three-dimensional object refers to the modeling representation of large-scale datasets or complex system overall spatiotemporal evolution of real-world objects, including: representing big datasets of zero-dimensional objects, one-dimensional objects, and two-dimensional objects as composites or composite processes, recording overall relationships or evolution processes of the objects in temporal and spatial dimensions, generating an overall spatiotemporal body structure based on objects and events of multiple spatiotemporal dimensions, and supporting comprehensive spatiotemporal analysis and prediction of complex systems.

[0021] Optionally, employing a multimodal spatiotemporal encoding algorithm to process the spatiotemporal data corresponding to the spatiotemporal objects, converting said spatiotemporal data into encoded data for subsequent training and application of the spatiotemporal intelligence model, comprising:

[0022] step Tl: based on the spatiotemporal objects, taking each fundamental object in the spatiotemporal data as a node in a spatiotemporal topology graph and a spatiotemporal adjacency between the nodes as an edge of the graph, defining weights of the edges between the nodes by calculating a spatiotemporal distance or association between the objects, and constructing a spatiotemporal topological graph structure, wherein the association between the objects is determined based on similarity, proximity, or temporal sequence in the spatiotemporal data, and the spatiotemporal data refers to multimodal spatiotemporal multidimensional data; and

[0023] step T2: performing data encoding based on a graph neural network, constructing a pretraining task through spatial association, and generating high-dimensional representation of the spatiotemporal topological graph structure.

[0024] Optionally, the performing data encoding based on a graph neural network, constructing a pre-training task through spatial association, and generating high-dimensional representation of the spatiotemporal topological graph structure includes:

[0025] by aggregating information from the nodes of the spatiotemporal topology graph and adjacent nodes, updating spatiotemporal features of each node to capture spatial and temporal associations between the multimodal spatiotemporal multidimensional data; and

[0026] performing deep learning on the multimodal spatiotemporal multidimensional data by using a multi-layer graph neural network to generate encoded representations with enhanced spatiotemporal associations for further model training.

[0027] Optionally, the training or fine-tuning the spatiotemporal intelligence by adopting a three-dimensional representation-based multimodal spatiotemporal training method includes:

[0028] step Nl: based on the multimodal spatiotemporal encoded data, collecting or generating data containing spatiotemporal information according to a specific application scenario and task requirements, and constructing the multimodal spatiotemporal training dataset, wherein the training dataset includes tag, and temporal and spatial attribute information of multimodal data, and is used for guiding adaptability and accuracy of the spatiotemporal intelligence in a specific task, so as to train a spatiotemporal association capability of the spatiotemporal intelligence; and

[0029] step N2: training the spatiotemporal intelligence by taking the constructed multimodal spatiotemporal training dataset as an input, and using a multidimensional spatiotemporal attention network model and the multimodal spatiotemporal training method, to capture temporal sequence relationships and spatial features annotated with zero-dimensional objects, one-dimensional objects, two-dimensional objects, and three-dimensional objects in the data.

[0030] Optionally, the multimodal spatiotemporal training method refers to multimodal spatiotemporal training combining three-dimensional representations of the zero-dimensional object, the one-dimensional object, the two-dimensional object, and the three-dimensional object, as well as the temporal dimension, an underlying method of the spatiotemporal intelligence adopts a three-dimensional representation method, and modeling is performed based on three-dimensional spatial representation and the temporal dimension, as well as the temporal sequence and the spatial features of the multimodal spatiotemporal multidimensional data, thereby accurately capturing a complex relationship in the temporal and spatial dimensions, reflecting a spatiotemporal dynamic change of the real world, and enhancing the performance of the spatiotemporal intelligence in spatiotemporal reasoning.

[0031] Optionally, during construction of the multimodal spatiotemporal training dataset, the focus is on acquiring data diversity and coverage, to ensure robustness and a generalization capability of the spatiotemporal intelligence in different contexts, wherein text data includes a plurality of types of content, including but not limited to, a natural language description, a command, and a dialogue, covering a broad semantic range and different degrees of complexity; image data includes, but is not limited to, static and dynamic images from different perspectives, in different environments, and in different time periods; and geospatial data includes, but is not limited to, formatted spatial data from different industries and with different precisions; and spatial relationships related to the spatiotemporal data are constructed.

[0032] Optionally, the spatiotemporal intelligence capability provided by the spatiotemporal intelligence reasoning service is based on a spatial relationship represented by the spatiotemporal objects, during reasoning of the spatiotemporal intelligence, the topological relationships between the objects recorded by the spatiotemporal object are used and outputted, including a knowledge graph and a service process logic related to the object, to reduce the computing complexity in reasoning of the spatiotemporal intelligence and efficiently complete a reasoning task.

[0033] Optionally, the spatiotemporal relationship representation includes a topological spatial relationship, a sequential spatial relationship, and a metric relationship, as well as a time mark or timestamp, the topological spatial relationship includes: association, adjacency, and containment, the sequential spatial relationship includes: orientation, front, rear, left, and right, and the metric relationship includes: length and angle;

[0034] the capability of semantic understanding involves analyzing and identifying a spatiotemporal object according to visual information processing or a digital input, and outputting a spatiotemporal object with zero-dimensional object, one-dimensional object, two-dimensional object, and three-dimensional object features, as well as spatial relationships between objects and between an object and an environment;

[0035] the capability of environmental perception involves understanding and predicting a relative relationship between spatiotemporal objects and a change trend based on the spatiotemporal object representation, to provide a support for further application of spatiotemporal relationship reasoning;

[0036] the capability of spatiotemporal reasoning involves understanding the spatial and temporal environments of objects in the real world, or the context scene and virtual environments of multimedia information based on the capability of environmental perception, reasoning and analyzing an association relationship between the objects, and generating a spatial object feature with a future temporal feature and an association relationship trend; and

[0037] the capability of spatial decision-making involves, based on semantic understanding, environmental perception, and spatiotemporal reasoning, achieving application of spatial decisionmaking in spatiotemporal scenarios through spatiotemporal reasoning and analysis or service process learning according to specific spatiotemporal scenarios and service processes.

[0038] According to the method for constructing an AGSTI of the present disclosure, in a first step, processing each object and a spatial relationship of the object in the real world using a multidimensional object abstraction representation manner, and then establishing a structured spatiotemporal object representation that can be uniformly processed by a spatiotemporal intelligence and data storage; then, processing a spatiotemporal object by a multimodal spatiotemporal encoding algorithm, and transforming spatiotemporal data corresponding to the spatiotemporal object into multimodal spatiotemporal encoded data for subsequent training and application of the spatiotemporal intelligence; constructing a multimodal spatiotemporal training dataset based on the multimodal spatiotemporal encoded data, and training or fine-tuning the spatiotemporal intelligence by adopting a three-dimensional representation-based multimodal spatiotemporal training method; and finally, deploying the trained or fine-tuned spatiotemporal intelligence as a spatiotemporal intelligence reasoning service to provide spatiotemporal intelligence capabilities.

[0039] According to the method for constructing an AGSTI proposed in the present disclosure, based on object representations of spatial zero-dimensional objects, one-dimensional objects, two-dimensional objects, and three-dimensional objects and a spatial relationship model, a unified and structured topological representation and storage of multimedia information, including text, graphics, images, audio, videos, and the like, in the real world of human society are established, a three-dimensional representation-based multimodal large model training method is employed to construct an AGSTI that can uniformly represent and process all objects in human society, and provide capabilities such as spatiotemporal relationship representation, semantic understanding, environmental perception, spatiotemporal reasoning, and spatial decision-making by combining spatiotemporal information processing, spatiotemporal relationship analysis, knowledge graphs, and service process logic reasoning, thereby solving the problems of the existing large language models, such as single-representation nature, lack of topological relationships, and difficulty handling three-dimensional spatiotemporal relationships in the real world. The AGSTI constructed based on the present disclosure can significantly enhance decision-making accuracy and adaptability of intelligence systems in complex scenarios, and can be widely applied in fields such as multimedia information processing, theoretical model derivation, autonomous driving, intelligent robotics, and industrial engineering, and provide a new technological support for these application fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become apparent to those of ordinary skill in the art. The drawings are provided only for the purpose of illustrating the preferred embodiments and are not intended to limit the present disclosure. Furthermore, the same reference numerals are used to denote the same components throughout the drawings. In the drawings:

[0041] FIG. 1 is a flowchart of a method for constructing an artificial general spatiotemporal intelligence (AGSTI) according to an embodiment of the present disclosure; and

[0042] FIG. 2 is a schematic diagram of an exemplary abstraction representation manner for real-world objects according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0043] To make objectives, features, and advantages of the present application clearer, the present disclosure will further be described in detail below with reference to the drawings and the embodiments. The specific embodiments described herein are merely illustrative of the present disclosure, are not all embodiments but only part of embodiments of the present disclosure, and are not intended to limit the present disclosure.

[0044] The present disclosure proposes a method for constructing an artificial general spatiotemporal intelligence (AGSTI). FIG. 1 is a flowchart of a method for constructing an AGSTI. The method includes:

[0045] Step 101: Processing each object and a spatial relationship of the object in the real world using a multidimensional object abstraction representation manner, and establishing a structured spatiotemporal object representation that can be uniformly processed by a spatiotemporal intelligence and data storage, wherein multidimensional object abstraction refers to zero-dimensional objects, one-dimensional objects, two-dimensional objects, and three-dimensional objects.

[0046] The present disclosure creatively proposes a theory of multidimensional object abstraction, which includes: zero-dimensional objects, one-dimensional objects, two-dimensional objects, and three-dimensional objects. To better understand the technical solution of the present disclosure, the theory of the multidimensional object abstraction will first be explained in detail.

[0047] The multidimensional object abstraction proposed by the present disclosure is a process of analyzing and abstracting various real-world objects in human society, as well as data obtained through the cognition of the objects, such as text, graphics, images, audio, and videos, into structured zero-dimensional objects, one-dimensional objects, two-dimensional objects, and three-dimensional objects. By adopting a representation and training method for a three-dimensional representation model, the limitations of large language models based solely on text representation are avoided from the perspective of data model representation and training, and computational power required for model training and reasoning is optimized.

[0048] The objects in the real world include: real-world physical geographic objects, virtual geographic objects, and multimedia objects in inherent relative space of multimedia contexts, and the multimedia objects include: text, graphics, images, audio, and videos. To better understand the above theory, a description is made by using some specific examples.

[0049] (1) Zero-dimensional object: it refers to a fundamental unit of the real world, such as vector nodes, the digital "1", the letter "A", the Chinese character "dL" (Mao), physical entities such as "a tree, a rose, a person, a tiger, an image, a snapshot, auditory and olfactory information at a specific time point, a part, or a building component", and mathematical symbols such as "the multiplication symbol n, summation symbol £, or sine function symbol sin". In the AGSTI of the present disclosure, they all participate in the representation and training of the model in the form of zero-dimensional objects.

[0050] Based on the above examples, the zero-dimensional object can be understood as an objectbased modeling representation of the fundamental units of the real world. The fundamental units refer to various physical objects with spatial locations in the real world, as well as multimedia information objects. The multimedia information objects include: text, graphics, images, audio, videos, etc. The spatial location includes absolute coordinates and local coordinates of the object, as well as relative index and page number, and relative coordinates of the object in a scene or context. The physical objects are represented based on spatial locations, environment, and spatial relationships of the objects. The multimedia information objects are represented based on digital features, scene or context, and relative space of the objects, including but not limited to a sequential location of text in a paragraph, page number of the text in a chapter, and relevance to other text, as well as sequential locations of image frames in a video and relevance to other frames.

[0051] (2) One-dimensional object: it refers to vector arc segments, a sequence of numbers or characters, a phrase or a sentence, physical entities such as "a row of trees, a row of roses, movement or running of a row of people or an individual person, movement or running of a row of tigers or an individual tiger, a segment of video, a segment of auditory or olfactory information", subsystems formed by multiple parts or building components, local computational formulas, etc. In the AGSTI of the present disclosure, they all participate in the representation and training of the model in the form of one-dimensional objects.

[0052] Based on the above examples, the one-dimensional object can be understood as the modeling representation of changes and evolution processes of real-world objects in the temporal dimension, including: representing a continuous change process or trajectory of an object in temporal and spatial dimensions as a one-dimensional object, or representing data sequences of zerodimensional objects or multiple similar zero-dimensional objects as a one-dimensional object, and recording a flow process, a logical relationship, and a topological adjacency relationship of the object in temporal and spatial dimensions based on topological node features. The data sequences include but are not limited to: coordinate sequences, character sequences, text paragraphs, mathematical formulas, image sets, and video frames.

[0053] (3) Two-dimensional object: it refers to vector polygons, multiple or multiple segments of data or strings, physical entities such as "a forest, a cluster of roses, a crowd of people, a group of tigers, a complete narrative segment of a movie or a complete segment of television video, a staged auditory and olfactory information set", fully functional machines or buildings, a complete computing or derivation process (a combination of multiple formulas), etc.

[0054] Based on the above examples, the two-dimensional object can be understood as the modeling representation of adjacency and aggregation of real-world objects, including: representing multiple similar zero-dimensional objector one-dimensional objects as a two-dimensional object, and recording adjacency relationships of the objects in temporal and spatial dimensions for larger-scale data analysis and spatiotemporal object modeling; and when processing adjacency and aggregation of real-world objects, fully recording intrinsic features of the fundamental units, as well as complex derivation processes or combined association features of multiple data based on the locations, types, or changing processes of the objects in temporal and spatial dimensions.

[0055] (4) Three-dimensional object: it refers to vector polyhedrons, datasets or text sets, parks or communities, zoos, physical entities such as "a farm, forest, or grassland, a movie or television series", workshops or factories, building complexes, mathematical, physical, and chemical monographs, etc.

[0056] Based on the above examples, the three-dimensional object can be understood as the modeling representation of large-scale datasets or complex system overall spatiotemporal evolution of real-world objects, including: representing multidimensional big datasets of zero-dimensional objects, one-dimensional objects, and two-dimensional objects as composites or composite processes, and recording overall relationships or evolution processes in temporal and spatial dimensions. An overall spatiotemporal body structure is generated based on objects and events of multiple spatiotemporal dimensions, and is used for supporting comprehensive spatiotemporal analysis and prediction of complex systems.

[0057] The theory of multidimensional object abstraction can be better understood by referring to FIG. 2, which is a schematic diagram of an exemplary abstraction representation manner for real-world objects: physical objects in the real world (e.g., the digital "1", ..., a string of numbers, ..., multiple segments of data,..., datasets or text sets... on the left of FIG. 2) can, based on their different types, be in a unified spatiotemporal representation: zero-dimensional objects, one-dimensional objects, two-dimensional objects, and three-dimensional objects, which correspond to vector nodes, arcs, polygons, and polyhedrons, each possessing temporal attributes and a variety of relationships, including spatial, topological, sequential, and metric relationships.

[0058] The above provides a detailed explanation and elaboration of the theory of multidimensional object abstraction. According to the method for constructing an AGSTI proposed in the present disclosure, first an object and a spatial relationship of the object in the real world are processed in a multidimensional object abstraction representation manner, and then a structured spatiotemporal object representation that can be uniformly processed by a spatiotemporal intelligence and data storage are established. A preferred method includes:

[0059] first collecting data of various types of objects in the real world, wherein the data includes but is not limited to multimedia information formats; identifying and analyzing spatiotemporal information features of each object in the multimedia information formats; and generating multimodal and structured spatiotemporal object representations of zero-dimensional objects, one-dimensional objects, two-dimensional objects, and three-dimensional objects based on the environment, scene or context, and relative space in which each object is located, as well as knowledge graph association and spatial relationships of each object, wherein the multimedia information formats include: text, graphic, image, audio, and video formats, and the spatiotemporal information features include: location, form, state, and temporality.

[0060] The spatiotemporal object representation includes features of location, temporality, and topological relationships to ensure consistency and association of objects in temporal and spatial dimensions. The relationships between the spatiotemporal objects can be represented through topological relationships, directional or sequential relationships, and metric relationships. The topological relationships include: association, adjacency, and containment; the directional or sequential relationships include: up, down, left, right, and north, south, east, west; the metric relationships include: distance, angle, or proximity.

[0061] Step 102: Processing a spatiotemporal object by a multimodal spatiotemporal encoding algorithm, to transform spatiotemporal data corresponding to the spatiotemporal object into multimodal spatiotemporal encoded data for subsequent training and application of the spatiotemporal intelligence.

[0062] After the structured spatiotemporal object representation that can be uniformly processed by the spatiotemporal intelligence and the data storage are established, the spatiotemporal object is required to be processed by the multimodal spatiotemporal encoding algorithm, to transform the spatiotemporal data corresponding to the spatiotemporal object into the multimodal spatiotemporal encoded data for subsequent training and application of the spatiotemporal intelligence. A preferred method includes:

[0063] step Tl: based on the spatiotemporal objects, taking each fundamental object in the spatiotemporal data as a node in a spatiotemporal topology graph and a spatiotemporal adjacency between the nodes as an edge of the graph, defining weights of the edge between the nodes by calculating a spatiotemporal distance or association between the objects, and constructing a spatiotemporal topological graph structure, wherein the association between the objects is determined based on similarity, proximity, or temporal sequence in the spatiotemporal data, and the spatiotemporal data refers to multimodal spatiotemporal multidimensional data; and

[0064] step T2: performing data encoding based on a graph neural network, constructing a pretraining task through spatial association, and generating high-dimensional representation of the spatiotemporal topological graph structure.

[0065] For the high-dimensional representation, a preferred method includes: first, by aggregating information from the nodes of the spatiotemporal topology graph and adjacent nodes, updating spatiotemporal features of each node to capture spatial and temporal associations between the multimodal spatiotemporal multidimensional data; then, performing deep learning on the multimodal spatiotemporal multidimensional data by using a multi-layer graph neural network to generate encoded representations with enhanced spatiotemporal associations for further model training.

[0066] Step 103: Constructing a multimodal spatiotemporal training dataset based on the multimodal spatiotemporal encoded data, and training or fine-tuning the spatiotemporal intelligence by adopting a three-dimensional representation-based multimodal spatiotemporal training method.

[0067] After the spatiotemporal data is transformed into the multimodal spatiotemporal encoded data for subsequent training and application of the spatiotemporal intelligence, the multimodal spatiotemporal training dataset can be constructed based on the multimodal spatiotemporal encoded data, and the spatiotemporal intelligence is trained or fine-tuned based on the multimodal spatiotemporal training dataset by adopting the three-dimensional representation-based multimodal spatiotemporal training method.

[0068] A preferred method for training or fine-tuning the spatiotemporal intelligence includes:

[0069] step Nl: based on the multimodal spatiotemporal encoded data, collecting or generating data containing spatiotemporal information according to a specific application scenario and task requirements, and constructing the multimodal spatiotemporal training dataset, wherein the training dataset includes tag, and temporal and spatial attribute information of multimodal data, and is used for guiding adaptability and accuracy of the spatiotemporal intelligence in a specific task, so as to train a spatiotemporal association capability of the spatiotemporal intelligence; and

[0070] step N2: training the spatiotemporal intelligence by taking the constructed multimodal spatiotemporal training dataset as an input, and using a multidimensional spatiotemporal attention network model and the spatiotemporal multimodal training method, to capture temporal sequence relationships and spatial features annotated with zero-dimensional objects, one-dimensional objects, two-dimensional objects, and three-dimensional objects in the data.

[0071] The multimodal spatiotemporal training method in step N2 refers to spatiotemporal multimodal training combining three-dimensional representations of the zero-dimensional object, the one-dimensional object, the two-dimensional object, and the three-dimensional object, as well as the temporal dimension. This method differs from the conventional training method for language models or multimodal language models, where an underlying method for conventional model training is typically based on sequence lengths, attention mechanisms, and tokenized one-dimensional representations. The two methods represent entirely different training concepts and architectures.

[0072] An underlying method of the spatiotemporal intelligence proposed by the present disclosure adopts a three-dimensional representation method, and modeling is performed based on three-dimensional spatial representation and the temporal dimension, as well as the temporal sequence and the spatial features of the multimodal spatiotemporal multidimensional data, thereby accurately capturing a complex relationship in the temporal and spatial dimensions, reflecting a spatiotemporal dynamic change of the real world, and enhancing the performance of the spatiotemporal intelligence in spatiotemporal reasoning.

[0073] During construction of the multimodal spatiotemporal training dataset, the focus is on acquiring data diversity and coverage, to ensure robustness and a generalization capability of the spatiotemporal intelligence in different contexts, wherein text data includes a plurality of types of content, including but not limited to, a natural language description, a command, and a dialogue, covering a broad semantic range and different degrees of complexity; image data includes, but is not limited to, static and dynamic images from different perspectives, in different environments, and in different time periods; and geospatial data includes, but is not limited to, formatted spatial data from different industries and with different precisions; and spatial relationships related to the spatiotemporal data are constructed based on the geospatial data.

[0074] Step 104: Deploying the trained spatiotemporal intelligence as a spatiotemporal intelligence reasoning service to provide spatiotemporal intelligence capabilities, which include: spatiotemporal relationship representation, semantic understanding, environmental perception, spatiotemporal reasoning, and spatial decision-making.

[0075] After the spatiotemporal intelligence is trained or fine-tuned, that is, after the spatiotemporal intelligence is well trained, the trained spatiotemporal intelligence is deployed as a spatiotemporal intelligence inference service to provide spatiotemporal intelligence capabilities, which include: spatiotemporal relationship representation, semantic understanding, environmental perception, spatiotemporal reasoning, and spatial decision-making.

[0076] The spatiotemporal intelligence capability provided by the spatiotemporal intelligence reasoning service is based on a spatial relationship represented by the spatiotemporal objects, during reasoning of the spatiotemporal intelligence, the topological relationships between the objects recorded by the spatiotemporal object are used and outputted, including a knowledge graph and a service process logic related to the object, to reduce the computing complexity in reasoning of the spatiotemporal intelligence and efficiently complete a reasoning task.

[0077] The spatiotemporal relationship representation includes a topological spatial relationship, a sequential spatial relationship, and a metric relationship, as well as a time mark or timestamp; the topological spatial relationship includes: association, adjacency, containment, and the like; the sequential spatial relationship includes: orientation, front, rear, left, right, and the like; and the metric relationship includes: length, angle, and the like.

[0078] The capability of semantic understanding involves analyzing and identifying a spatiotemporal object according to visual information processing or a digital input, and outputting a spatiotemporal object with zero-dimensional object, one-dimensional object, two-dimensional object, and three-dimensional object features, as well as spatial relationships between objects and between an object and an environment.

[0079] The capability of environmental perception involves understanding and predicting a relative relationship between spatiotemporal objects and a change trend based on the spatiotemporal object representation, to provide a support for further application of spatiotemporal relationship reasoning.

[0080] The capability of spatiotemporal reasoning involves understanding the spatial and temporal environments of objects in the real world, or the context scene and virtual environments of multimedia information based on the capability of the environmental perception, reasoning and analyzing a relationship between the objects, and generating a spatial object feature with a future temporal feature and a relationship trend.

[0081] The capability of spatial decision-making involves, based on semantic understanding, environmental perception, and spatiotemporal reasoning, achieving application of spatial decisionmaking in spatiotemporal scenarios through spatiotemporal reasoning and analysis or service process learning according to specific spatiotemporal scenarios and service processes.

[0082] The AGSTI proposed by the present disclosure has a wide range of applications and can be utilized in any field requiring the aforementioned spatiotemporal intelligence capabilities.

[0083] For example, during multimedia information processing, multimedia data is converted into spatial representations of the spatiotemporal intelligence, semantic understanding and reasoning results are generated, and contextually consistent and highly associated answers are provided. In other words, in a multimedia information processing application scenario, a question is understood and semantically analyzed based on spatiotemporal data, and multimedia input is converted into spatiotemporal object representations with spatiotemporal associations, and the spatiotemporal intelligence then analyzes temporal and spatial features of the question and generates answers consistent with the spatiotemporal context. Based on multimodal data, reasoning results with spatiotemporal accuracy are provided during question answering, to ensure consistency of the answers in temporal and spatial dimensions. For time-dependent question answering tasks, the spatiotemporal intelligence can infer future possible contexts based on historical spatiotemporal data and provide predictive answers.

[0084] In an autonomous driving scenario, the spatiotemporal intelligence processes and analyzes spatiotemporal data of each object in the environment in real time. The spatiotemporal intelligence predicts motion trajectories of surrounding objects, and generates path planning and driving decisions based on current spatiotemporal information. Based on multimodal data, the spatiotemporal intelligence enables comprehensive perception of a surrounding environment of a vehicle, forms a driving scene analysis with spatiotemporal associations to support autonomous driving decisionmaking under complex road conditions and in dynamic environments. During autonomous driving, the spatiotemporal intelligence can further predict potential dangers and adjust the driving strategy in real time to improve driving safety and response capability. When applied to autonomous driving systems, the spatiotemporal intelligence identifies and predicts motion trajectories and state changes of surrounding objects by analyzing and processing spatiotemporal data of the environment, to support driving decision-making.

[0085] In an intelligent robotics scenario, the spatiotemporal intelligence processes multimodal spatiotemporal data of an environment in which a robot is located in real time, to enhance perception of the robot to environmental changes. By using the spatiotemporal intelligence for path planning and autonomous decision-making, the robot can accurately judge spatiotemporal relationships of objects in complex environments and make reasonable actions in response. The spatiotemporal intelligence also facilitates behavior prediction, helps the robot understand spatiotemporal change trends of surrounding objects, and supports task execution and dynamic adjustments. During task execution, the robot makes real-time adjustments based on the spatiotemporal intelligence, and ensures its decision-making process accords with the temporal and spatial features of objects in the environment, thereby improving autonomy and accuracy in task completion. Real-time processing and analysis of multimodal data of the environment enhances the ability of the robot to perceive complex spatiotemporal environments and also enhances an independent decision-making ability of the robot.

[0086] In an industrial engineering application scenario, the spatiotemporal intelligence can intelligently process multidimensional engineering thematic graphics, documents, and the like. In summary, any field or scenario that requires the spatiotemporal intelligence capabilities described above can benefit from the application of the AGSTI proposed by the present disclosure.

[0087] In conclusion, according to the method for constructing an AGSTI proposed by the present disclosure, based on object representations of spatial zero-dimensional objects, one-dimensional objects, two-dimensional objects, and three-dimensional objects and a spatial relationship model, a unified and structured topological representation and storage of multimedia information, including text, graphics, images, audio, videos, and the like, in the real world of human society are established, a three-dimensional representation-based multimodal large model training method is employed to construct an AGSTI that can uniformly represent and process all objects in human society, and provide capabilities such as spatiotemporal relationship representation, semantic understanding, environmental perception, spatiotemporal reasoning, and spatial decision-making by combining spatiotemporal information processing, spatiotemporal relationship analysis, knowledge graphs, and service process logic reasoning, thereby solving the problems of the existing large language models, such as single-representation nature, lack of topological relationships, and difficulty in handling three-dimensional spatiotemporal relationships in the real world. The ASTI can be widely applied in fields such as multimedia information processing, theoretical model derivation, autonomous driving, intelligent robotics, and industrial engineering, and has high practicability.

[0088] Although the preferred embodiments of the present disclosure have been described, those skilled in the art, upon learning the fundamental inventive concept, may make additional changes and modifications to these embodiments. Accordingly, the appended claims are intended to be construed as encompassing the preferred embodiments as well as all changes and modifications that fall within the scope of the present disclosure.

[0089] Finally, relational terms such as first and second may be used solely to distinguish one entity or operation from another entity or operation without necessarily requiring or implying any actual relationship or order between such entities or operation. Furthermore, the terms "include", "comprise", or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or terminal device that includes a list of elements does not include only those elements but may include other elements not expressly listed or inherent to such process, method, article, or terminal device. An element proceeded by "comprise a ..." does not, without more constraints, preclude the existence of additional identical elements in the process, method, article, or terminal device that includes the element.

[0090] The embodiments of the present disclosure are described above with reference to the accompanying drawings, but the present disclosure is not limited to the above specific implementations, the above specific implementations are merely schematic and not limiting, and under the enlightenment of the present disclosure, those of ordinary skill in the art can make many modifications without departing from the spirit of the present disclosure and the scope of protection of the claims, and all the modification fall within the scope of protection of the present disclosure.

Claims

1. A method for constructing a large model for artificial general spatiotemporal intelligence (AGSTI), comprising:processing each object and a spatial relationship of the object in the real world using a multidimensional object abstraction representation manner, and establishing a structured spatiotemporal object representation and data storage capable of being processed by the artificial general spatiotemporal intelligence, wherein multidimensional object abstraction refers to zerodimensional objects, one-dimensional objects, two-dimensional objects, and three-dimensional objects;processing a spatiotemporal object by a multimodal spatiotemporal encoding algorithm, to transform spatiotemporal data corresponding to the spatiotemporal object into multimodal spatiotemporal encoded data for subsequent training and application of the artificial general spatiotemporal intelligence large model;constructing a multimodal spatiotemporal training dataset based on the multimodal spatiotemporal encoded data, and training the artificial general spatiotemporal intelligence large model by adopting a three-dimensional representation-based multimodal spatiotemporal training method; anddeploying the trained artificial general spatiotemporal intelligence as a spatiotemporal intelligence inference service to provide spatiotemporal intelligence capability, wherein the spatiotemporal intelligence capability comprises: spatiotemporal relationship representation, semantic understanding, environmental perception, spatiotemporal reasoning, and spatial decisionmaking;the objects in the real world comprise: real-world physical geographic objects, virtual geographic objects, and multimedia objects in inherent relative space of multimedia contexts, and the multimedia objects comprise: text, graphics, images, audio, and videos;the zero-dimensional object refers to the object-based modeling representation of fundamental units in the real world; the fundamental units refer to various physical objects with spatial locations in the real world, as well as multimedia information objects which comprise: text, graphics, images, audio, and videos; the representation of the physical objects is based on spatial locations, environment, and spatial relationships of the objects, the representation of the multimedia information objects is based on digital features, scene or context, and relative space of the objects, comprising but not limited to a sequential location of text in a paragraph and relevance to other text, graphic vector nodes, as well as sequential locations of image frames in a video and relevance to other frames;the one-dimensional object refers to the modeling representation of changes and evolution processes of real-world objects in the temporal dimension, comprising: representing a continuous change process or trajectory of an object in temporal and spatial dimensions as a one-dimensional object, or representing data sequences of zero-dimensional objects or multiple similar zerodimensional objects as a one-dimensional object, and recording a flow process, a logical relationship, and a topological adjacency relationship of the object in temporal and spatial dimensions, wherein the data sequences comprise but are not limited to: coordinate sequences, character sequences, text paragraphs, mathematical formulas, graphic vector segments, image sets, and video frames;the two-dimensional object refers to the modeling representation of adjacency and aggregation of real-world objects, comprising: representing similar one-dimensional objects as a two-dimensional object, and recording adjacency relationships of the objects in temporal and spatial dimensions for larger-scale data analysis and association modeling; and when processing adjacency and aggregation of real-world objects, fully recording intrinsic features of the fundamental units, as well as complex derivation processes or combined association features of multiple data based on the locations, types, or changing processes of the objects in temporal and spatial dimensions; andthe three-dimensional object refers to the modeling representation of large-scale datasets or complex system overall spatiotemporal evolution of real-world objects, comprising: representing big datasets of zero-dimensional objects, one-dimensional objects, and two-dimensional objects as composites or composite processes, recording overall relationships or evolution processes of the objects in temporal and spatial dimensions, generating an overall spatiotemporal body structure based on objects and events of multiple spatiotemporal dimensions, and supporting comprehensive spatiotemporal analysis and prediction of complex systems.

2. The method according to claim 1, wherein the processing each object and a spatial relationship of the object in the real world using a multidimensional object abstraction representation manner comprises:collecting data of each object in the real world, identifying and analyzing spatiotemporal information features of each object in the multimedia information formats, combining the environment, scene, or context in which each object is located, the relative spatial positions of objects, as well as knowledge graph associations and spatial relationships of the objects, to generate multimodal and structured spatiotemporal object representations in the form of zero-dimensional objects, one-dimensional objects, two-dimensional objects, and three-dimensional volumes, wherein the data comprises multimedia information formats, the multimedia information formats comprise:text, graphic, image, audio, and video formats, and the spatiotemporal information features comprise: location, form, state, and temporality; andthe spatiotemporal object representation comprises features of location, temporality, and spatial relationships to ensure consistency and association of objects in temporal and spatial dimensions, the relationships between the spatiotemporal objects are represented through topological relationships, directional or sequential relationships, and metric relationships, the topological relationships comprise: association, adjacency, and containment, the directional or sequential relationships comprise: up, down, left, right, and north, south, east, west, and the metric relationships comprise: distance, angle, or proximity.

3. The method according to claim 1, wherein the processing a spatiotemporal object by a multimodal spatiotemporal encoding algorithm, to transform spatiotemporal data corresponding to the spatiotemporal object into multimodal spatiotemporal encoded data for subsequent training and application of the spatiotemporal intelligence comprises:step Tl: based on the spatiotemporal objects, taking each fundamental object in the spatiotemporal data as a node in a spatiotemporal topology graph and a spatiotemporal adjacency between the nodes as an edge of the graph, defining weights of the edges between the nodes by calculating a spatiotemporal distance or association between the objects, and constructing a spatiotemporal topological graph structure, wherein the association between the objects is determined based on similarity, proximity, or temporal sequence in the spatiotemporal data, and the spatiotemporal data refers to multimodal spatiotemporal multidimensional data; andstep T2: performing data encoding based on a graph neural network, constructing a pre-training task through spatial association, and generating high-dimensional representation of the spatiotemporal topological graph structure.

4. The method according to claim 3, wherein the performing data encoding based on a graph neural network, constructing a pre-training task through spatial association, and generating highdimensional representation of the spatiotemporal topological graph structure comprises:by aggregating information from the nodes of the spatiotemporal topology graph and adjacent nodes, updating spatiotemporal features of each node to capture spatial and temporal associations between the multimodal spatiotemporal multidimensional data; andperforming deep learning on the multimodal spatiotemporal multidimensional data by using a multi-layer graph neural network to generate encoded representations with enhanced spatiotemporal associations for further model training.

5. The method according to claim 1, wherein the training or fine-tuning the spatiotemporal intelligence by adopting a three-dimensional representation-based multimodal spatiotemporal training method comprises:step Nl: based on the multimodal spatiotemporal encoded data, collecting or generating data containing spatiotemporal information according to a specific application scenario and task requirements, and constructing the multimodal spatiotemporal training dataset, wherein the training dataset comprises tag, and temporal and spatial attribute information of multimodal data, and is used for guiding adaptability and accuracy of the spatiotemporal intelligence in a specific task, so as to train a spatiotemporal association capability of the spatiotemporal intelligence; andstep N2: training the spatiotemporal intelligence by taking the constructed multimodal spatiotemporal training dataset as an input, and using a multidimensional spatiotemporal attention network model and the multimodal spatiotemporal training method, to capture temporal sequence relationships and spatial features annotated with zero-dimensional objects, one-dimensional objects, two-dimensional objects, and three-dimensional objects in the data.

6. The method according to claim 5, wherein the multimodal spatiotemporal training method refers to multimodal spatiotemporal training combining three-dimensional representations of the zero-dimensional object, the one-dimensional object, the two-dimensional object, and the three-dimensional object, as well as the temporal dimension, an underlying method of the spatiotemporal intelligence adopts a three-dimensional representation method, and modeling is performed based on three-dimensional spatial representation and the temporal dimension, as well as the temporal sequence and the spatial features of the multimodal spatiotemporal multidimensional data, to accurately capture a complex relationship in the temporal and spatial dimensions, reflect a spatiotemporal dynamic change of the real world, and enhance the performance of the spatiotemporal intelligence in spatiotemporal reasoning.

7. The method according to claim 5, wherein during construction of the multimodal spatiotemporal training dataset, the focus is on acquiring data diversity and coverage, to ensure robustness and a generalization capability of the spatiotemporal intelligence in different contexts, wherein text data comprises a plurality of types of content, comprising but not limited to, a natural language description, a command, and a dialogue, covering a broad semantic range and different degrees of complexity; image data comprises, but is not limited to, static and dynamic images from different perspectives, in different environments, and in different time periods; and geospatial datacomprises, but is not limited to, formatted spatial data from different industries and with different precisions; and spatial relationships related to the spatiotemporal data are constructed.

8. The method according to claim 1, wherein the spatiotemporal intelligence capability provided by the spatiotemporal intelligence reasoning service is based on a spatial relationship represented by the spatiotemporal objects, during reasoning of the spatiotemporal intelligence, the topological relationships between the objects recorded by the spatiotemporal object are used and outputted, comprising a knowledge graph and a service process logic related to the object, to reduce the computing complexity in reasoning of the spatiotemporal intelligence and efficiently complete a reasoning task.

9. The method according to claim 8, wherein the spatiotemporal relationship representation comprises a topological spatial relationship, a sequential spatial relationship, and a metric relationship, as well as a time mark or timestamp, the topological spatial relationship comprises: association, adjacency, and containment, the sequential spatial relationship comprises: orientation, front, rear, left, and right, and the metric relationship comprises: length and angle;the capability of semantic understanding involves analyzing and identifying a spatiotemporal object according to visual information processing or a digital input, and outputting a spatiotemporal object with zero-dimensional object, one-dimensional object, two-dimensional object, and three-dimensional object features, as well as spatial relationships between objects and between an object and an environment;the capability of environmental perception involves understanding and predicting a relative relationship between spatiotemporal objects and a change trend based on the spatiotemporal object representation, to provide a support for further application of spatiotemporal relationship reasoning;the capability of spatiotemporal reasoning involves understanding the spatial and temporal environments of objects in the real world, or the context scene and virtual environments of multimedia information based on the capability of environmental perception, reasoning and analyzing an association relationship between the objects, and generating a spatial object feature with a future temporal feature and an association relationship trend; andthe capability of spatial decision-making involves, based on semantic understanding, environmental perception, and spatiotemporal reasoning, achieving application of spatial decisionmaking in spatiotemporal scenarios through spatiotemporal reasoning and analysis or service process learning according to specific spatiotemporal scenarios and service processes.