Method for constructing a large-scale, general-purpose artificial spatiotemporal intelligence model
A general-purpose artificial spatiotemporal intelligence model addresses the limitations of large-scale language models by integrating multidimensional object abstraction and spatiotemporal coding, enhancing decision-making and adaptability in complex scenarios.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- BEIJING LONGRUAN TECHNOLOGIES INC
- Filing Date
- 2025-02-21
- Publication Date
- 2026-04-27
AI Technical Summary
Large-scale language models lack the ability to effectively process and integrate multidimensional spatiotemporal information, making them unsuitable for applications requiring three-dimensional spatial and temporal interactions, such as autonomous driving and robotics, due to their one-dimensional data structure and limited representation of spatial relationships.
A method for constructing a general-purpose artificial spatiotemporal intelligence large-scale model using a multidimensional object abstraction representation scheme, spatiotemporal coding algorithm, and multimodal training to process and represent 0-dimensional, 1-dimensional, 2-dimensional, and 3-dimensional objects, incorporating spatiotemporal relationships and data storage.
The model enhances decision-making accuracy and adaptability in complex scenarios by providing spatiotemporal intelligence capabilities, enabling comprehensive processing of multimedia data and supporting applications like autonomous driving, intelligent robots, and industrial engineering.
Smart Images

Figure 2026070438000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and particularly to a method for constructing a general artificial spatio-temporal intelligence large-scale model.
Background Art
[0002] In recent years, artificial intelligence technology has been developing rapidly. Large language models (LLMs) have achieved remarkable results in fields such as natural language processing, speech recognition, and computer vision due to their powerful representation learning and generalization capabilities. They have become a popular area of research and development in the field of artificial intelligence and are widely used in industries such as literature, art, education, and finance, having a significant impact on the development of human society.
[0003] Currently, the parameter volume of large language models usually reaches several billion or even several trillion, and the requirements for computing resources and training technologies are extremely high. Training such large-scale models requires a large amount of data training, acquisition, cleaning, and annotation, which is a process that requires a great deal of cost and time, and will soon face bottlenecks in computing resources and data resources.
[0004] Furthermore, large-scale language models, while essentially processing linguistic text data and capable of handling multimedia data such as images and videos through encoding schemes, are fundamentally based on a one-dimensional data structure with a relatively simple structure. This structure consists of tokenized text sequences and lacks representation of spatial relationships. Moreover, large-scale models are primarily trained on text data and lack direct experience of interaction with the physical world. As such, they cannot learn spatial, temporal, and physical interactions through their own experience like humans, and cannot effectively integrate multidimensional information of time and space in the real world. This makes it difficult to support the perception, understanding, and manipulation of three-dimensional space, and they cannot be directly used in fields that require processing the interaction of multidimensional spatiotemporal information, such as industrial production and civil engineering. For example, in applications such as autonomous driving and robotics, it is necessary to accurately model and predict the spatiotemporal changes of various objects in the environment. However, since the processing methods of large-scale language models are based on processing one-dimensional data structures, they require extremely high computational power support. Furthermore, the lack of a framework for analyzing temporal and spatial dimensions together makes it difficult to provide accurate and consistent spatiotemporal relationship analysis results.
[0005] Currently, large-scale language models are primarily applied to intelligent processing in the humanities, such as text, images, and video—in other words, "humanities intelligence"—and are not suitable for processing data in the sciences and engineering fields, particularly the processing and interpretation of spatiotemporal data in the field of human and social engineering, such as the intelligent processing of multidimensional design graphics for various projects. Graphic design and applications in the engineering field are problems that must be addressed in an industrialized and intelligent society; otherwise, the application scenarios of artificial intelligence will remain partial and incomplete.
[0006] The human real world is composed of three-dimensional physical objects or multimedia information objects with temporal and spatial characteristics. The three-dimensional world follows physical laws and possesses unique structures, attributes, and spatiotemporal relationships. Due to the inherent limitations of existing large-scale language models' one-dimensional data structure principles and data modeling methods, there is an urgent need to find a unified, universal method for representing and processing large-scale models, and to research large-scale models better suited to representing, perceiving, reasoning, and interactive decision-making in the real universe. Such models would comprehensively process multimedia data such as text, graphics, images, audio, and video that human society faces, and address problems across various specialized fields, including humanities, sciences, and engineering. [Overview of the project] [Problems that the invention aims to solve]
[0007] In light of the above problems, the present invention proposes a method for constructing a general-purpose artificial spatiotemporal intelligence large-scale model. [Means for solving the problem]
[0008] The present invention provides a multidimensional object abstraction representation scheme for processing various objects and spatial relationships in the real world, and establishes a structured spatiotemporal object representation and data storage that can be processed collectively by a spatiotemporal intelligence large-scale model, wherein the multidimensional object abstraction refers to 0-dimensional objects, 1-dimensional objects, 2-dimensional objects, and 3-dimensional objects. The steps include processing the aforementioned spatiotemporal object with a multimodal spatiotemporal coding algorithm and converting the spatiotemporal data corresponding to the spatiotemporal object into coded data for subsequent training and application of spatiotemporal intelligence large-scale models, The steps include creating a multimodal spatiotemporal training dataset based on multimodal spatiotemporal coded data, training or fine-tuning it using a spatiotemporal multimodal training method with three-dimensional characterization to form the spatiotemporal intelligent large-scale model, and The present invention discloses a method for constructing a general-purpose artificial spatiotemporal intelligence model, comprising the steps of: issuing the trained spatiotemporal intelligence model as a spatiotemporal intelligence model inference service and providing spatiotemporal intelligence capabilities, wherein the spatiotemporal intelligence capabilities include spatiotemporal relational representation, semantic understanding, environmental perception, spatiotemporal reasoning, and spatial decision-making.
[0009] Optionally, the step of processing various objects and spatial relationships in the real world using a multidimensional object abstraction representation scheme is: The process includes the steps of collecting, recognizing, and analyzing data of various real-world objects, including but not limited to multimedia information formats; obtaining spatiotemporal information features of the various objects in the multimedia information formats; and combining the environment, scene, or context in which the various objects exist, relative space, and knowledge graphs and spatial relationships of the various objects to generate multimodal structured spatiotemporal object representations of 0-dimensional objects, 1-dimensional objects, 2-dimensional objects, and 3-dimensional objects, wherein the multimedia information formats include text formats, graphics formats, image formats, audio formats, and video formats, and the spatiotemporal information features include location, form, state, and tense. A spatiotemporal object representation simultaneously possesses position, tense, and spatial relational features, ensuring the consistency and relevance of objects in the temporal and spatial dimensions, wherein the spatiotemporal objects are represented by topological relationships, orientation or order relationships, and metric relationships, where the topological relationships include association, adjacency, and inclusion; the orientation or order relationships include up, down, left, right, east, west, north, and south; and the metric relationships include distance, angle, or proximity.
[0010] Optionally, the various objects of the real world include actual geospatial objects in the real world, virtual geospatial objects, and multimedia type objects in relative space specific to the multimedia context, and the multimedia type objects include various types such as text, graphics, images, audio, and video. The aforementioned 0-dimensional object is an objectification modeling representation of a fundamental unit of the real world, where the object of the fundamental unit refers to various physical objects and multimedia information objects that have spatial locations in the real world, where the multimedia information objects include text, graphics, images, audio, and video, where the various physical objects are represented by the spatial location, environment, and spatial relationships of the object, where the multimedia information objects are represented by the numerical features, scene or context, and relative space of the object, and include, but are not limited to, the sequence position of a character within a paragraph and the correlation relationship between other characters, graphics vector nodes, and the sequence position of an image frame within a video and the correlation relationship between other frames. The aforementioned one-dimensional object is a modeling representation of the process of change and evolution of a real-world object in the time dimension, representing the process or trajectory of a continuous change of an object in the time and spatial dimensions as a one-dimensional object, or including a data sequence of zero-dimensional objects, representing multiple zero-dimensional objects of the same kind as one-dimensional objects, and recording their flow process, logical relationships, and topological adjacencies in spacetime, the data sequence includes, but is not limited to, coordinate sequences, character sequences, text paragraphs, mathematical formulas, graphics vector arc segments, image sets, and video frames. The aforementioned two-dimensional object is a modeling representation of the adjacency and aggregation of objects in the real world, representing one-dimensional objects of the same type as two-dimensional objects, recording their temporal and spatial adjacency relationships, and using them for broader data analysis and related modeling. When processing the adjacency and aggregation of objects in the real world, it includes fully recording the characteristics of the base unit itself and complex derivation processes or combinations and associations of multiple data, depending on the object's position, type, or transformation process in the temporal or spatial dimension. The aforementioned three-dimensional object is a modeling representation of the overall spatiotemporal evolution of a large dataset or complex system in real-world objects, and includes displaying a big dataset of multidimensional zero-dimensional, one-dimensional, and two-dimensional objects in a composite or composite process, recording its overall relationships or transitional processes in spatiotemporaneity, and combining objects and events from multiple spatiotemporal dimensions to generate an overall chronotope structure that supports comprehensive spatiotemporal analysis and prediction of complex systems.
[0011] The optional step of processing the spatiotemporal data corresponding to the spatiotemporal object using a multimodal spatiotemporal coding algorithm to convert it into coded data for subsequent spatiotemporal intelligence large-scale model training and application is: Step T1 involves constructing a spatiotemporal topology diagram structure by considering each basic object in the spatiotemporal data as a node in a spatiotemporal topology diagram based on the aforementioned spatiotemporal objects, defining spatiotemporal proximity between each node as an edge in the diagram, calculating the spatiotemporal distance or correlation between objects, and defining the edge weights between nodes, wherein the correlation between objects is determined by similarity, neighborliness, or time-series order in the spatiotemporal data, and the spatiotemporal data refers to multimodal spatiotemporal multidimensional data. The method includes step T2, which involves encoding data based on a graph neural network, constructing a pre-trained task through spatial association, and representing the spatiotemporal topology diagram structure in high dimensions.
[0012] The optional steps include encoding data based on a graph neural network, constructing a pre-trained task through spatial association, and representing the spatiotemporal topology diagram structure in high dimensions. The steps include: updating the spatiotemporal characteristics of each node by aggregating information on nodes and neighboring nodes in a spatiotemporal topology diagram, and capturing the spatial and temporal relationships between the multimodal spatiotemporal multidimensional data; The method includes the step of deep learning the multimodal spatiotemporal multidimensional data using a multilayer graph neural network to generate an encoded representation that enhances spatiotemporal relationships for further model training.
[0013] The step of training or fine-tuning the spatiotemporal intelligence large-scale model using a spatiotemporal multimodal training method with three-dimensional characterization, at the discretion of the user, is: Step N1 involves collecting or generating data containing spatiotemporal information based on the aforementioned multimodal spatiotemporal coded data, combining specific application scenes and task requirements, creating a multimodal spatiotemporal training dataset, the training dataset containing labels and spatiotemporal attribute information of the multimodal data, guiding the adaptability and accuracy of the spatiotemporal intelligent large-scale model in a specific task, and training the spatiotemporal-related capabilities of the spatiotemporal intelligent large-scale model. Step N2 includes training a spatiotemporal intelligence large-scale model using a spatiotemporal multimodal training method with a multidimensional spatiotemporal attention network model as input to a constructed multimodal spatiotemporal training dataset, and capturing time-sequence relationships and spatial features annotated by 0-dimensional, 1-dimensional, 2-dimensional, and 3-dimensional objects from the data.
[0014] Optionally, the spatiotemporal multimodal training method refers to spatiotemporal multimodal training of 0-dimensional objects, 1-dimensional objects, 2-dimensional objects, and 3-dimensional objects with 3D characterization + time dimension, employing the 3D characterization method in the base layer of the spatiotemporal intelligent large-scale model, and modeling by combining the time sequence and spatial features of the multimodal spatiotemporal multidimensional data based on the 3D spatial representation and time dimension, thereby more accurately capturing complex relationships in time and space, reflecting spatiotemporal dynamic changes in the real world, and improving the spatiotemporal inference representation of the spatiotemporal intelligent large-scale model.
[0015] Optionally, when creating the multimodal spatiotemporal training dataset, in order to ensure the robustness and generalization ability of the spatiotemporal intelligent large-scale model in various situations, the data diversity and coverage are primarily obtained, and the text data includes various types of content, including but not limited to natural language descriptions, commands, and dialogues, covering a broad semantic range and various levels of complexity; the image data includes but not limited to still and moving images from various viewpoints, various environments, and various time periods; and the geospatial data includes but not limited to spatial data in various industries and formats of varying precision, thereby constructing spatial relationships with respect to the spatiotemporal data.
[0016] Optionally, the spatiotemporal intelligence capabilities of the spatiotemporal intelligence large-scale model inference service output, based on the spatial relationships of spatiotemporal object representations, using topological relationships between objects recorded by the spatiotemporal objects, including knowledge graphs and business flow logic about the objects, in the spatiotemporal intelligence large-scale model inference process, thereby reducing the computational complexity of the inference by the spatiotemporal intelligence large-scale model and efficiently performing the inference task.
[0017] Optionally, the spatiotemporal relationship representation includes a topological spatial relationship, a sequential spatial relationship, a metric relationship, and a time mark or timestamp, the topological spatial relationship includes association, adjacency, and inclusion, the sequential spatial relationship includes orientation and front / back / left / right, and the metric relationship includes length and angle. The aforementioned semantic comprehension ability involves analyzing and recognizing spatiotemporal objects through vision information processing or digitized input, and outputting spatiotemporal objects with 0-dimensional, 1-dimensional, 2-dimensional, and 3-dimensional object features, as well as spatial relationships between objects and between objects and the environment. The aforementioned environmental perception ability is to understand and predict the relative relationships and trends of change between spatiotemporal objects based on the spatiotemporal object representation, and to provide support for further spatiotemporal relationship inference applications. The aforementioned spatiotemporal reasoning ability, based on the aforementioned environmental perception ability, understands the space in which objects exist in the real world, the temporal environment, or the contextual scene of multimedia information, or the virtual environment, and infers and analyzes the relationships between objects to generate trends in spatial object features and relationships of future temporal features. The aforementioned spatial decision-making capability involves combining specific spatiotemporal scenes and business flows based on semantic understanding, environmental perception, and spatiotemporal reasoning, and realizing the application of spatial decision-making in spatiotemporal scenes through spatiotemporal reasoning analysis or business flow learning. [Effects of the Invention]
[0018] The present invention provides a method for constructing a general-purpose artificial spatiotemporal intelligence large-scale model. First, a multidimensional object abstraction representation scheme is used to process various objects and spatial relationships in the real world, establishing a structured spatiotemporal object representation and data storage that can be processed collectively by the spatiotemporal intelligence large-scale model. Next, spatiotemporal objects are processed using a multimodal spatiotemporal coding algorithm, and the spatiotemporal data corresponding to the spatiotemporal objects is converted into coded data for subsequent training and application of the spatiotemporal intelligence large-scale model. Furthermore, a multimodal spatiotemporal training dataset is created based on the multimodal spatiotemporal coded data, and the spatiotemporal intelligence large-scale model is trained or fine-tuned using a spatiotemporal multimodal training method with three-dimensional characterization to form a spatiotemporal intelligence large-scale model. Finally, the trained spatiotemporal intelligence large-scale model is issued as a spatiotemporal intelligence large-scale model inference service, providing spatiotemporal intelligence capabilities.
[0019] A construction method of a general artificial spatio-temporal intelligence model (which can be defined as Artificial General SpatioTemporal Intelligence and abbreviated as AGSTI) according to the present invention is proposed. This construction method is based on the object representation methods of 0-dimensional objects, 1-dimensional objects, 2-dimensional objects, and 3-dimensional objects in space and the spatial relationship model, and establishes a unified structured topological representation and storage of various multimedia information such as texts, graphics, images, audios, and videos in the real world of human society. Using a multi-modal large-scale model training method with three-dimensional characterization, an artificial spatio-temporal large-scale model that can comprehensively represent and process all objects in human society is constructed. By combining spatio-temporal information processing, spatio-temporal relationship analysis, knowledge graphs, and business process logical inferences, capabilities such as spatio-temporal relationship representation, semantic understanding, environmental perception, spatio-temporal inference, and spatial decision-making are provided, solving the difficult problems of conventional large-scale language models that have a single representation, lack topological relationships, and are difficult to handle the problems of the relationship between time and space in the three-dimensional real world. The general artificial spatio-temporal intelligence large-scale model constructed based on the present invention can greatly improve the decision-making accuracy and adaptability of intelligent systems in complex scenarios, and can be fully applied to application fields such as multimedia information processing, theoretical model derivation, autonomous driving, intelligent robots, and industrial engineering, providing new technical support for these application fields.
Brief Description of the Drawings
[0020] Various other advantages and superiority will become clear to those skilled in the art by reading the detailed description of the following preferred embodiments. The drawings are for the purpose of showing the preferred embodiments and are not intended to limit the present invention. Also, throughout the drawings, the same reference numerals are assigned to the same components. [Figure 1] It is a flowchart of a construction method of a general artificial spatio-temporal intelligence large-scale model according to an embodiment of the present invention. [Figure 2] It is a schematic diagram of an example of an abstraction representation method of an object in the real world in an embodiment of the present invention. [Modes for carrying out the invention]
[0021] To make the above-mentioned objectives, features, and advantages of the present invention clearer and easier to understand, the present invention will be described in further detail with reference to the drawings and specific embodiments. It should be understood that the specific embodiments described herein are used solely to illustrate the present invention and are not all embodiments of the present invention, but only a selection of embodiments, and are not intended to limit the present invention.
[0022] The method for constructing a general-purpose artificial spatiotemporal intelligence large-scale model according to the present invention includes the following steps 101 to 104, referring to the flowchart of the method for constructing a general-purpose artificial spatiotemporal intelligence large-scale model shown in Figure 1.
[0023] Step 101: Using a multidimensional object abstraction representation scheme, we will process various objects and spatial relationships in the real world and establish a structured spatiotemporal object representation and data storage that can be processed collectively by a spatiotemporal intelligence large-scale model. Multidimensional object abstraction refers to 0-dimensional objects, 1-dimensional objects, 2-dimensional objects, and 3-dimensional objects.
[0024] In this invention, an innovative theory of multidimensional object abstraction is proposed, which includes 0-dimensional objects, 1-dimensional objects, 2-dimensional objects, and 3-dimensional objects. To better understand the technical solutions of this invention, the theory of multidimensional object abstraction will first be described in detail.
[0025] The multi-dimensional object abstraction according to the present invention analyzes all kinds of physical objects in human society and data such as characters, graphics, images, audio, videos, etc. obtained by recognizing various objects, and abstracts them into structured 0-dimensional objects, 1-dimensional objects, 2-dimensional objects, and 3-dimensional objects. By using a 3D feature characterization model representation and training method, at the level of data model representation and training, it solves the limitations of large language models based on text representation and optimizes the computational power requirements for model training and inference.
[0026] The so-called various objects in the real world include actual geographical space objects in the real world, virtual geographical space objects, and multimedia type objects of relative space unique to the multimedia context. This multimedia type object includes each type such as text, graphics, images, audio, videos, etc. To better understand the above theory, some specific examples will be combined and described.
[0027] (1) 0-dimensional object: vector node, number "1", character "A", Chinese character "Mao", "one tree, one rose, one person, one tiger, one image, one snapshot, auditory and olfactory information at one point in time, or one part or building component" when a physical entity exists, etc., and basic units of the real world such as mathematical operation symbols "product symbol Π or summation symbol Σ or sine symbol sin", etc. In the general artificial spacetime intelligent large-scale model of the present invention, all of these participate in the model's representation and training in the form of 0-dimensional objects.
[0028] Combined with the above example, a 0-dimensional object can be understood as an objectification modeling representation of the fundamental units of the real world. Here, fundamental units of objects refer to various physical objects with spatial locations in the real world, and multimedia information objects, which include text, graphics, images, audio, and video. Spatial locations include the absolute and local coordinates of various objects, as well as the relative sequence and page numbers of the scene and context in which the object is located, and relative coordinates. Various physical objects are represented by the spatial location, environment, and spatial relationships of the object. Multimedia information objects are represented by the numerical features of the object, the scene or context, and the relative space, and include, but are not limited to, the sequence location of text within a paragraph, the correlation relationship between page numbers in a chapter and other text, and the correlation relationship between the sequence location of an image frame in a video and other frames.
[0029] (2) One-dimensional objects: Vector arch segments, rows of numbers or letters, phrases or sentences, subsystems formed by multiple parts or building components, local calculation formulas, etc., where physical entities exist, such as "a row of trees, a row of roses, a row of people or a person walking or running, a row of tigers or a tiger walking or running, a video segment, a segment of auditory or olfactory information." In the general-purpose artificial spatiotemporal intelligence large-scale model of the present invention, all of these participate in the representation and training of the model in the form of one-dimensional objects.
[0030] Combined with the above examples, a one-dimensional object can be understood as a model and representation of the process of change and evolution of a real-world object in the time dimension, representing the continuous process or trajectory of change of an object in the time and spatial dimensions as a one-dimensional object, or a data sequence of zero-dimensional objects, representing multiple zero-dimensional objects of the same kind as a one-dimensional object, and combining topological node features to record its flow process, logical relationships, and topological adjacencies in spacetime. Data sequences include, but are not limited to, coordinate sequences, character sequences, text paragraphs, mathematical formulas, image sets, and video frames.
[0031] (3) Two-dimensional objects: vector polygons, multi- or multi-segment data or strings, physical entities such as "a forest, a rose bush, a crowd of people, a pack of tigers, a complete story of a film epic or a complete video of television, a stepwise auditory and olfactory information set", a fully functional machine or building, a complete calculation or derivation process (combination of multiple formulas), etc.
[0032] Combining the above examples, a 2D object is a modeling representation of the adjacency and aggregation of objects in the real world. It represents multiple 0-dimensional and 1-dimensional objects of the same type as 2D objects, recording their temporal and spatial adjacency relationships for use in broader data analysis and spatiotemporal object modeling. When processing the association and aggregation of objects in the real world, it involves fully recording the characteristics of the base unit itself and complex derivation processes or combinations and associations of multiple data, depending on the object's position, type, or transformation process in the temporal or spatial dimension.
[0033] (4) Three-dimensional objects: vector polyhedra, datasets or text sets, parks or communities, zoos, "fields of farmland or forests or grasslands, a movie or television series" in which physical entities exist, workshops or factories, clusters of buildings, a monograph on mathematics, physics, or chemistry, etc.
[0034] Combining the above examples, a 3D object is a modeling representation of the overall spatiotemporal evolution of a large dataset or complex system in the real world, displaying a complex or composite process of multidimensional 0D, 1D, and 2D objects in a big dataset, recording its overall relationships or transitional processes in spatiotemporaneity. This includes combining objects and events from multiple spatiotemporal dimensions to generate an overall chronotope structure that supports comprehensive spatiotemporal analysis and prediction of complex systems.
[0035] The multidimensional object abstraction theory described above can be understood more intuitively when combined with schematic diagrams of abstract representations of exemplary real-world objects, as shown in Figure 2. Physical objects within real-world objects (such as the example number "1" on the left side of Figure 2, a column of numbers, data with multiple segments, a dataset, or a set of text) can be represented, depending on their type, as a unified space-time representation, i.e., a 0-dimensional object, a 1-dimensional object, a 2-dimensional object, or a 3-dimensional object. These correspond to vector nodes, arc segments, polygons, and polyhedra, possess temporal attributes, and have multiple relationships such as spatial relationships, topology, order, and metrics.
[0036] The above provides a detailed interpretation and explanation of the theory of multidimensional object abstraction. To construct a general-purpose artificial spatiotemporal intelligence model (which can be defined as Artificial General SpatioTemporal Intelligence, abbreviated as AGSTI) according to the present invention, a multidimensional object abstraction representation scheme is first used to process various objects and spatial relationships in the real world, establishing a structured spatiotemporal object representation and data storage that can be processed uniformly by a large-scale spatiotemporal intelligence big model. A more preferred method includes the following:
[0037] First, data on various real-world objects, including but not limited to multimedia information formats, is collected. Next, the spatiotemporal information features of the various objects in the obtained multimedia information formats are recognized and analyzed. Furthermore, the environment, scene or context in which the various objects exist, relative space, and knowledge graph relationships and spatial relationships of the various objects are combined to generate multimodal and structured spatiotemporal object representations of 0-dimensional, 1-dimensional, 2-dimensional, and 3-dimensional objects. Here, multimedia information formats include text formats, graphics formats, image formats, audio formats, and video formats, and spatiotemporal information features include location, form, state, and tense.
[0038] A spatiotemporal object representation simultaneously possesses positional, temporal, and topological relational features, thus ensuring the consistency and relevance of objects in the temporal and spatial dimensions. Spatiotemporal objects may be represented by topological relationships, azimuthal or orthogonal relationships, or metric relationships, where topological relationships include association, adjacency, and inclusion; azimuthal or orthogonal relationships include up, down, left, right, east, west, north, and south; and metric relationships include distance, angle, or proximity.
[0039] Step 102: Process the spatiotemporal objects using a multimodal spatiotemporal coding algorithm to convert the spatiotemporal data corresponding to the spatiotemporal objects into multimodal spatiotemporal coded data for subsequent spatiotemporal intelligence large-scale model training and application.
[0040] After establishing a structured spatiotemporal object representation and data storage that can be processed in bulk by spatiotemporal intelligence large-scale models, it is necessary to process the spatiotemporal objects with a multimodal spatiotemporal coding algorithm, thereby converting the spatiotemporal data corresponding to the spatiotemporal objects into multimodal spatiotemporal coded data for subsequent training and application of spatiotemporal intelligence large-scale models. A preferred method includes the following steps:
[0041] Step T1: Based on the spatiotemporal objects, each basic object in the spatiotemporal data is considered a node in the spatiotemporal topology diagram, spatiotemporal proximity between each node is defined as an edge in the diagram, and the spatiotemporal distance or correlation between objects is calculated to define edge weights between nodes, thereby constructing the spatiotemporal topology diagram structure. The correlation between objects is determined by similarity, neighborliness, or time-series order in the spatiotemporal data, and the spatiotemporal data refers to multimodal spatiotemporal multidimensional data.
[0042] Step T2: Encode data based on a graph neural network, construct a pre-trained task through spatial association, and represent the spatiotemporal topology diagram structure in high dimensions.
[0043] With respect to higher-dimensional representations, a preferred method includes the steps of first updating the spatiotemporal features of each node by aggregating information on nodes and neighboring nodes in a spatiotemporal topology diagram to capture the spatial and temporal relationships between the multimodal spatiotemporal multidimensional data, and second deep learning the multimodal spatiotemporal multidimensional data using a multilayer graph neural network to generate an encoded representation that enhances the spatiotemporal relationships for further model training.
[0044] Step 103: Based on multimodal spatiotemporal coded data, create a multimodal spatiotemporal training dataset, and train or fine-tune it using a spatiotemporal multimodal training method with 3D characterization to form a spatiotemporal intelligent large-scale model.
[0045] After converting spatiotemporal data into multimodal spatiotemporal coded data for training and application of subsequent spatiotemporal intelligent large-scale models, a multimodal spatiotemporal training dataset may be created based on the multimodal spatiotemporal coded data. Then, based on this multimodal spatiotemporal training dataset, a spatiotemporal multimodal training method with three-dimensional characterization may be used to train or fine-tune the spatiotemporal intelligent large-scale model.
[0046] A preferred method for training or fine-tuning a spatiotemporal intelligence large-scale model includes the following steps:
[0047] Step N1: Based on multimodal spatiotemporal coded data, collect or generate data containing spatiotemporal information by combining specific application scenes and task requirements to create a multimodal spatiotemporal training dataset. This training dataset includes labels and spatiotemporal attribute information of the multimodal data, and guides the adaptability and accuracy of the spatiotemporal intelligent large-scale model in specific tasks, and trains the spatiotemporal-related capabilities of the spatiotemporal intelligent large-scale model.
[0048] Step N2: Using the constructed multimodal spatiotemporal training dataset as input, train a spatiotemporal intelligence large-scale model using a spatiotemporal multimodal training method with a multidimensional spatiotemporal attention network model to capture time-sequence relationships and spatial features annotated by 0-dimensional, 1-dimensional, 2-dimensional, and 3-dimensional objects in the data.
[0049] The spatiotemporal multimodal training method in step N2 refers to the three-dimensional characterization of 0-dimensional, 1-dimensional, 2-dimensional, and 3-dimensional objects, plus spatiotemporal multimodal training of the time dimension. Unlike the training of conventional language models or multimodal language models, the foundational layer of conventional model training is usually based on sequence length and attention mechanisms and uses tokenized one-dimensional representations. The training concepts and architectures of the two are completely different.
[0050] The present invention employs a three-dimensional characterization method in the foundational layer of spatiotemporal intelligent large-scale models. Based on three-dimensional spatial representations and time dimensions, it combines time sequences and spatial features of multimodal spatiotemporal multidimensional data to model, more accurately capturing complex relationships in time and space, reflecting spatiotemporal dynamic changes in the real world, and improving the spatiotemporal inference representation of spatiotemporal intelligent large-scale models.
[0051] On the other hand, when creating multimodal spatiotemporal training datasets, the primary goal is to acquire data diversity and coverage to ensure the robustness and generalization capabilities of spatiotemporal intelligence large-scale models in various situations. Text data includes various types of content, including but not limited to natural language descriptions, commands, and dialogues, covering a broad semantic range and varying levels of complexity; image data includes but not limited to still and moving images from various viewpoints, environments, and time periods; and geospatial data includes but not limited to spatial data from various industries and in various precision formats. This geospatial data also constructs spatial relationships with respect to spatiotemporal data.
[0052] Step 104: The trained spatiotemporal intelligence large-scale model is issued as a spatiotemporal intelligence large-scale model inference service, providing spatiotemporal intelligence capabilities, which include spatiotemporal relational representation, semantic understanding, environmental perception, spatiotemporal reasoning, and spatial decision-making.
[0053] After training or fine-tuning to form a spatiotemporal intelligence large-scale model, i.e., after training the spatiotemporal intelligence large-scale model, this model can be published as a spatiotemporal intelligence large-scale model inference service, thereby providing spatiotemporal intelligence capabilities, which include spatiotemporal relational representation, semantic understanding, environmental perception, spatiotemporal reasoning, and spatial decision-making.
[0054] The spatiotemporal intelligence capabilities of the spatiotemporal intelligence large-scale model inference service, based on the spatial relationships of spatiotemporal object representations, output using inter-object topology relationships recorded by spatiotemporal objects, including knowledge graphs and business flow logic about the objects, in the spatiotemporal intelligence large-scale model inference process. This reduces the computational complexity of inference by the spatiotemporal intelligence large-scale model and efficiently performs inference tasks.
[0055] So-called spatiotemporal relationship representations include topological relationships, sequential relationships, metric relationships, and time marks or timestamps. Topological relationships include association, adjacency, inclusion, etc. Sequential relationships include orientation and front / back / left / right, etc. Metric relationships include length and angle, etc.
[0056] So-called semantic comprehension ability involves analyzing and recognizing spatiotemporal objects through vision information processing or digitized input, and outputting spatiotemporal objects with 0-dimensional, 1-dimensional, 2-dimensional, and 3-dimensional object features, as well as spatial relationships between objects and between objects and the environment.
[0057] So-called environmental perception ability involves understanding and predicting the relative relationships and trends of change between spatiotemporal objects based on spatiotemporal object representations, and providing support for further spatiotemporal relationship inference applications.
[0058] So-called spatiotemporal reasoning ability is the ability to understand the space in which objects exist in the real world, the temporal environment, or the contextual scene of multimedia information, or the virtual environment, based on the aforementioned environmental perception ability, and to infer and analyze the relationships between objects to generate future temporal features, spatial object features, and relationship trends.
[0059] So-called spatial decision-making ability is realized by combining specific spatiotemporal scenes and business flows, based on the aforementioned semantic understanding, environmental perception, and spatiotemporal reasoning, and through spatiotemporal reasoning analysis or business flow learning, thereby realizing the application of spatial decision-making in spatiotemporal scenes.
[0060] The general-purpose artificial spatiotemporal intelligence large-scale model according to the present invention has a wide range of specific applications and can be applied in all fields that require the above-mentioned spatiotemporal intelligence capabilities.
[0061] For example, in multimedia information processing, multimedia data is transformed into a spatial representation of a spatiotemporal intelligent large-scale model, generating results of semantic understanding and inference, and obtaining contextually relevant and highly relevant answers. In other words, in application scenarios of multimedia information processing, problems are understood and their meaning analyzed based on spatiotemporal data, multimedia input is transformed into a spatiotemporal object representation with spatiotemporal relationships, the spatiotemporal intelligent large-scale model analyzes the spatiotemporal features of the problem, and answers that match the spatiotemporal context are generated. By combining multimodal data, inference results with spatiotemporal accuracy are provided in the question-and-answer process, thereby ensuring the spatiotemporal consistency of the answers. In the case of time-dependent question-and-answer tasks, the spatiotemporal intelligent large-scale model can estimate possible future situations based on historical spatiotemporal data and provide predicted answers.
[0062] In autonomous driving scenarios, spatiotemporal data of each object in the environment is processed and analyzed in real time based on a spatiotemporal intelligence large-scale model. The spatiotemporal intelligence large-scale model predicts the movement trajectories of surrounding objects, and this information is combined with current spatiotemporal data to generate route planning and driving decisions. Multimodal data is used to comprehensively perceive the environment around the vehicle, analyze driving scenes with spatiotemporal relationships, and support autonomous driving decision-making in complex road conditions and dynamic environments. In the autonomous driving process, it is also possible to improve driving safety and responsiveness by predicting potential hazards and adjusting the driving strategy in real time. Applied to autonomous driving systems, it supports driving decision-making by analyzing and processing spatiotemporal data in the environment and recognizing and predicting the movement trajectories and state changes of surrounding objects.
[0063] In intelligent robot scenarios, large-scale spatiotemporal intelligence models (SUIs) can be used to process multimodal spatiotemporal data in the robot's environment in real time, enhancing the robot's ability to perceive environmental changes. SUIs enables robot path planning and autonomous decision-making, allowing robots to accurately determine spatiotemporal relationships of objects in complex environments and perform rational actions and reactions. By using SUIs to predict actions, robots can understand trends in spatiotemporal changes of surrounding objects, supporting task execution and dynamic adjustments. When performing tasks, robots make real-time adjustments through SUIs, improving autonomy and task execution accuracy by ensuring the decision-making process adapts to the spatiotemporal characteristics of objects in the environment. Processing and analyzing multimodal environmental data in real time can improve the robot's perception and autonomous decision-making capabilities in complex spatiotemporal environments.
[0064] In industrial engineering applications, the spatiotemporal intelligence large-scale model can intelligently process graphics and documents related to multidimensional engineering themes. In short, the general-purpose artificial spatiotemporal intelligence model according to the present invention can be successfully applied to any field and scenario requiring the use of the above-mentioned spatiotemporal intelligence functions.
[0065] As described above, the method for constructing a general-purpose artificial spatiotemporal intelligence large-scale model according to the present invention establishes a unified structured topological representation and storage of various multimedia information such as text, graphics, images, audio, and video in the real world of human society, based on object representation and spatial relationship models using 0-dimensional, 1-dimensional, 2-dimensional, and 3-dimensional object systems. Using a multimodal large-scale model training method based on 3D characterization, it constructs an artificial spatiotemporal intelligence large-scale model that can collectively represent and process all objects in human society. By combining spatiotemporal information processing, spatiotemporal relationship analysis, knowledge graphs, and business flow logical reasoning, it provides capabilities such as spatiotemporal relationship representation, semantic understanding, environmental perception, spatiotemporal reasoning, and spatial decision-making. It solves the problems of conventional large-scale language models, which have a single representation, lack topological relationships, and have difficulty dealing with problems of time and space relationships in the 3D real world. It can be widely applied to fields such as multimedia information processing, theoretical model derivation, autonomous driving, intelligent robots, and industrial engineering, and has high practicality.
[0066] While preferred embodiments of the present invention have been described, those skilled in the art, knowing the basic creative concepts, can make additional changes and modifications to these embodiments. Therefore, the appended claims are intended to be construed as encompassing all changes and modifications that fall within the scope of the preferred embodiments and embodiments of the present invention.
[0067] In this specification, relational terms such as those in the first and second paragraphs are used solely to distinguish one entity or operation from another, and do not necessarily require or imply that such an actual relationship or order exists between these entities or operations. Furthermore, the terms “includes,” “incorporates,” or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device containing a set of elements includes not only those elements but also other elements not expressly described, or elements specific to such a process, method, article, or terminal device. Unless further limited, an element defined by the phrase “includes…” does not preclude the existence of other identical elements in a process, method, article, or terminal device containing that element.
[0068] Although embodiments of the present invention have been described above with reference to the drawings, the present invention is not limited to the above-described specific embodiments. The above-described specific embodiments are schematic and not limiting, and those skilled in the art can create many more forms without departing from the spirit and claims of the present invention, under the guidance of the present invention, and all of these are within the scope of the protection of the present invention.
Claims
1. A step of establishing a structured spatiotemporal object representation and data storage that can be processed collectively by a spatiotemporal intelligence large-scale model, using a multidimensional object abstraction representation scheme to process various objects and spatial relationships in the real world, wherein the multidimensional object abstraction refers to zero-dimensional objects, one-dimensional objects, two-dimensional objects, and three-dimensional objects. The steps include processing the aforementioned spatiotemporal object with a multimodal spatiotemporal coding algorithm and converting the spatiotemporal data corresponding to the aforementioned spatiotemporal object into multimodal spatiotemporal coded data for subsequent training and application of spatiotemporal intelligence large-scale models, The steps include creating a multimodal spatiotemporal training dataset based on the aforementioned multimodal spatiotemporal coded data, training or fine-tuning it using a spatiotemporal multimodal training method based on three-dimensional characterization, and forming the aforementioned spatiotemporal intelligent large-scale model. The steps include: issuing the trained spatiotemporal intelligence large-scale model as a spatiotemporal intelligence large-scale model inference service and providing spatiotemporal intelligence capabilities, wherein the spatiotemporal intelligence capabilities include spatiotemporal relational representation, semantic understanding, environmental perception, spatiotemporal reasoning, and spatial decision-making; The various objects of the real world include actual geospatial objects in the real world, virtual geospatial objects, and multimedia type objects in relative space specific to the multimedia context, and the multimedia type objects include various types such as text, graphics, images, audio, and video. The 0-dimensional object is an objectification modeling representation of a fundamental unit of the real world, where the object of the fundamental unit refers to various physical objects and multimedia information objects that have spatial locations in the real world, where the multimedia information objects include text, graphics, images, audio, and video, where the various physical objects are represented by the spatial location, environment, and spatial relationships of the object, where the multimedia information objects are represented by the numerical features, scene or context, and relative space of the object, including, but not limited to, the sequence position of a character within a paragraph and the correlation relationship between other characters, graphics vector nodes, and the sequence position of an image frame within a video and the correlation relationship between other frames. The aforementioned one-dimensional object is a modeling representation of the process of change and evolution of a real-world object in the time dimension, representing the process or trajectory of a continuous change of an object in the time and spatial dimensions as a one-dimensional object, or including a data sequence of zero-dimensional objects, representing multiple zero-dimensional objects of the same kind as one-dimensional objects, and recording their flow process, logical relationships, and topological adjacencies in spacetime, the data sequence includes, but is not limited to, coordinate sequences, character sequences, text paragraphs, mathematical formulas, graphics vector arc segments, image sets, and video frames. The aforementioned two-dimensional object is a modeling representation of the adjacency and aggregation of objects in the real world, representing one-dimensional objects of the same type as two-dimensional objects, recording their temporal and spatial adjacency relationships, and using them for broader data analysis and related modeling. When processing during the adjacency and aggregation of objects in the real world, this includes fully recording the characteristics of the base unit itself and complex derivation processes or combinations and associations of multiple data, depending on the object's position, type, or transformation process in the temporal or spatial dimension. A method for constructing a general-purpose artificial spatiotemporal intelligence large-scale model, characterized in that the three-dimensional object is a modeling representation of a large dataset or the overall spatiotemporal evolution of a complex system in a real-world object, and includes displaying a big dataset of multidimensional zero-dimensional objects, one-dimensional objects, and two-dimensional objects in a complex or complex process, recording its overall relationships or transitional processes in spatiotemporaneity, and combining objects and events of multiple spatiotemporal dimensions to generate an overall chronotope structure that supports comprehensive spatiotemporal analysis and prediction of a complex system.
2. The step of processing various objects and spatial relationships in the real world using a multidimensional object abstraction representation scheme is: The process includes the steps of collecting, recognizing, and analyzing data of various real-world objects, including but not limited to multimedia information formats; obtaining spatiotemporal information features of the various objects in the multimedia information formats; and generating multimodal structured spatiotemporal object representations of zero-dimensional, one-dimensional, two-dimensional, and three-dimensional objects by combining the environment, scene, or context in which the various objects exist, relative space, and knowledge graphs and spatial relationships of the various objects, wherein the multimedia information formats include text formats, graphics formats, image formats, audio formats, and video formats, and the spatiotemporal information features include location, form, state, and tense. The construction method according to claim 1, characterized in that the spatiotemporal object representation simultaneously possesses position, tense, and spatial relational features, ensuring consistency and relevance of objects in the temporal and spatial dimensions, and the spatiotemporal objects are represented by topological relationships, orientation or order relationships, and metric relationships, wherein the topological relationships include association, adjacency, and inclusion, the orientation or order relationships include up, down, left, right, and east, west, north, and south, and the metric relationships include distance, angle, or proximity.
3. The step of processing with a multimodal spatiotemporal coding algorithm and converting the spatiotemporal data corresponding to the spatiotemporal object into multimodal spatiotemporal coded data for subsequent spatiotemporal intelligence large-scale model training and application is: Step T1 involves constructing a spatiotemporal topology diagram structure by considering each basic object in the spatiotemporal data as a node in a spatiotemporal topology diagram based on the aforementioned spatiotemporal objects, defining spatiotemporal proximity between nodes as edges in the diagram, calculating spatiotemporal distance or correlation between objects, and defining edge weights between nodes, wherein the correlation between objects is determined by similarity, neighborliness, or time-series order in the spatiotemporal data, and the spatiotemporal data refers to multimodal spatiotemporal multidimensional data. The construction method according to claim 1, comprising step T2 of encoding data based on a graph neural network, constructing a pre-trained task by spatial association, and representing the spatiotemporal topology diagram structure in a high-dimensional manner.
4. The steps of encoding data based on a graph neural network, constructing a pre-trained task through spatial association, and representing the spatiotemporal topology diagram structure in high dimensions are: The steps include: updating the spatiotemporal characteristics of each node by aggregating information on nodes and neighboring nodes in a spatiotemporal topology diagram, and capturing the spatial and temporal relationships between the multimodal spatiotemporal multidimensional data; The construction method according to claim 3, comprising the step of deep learning the multimodal spatiotemporal multidimensional data using a multilayer graph neural network to generate an encoded representation that enhances spatiotemporal relationships for further model training.
5. The step of training or fine-tuning the spatiotemporal intelligence large-scale model using a spatiotemporal multimodal training method based on three-dimensional characterization is: Step N1 involves collecting or generating data containing spatiotemporal information based on the multimodal spatiotemporal encoded data, combining specific application scenes and task requirements, creating a multimodal spatiotemporal training dataset, the training dataset containing labels and spatiotemporal attribute information of the multimodal data, guiding the adaptability and accuracy of the spatiotemporal intelligent large-scale model in specific tasks, and training the spatiotemporal-related capabilities of the spatiotemporal intelligent large-scale model. The construction method according to claim 1, comprising step N2 of training a spatiotemporal intelligence large-scale model using a spatiotemporal multimodal training method with a multidimensional spatiotemporal attention network model as input to a constructed multimodal spatiotemporal training dataset, and capturing time sequence relationships and spatial features annotated with 0-dimensional objects, 1-dimensional objects, 2-dimensional objects, and 3-dimensional objects from the data.
6. The spatiotemporal multimodal training method refers to spatiotemporal multimodal training of 0-dimensional objects, 1-dimensional objects, 2-dimensional objects, and 3-dimensional objects with 3D characterization + time dimension, characterized in that the 3D characterization method is adopted in the base layer of the spatiotemporal intelligent large-scale model, and the time sequence and spatial features of multimodal spatiotemporal multidimensional data are combined and modeled based on the 3D spatial representation and time dimension, thereby more accurately capturing complex relationships in time and space, reflecting spatiotemporal dynamic changes in the real world, and improving the spatiotemporal inference representation of the spatiotemporal intelligent large-scale model.
7. When creating the aforementioned multimodal spatiotemporal training dataset, in order to ensure the robustness and generalization ability of the spatiotemporal intelligence large-scale model under various conditions, the main objectives are to acquire data diversity and coverage. The construction method according to claim 5, characterized in that the text data includes various types of content, including but not limited to natural language descriptions, commands, and dialogues, covering a broad range of meanings and various levels of complexity; the image data includes but not limited to still and moving images from various viewpoints, various environments, and various time periods; and the geospatial data includes but not limited to spatial data in various industries and formats of varying precision, and constructs spatial relationships relating to spatiotemporal data.
8. The construction method according to claim 1, characterized in that the spatiotemporal intelligence capability provided by the spatiotemporal intelligence large-scale model inference service outputs, based on the spatial relationships of spatiotemporal object representations, using topological relationship relationships between spatiotemporal objects recorded by the spatiotemporal objects, including a knowledge graph and business flow logic about the objects, in the inference process of the spatiotemporal intelligence large-scale model, thereby reducing the computational complexity of the inference by the spatiotemporal intelligence large-scale model and efficiently performing the inference task.
9. The spatiotemporal relationship representation includes a topological spatial relationship, a sequential spatial relationship, a metric relationship, and a time mark or timestamp, the topological spatial relationship includes association, adjacency, and inclusion, the sequential spatial relationship includes orientation and front / back / left / right, and the metric relationship includes length and angle. The aforementioned semantic comprehension ability involves analyzing and recognizing spatiotemporal objects through vision information processing or digitized input, and outputting spatiotemporal objects with 0-dimensional, 1-dimensional, 2-dimensional, and 3-dimensional object features, as well as spatial relationships between objects and between objects and the environment. The aforementioned environmental perception ability is to understand and predict the relative relationships and trends of change between spatiotemporal objects based on the spatiotemporal object representation, and to provide support for further spatiotemporal relationship inference applications. The aforementioned spatiotemporal reasoning ability, based on the aforementioned environmental perception ability, understands the space in which objects exist in the real world, the temporal environment, or the contextual scene of multimedia information, or the virtual environment, and infers and analyzes the relationships between objects to generate trends in spatial object features and relationships of future temporal features. The construction method according to claim 8, characterized in that the spatial decision-making capability is realized by combining specific spatiotemporal scenes and business flows based on semantic understanding, environmental perception, and spatiotemporal reasoning, and by spatiotemporal reasoning analysis or business flow learning to realize the application of spatial decision-making in spatiotemporal scenes.
Citation Information
Patent Citations
Multi-modal knowledge graph construction method
CN112200317A
Scene-aware video dialogue
JP2023510430A
Multidimensional Deep Neural Networks
JP2023539954A
Method and system for intelligent analysis of bills based on semantic graph model
JP7579022B1