Methods for constructing large-scale general-purpose artificial spatiotemporal intelligence models
The construction of a general-purpose artificial spatiotemporal intelligence model addresses the limitations of large-scale language models by incorporating multidimensional object abstraction and spatiotemporal training, enhancing decision-making and adaptability in complex environments.
Patent Information
- Application Number
- JP2025026489
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2024-10-15
- Filing Date
- 2025-02-21
- Publication Date
- 2026-01-15
- Estimated Expiration
- 2045-02-21
AI Technical Summary
Large-scale language models lack the capability to effectively integrate multidimensional spatiotemporal interactions, making them unsuitable for applications requiring accurate modeling and prediction of spatiotemporal changes in the real world, such as autonomous driving and robotics, due to their one-dimensional data structure and lack of spatial and temporal relationship representations.
A method for constructing a general-purpose artificial spatiotemporal intelligence large-scale model using multidimensional object abstraction, multimodal space-time encoding, and spatiotemporal multimodal training to process and represent spatiotemporal data, enabling spatiotemporal relationship representation, semantic understanding, environmental perception, and spatial decision-making.
The model provides enhanced decision-making accuracy and adaptability in complex scenes by integrating spatiotemporal information, supporting applications like multimedia processing, autonomous driving, intelligent robots, and industrial engineering.
Smart Images

Figure 0007799295000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to the technical field of artificial intelligence, and in particular to a method for constructing a general-purpose artificial spatiotemporal intelligence large-scale model. [Background technology]
[0002] In recent years, artificial intelligence technology has developed rapidly. Large-scale language models (LLMs) have achieved remarkable results in fields such as natural language processing, speech recognition, and computer vision due to their powerful representation learning and generalization capabilities. They have become a popular research and development topic in the field of artificial intelligence and are widely used in industries such as literature, art, education, and finance, having a significant impact on the development of human society.
[0003] Currently, the number of parameters in large-scale language models typically reaches billions or even hundreds of billions, placing extremely high demands on computing resources and training techniques. Training such large-scale models requires training, acquiring, cleaning, and annotating large amounts of data, which is a costly and time-consuming process that quickly leads to bottlenecks in computing and data resources.
[0004] Large-scale language models essentially process linguistic text data and can also process multimedia data such as images and videos using encoding methods. However, their inherently one-dimensional data structure is relatively simple: they are tokenized text sequences and lack spatial relationship representations. Furthermore, large-scale models are primarily trained on text data and lack experience in direct interactions with the physical world. Therefore, they cannot learn spatial, temporal, and physical interactions through their own experiences, as humans do. They are therefore unable to effectively integrate multidimensional time and space information from the real world, making it difficult to support the perception, understanding, and manipulation of three-dimensional space. This makes them unsuitable for applications that require the processing of multidimensional spatiotemporal interactions, such as industrial manufacturing and civil engineering and construction. For example, autonomous driving and robotics applications require accurate modeling and prediction of spatiotemporal changes in various objects in the environment. Large-scale language model processing methods, however, are based on processing one-dimensional data structures, requiring extremely high computational power. Furthermore, the lack of a framework for analyzing both the time and spatial dimensions together makes it difficult to provide accurate and consistent spatiotemporal relationship analysis results.
[0005] Currently, large-scale language models are primarily used in intelligent processing of text, images, and videos, i.e., "humanities intelligence." They are not suitable for processing data in science and engineering, particularly for processing and interpreting spatiotemporal data in the field of humanities and social engineering, such as the intelligent processing of multidimensional design graphics for various projects. Graphic design and applications in the engineering field are issues that must be addressed in an industrialized and intelligent society; otherwise, the application of artificial intelligence will remain partial and incomplete.
[0006] The human real world is composed of three-dimensional physical objects or multimedia information objects with time and space characteristics. The three-dimensional world follows the laws of physics and has its own unique structure, attributes, and time-space relationships. Due to the inherent limitations of existing one-dimensional data structure principles and data modeling methods in large-scale language models, there is an urgent need to find a unified and universal method for representing and processing large-scale models. Research is currently being conducted on large-scale models that are more suitable for representing, perceiving, reasoning, and interactive decision-making in the real universe. Such models can comprehensively process multimedia data, such as text, graphics, images, audio, and video, facing human society, and address a wide range of problems in the humanities, sciences, and engineering. Summary of the Invention [Problem to be solved by the invention]
[0007] In view of the above problems, the present invention proposes a method for constructing a general-purpose artificial spatiotemporal intelligence large-scale model. [Means for solving the problem]
[0008] The present invention includes the steps of: using a multidimensional object abstraction representation method to process various objects and spatial relationships in the real world, and establishing a structured spatiotemporal object representation and data storage that can be collectively processed by a spatiotemporal intelligence large-scale model, where the multidimensional object abstraction refers to a 0-dimensional object, a 1-dimensional object, a 2-dimensional object, and a 3-dimensional object; processing the spatio-temporal objects with a multimodal space-time coding algorithm to convert the spatio-temporal data corresponding to the spatio-temporal objects into coded data for subsequent training and application of a spatio-temporal intelligence large-scale model; Creating a multimodal spatiotemporal training dataset based on the multimodal spatiotemporal coding data, and using a spatiotemporal multimodal training method with three-dimensional characterization to train or fine-tune the spatiotemporal intelligent large-scale model; A method for constructing a general-purpose artificial spatio-temporal intelligence large-scale model is disclosed, which includes: publishing the trained spatio-temporal intelligence large-scale model as a spatio-temporal intelligence large-scale model inference service to provide spatio-temporal intelligence capabilities, wherein the spatio-temporal intelligence capabilities include spatio-temporal relationship representation, semantic understanding, environmental perception, spatio-temporal reasoning, and spatial decision-making.
[0009] Optionally, the step of processing various objects and spatial relationships in the real world using a multidimensional object abstraction representation scheme comprises: collecting, recognizing and analyzing data of various objects in the real world, including but not limited to multimedia information formats, obtaining spatiotemporal information features of the various objects in the multimedia information formats, and combining the environment, scene or context in which the various objects exist, relative space, and knowledge graphs and spatial relationships of the various objects to generate multimodally structured spatiotemporal object representations of 0-dimensional objects, 1-dimensional objects, 2-dimensional objects and 3-dimensional objects, wherein the multimedia information formats include text formats, graphics formats, image formats, audio formats and video formats, and the spatiotemporal information features include position, form, state and tense; The spatiotemporal object representation simultaneously possesses location, tense, and spatial relationship characteristics, ensuring the consistency and association of objects in the temporal and spatial dimensions, and the spatiotemporal objects are represented by topological relationships, directional or order relationships, and metric relationships, where the topological relationships include association, adjacency, and containment, the directional or order relationships include up, down, left, right, east, west, north, south, and the metric relationships include distance, angle, or perspective.
[0010] Optionally, the various real-world objects include actual geospatial objects in the real world, virtual geospatial objects, and multimedia type objects of relative space specific to a multimedia context, the multimedia type objects including types such as text, graphics, images, audio, and video; The zero-dimensional object is an object modeling representation of a basic unit of the real world, and the basic unit object refers to various physical objects and multimedia information objects that have spatial positions in the real world, and the multimedia information objects include text, graphics, images, audio, and video, and the various physical objects are represented by the spatial positions, environments, and spatial relationships of the objects, and the multimedia information objects are represented by the numerical features, scenes or contexts, and relative spaces of the objects, including, but not limited to, the sequence position of a character in a paragraph and its correlation with other characters, graphics vector nodes, and the sequence position of an image frame in a video and its correlation with other frames; The one-dimensional object is a modeling representation of the change and evolution process of a real-world object in the time dimension, and includes representing the process or trajectory of continuous change of an object in the time dimension and the space dimension as a one-dimensional object, or a data sequence of a zero-dimensional object, representing multiple zero-dimensional objects of the same type as a one-dimensional object, and recording their flow process, logical relationship, and topological adjacent relationship in time and space, and the data sequence includes, but is not limited to, a coordinate sequence, a character sequence, a character paragraph, a mathematical formula, a graphic vector arc segment, an image set, and a video frame; The two-dimensional object is a modeling representation of real-world object adjacencies and object aggregation, and represents the same type of one-dimensional object as a two-dimensional object, records their temporal and spatial adjacency relationships, and uses them for more extensive data analysis and association modeling. When processing real-world object adjacencies and object aggregation, the characteristics of the basic unit itself and the complex derivation process or the combination and association characteristics of multiple data are completely recorded according to the position, type, or change process of the object in the time or spatial dimension. The three-dimensional object is a modeling representation of a large-scale dataset of real-world objects or the overall spatiotemporal evolution of a complex system, including displaying a big dataset of multidimensional zero-dimensional objects, one-dimensional objects, and two-dimensional objects in a complex or composite process, recording their overall relationship or transition process in space-time, and combining objects and events of multiple spatiotemporal dimensions to generate an overall chronotope structure that supports comprehensive spatiotemporal analysis and prediction of complex systems.
[0011] Optionally, the step of processing with a multimodal space-time encoding algorithm to convert the space-time data corresponding to the space-time object into encoded data for subsequent training and application of a space-time intelligence large-scale model includes: Step T1: based on the spatio-temporal objects, construct a spatio-temporal topology diagram structure by regarding each basic object in the spatio-temporal data as a node of a spatio-temporal topology diagram, the spatio-temporal adjacency between each node as an edge of the diagram, calculating the spatio-temporal distance or correlation between objects, and defining the edge weight between nodes, where the correlation between objects is determined by similarity, proximity, or chronological order in the spatio-temporal data, and the spatio-temporal data refers to multi-modal spatio-temporal multidimensional data; and step T2 of encoding data based on a graph neural network, constructing a pre-training task by spatial association, and expressing the spatio-temporal topological diagram structure in high dimensions.
[0012] Optionally, the step of encoding data based on a graph neural network, constructing a pre-training task by spatial association, and expressing the spatio-temporal topological diagram structure in a high dimension includes: updating the spatiotemporal features of each node by aggregating information of the nodes and neighboring nodes in the spatiotemporal topology diagram, and capturing the space-time associations between the multimodal spatiotemporal multidimensional data; and deep learning the multi-modal spatio-temporal multi-dimensional data with a multi-layer graph neural network to generate a coded representation that reinforces spatio-temporal associations for further model training.
[0013] Optionally, training or fine-tuning to form the spatio-temporal intelligent large-scale model using a spatio-temporal multimodal training method with three-dimensional characterization includes: Step N1: based on the multimodal space-time coding data, by combining specific application scenarios and task requirements, collect or generate data containing space-time information, and create a multimodal space-time training dataset, the training dataset including labels and time-space attribute information of multimodal data, to guide the adaptability and accuracy of the space-time intelligence large-scale model in specific tasks, and train the space-time related ability of the space-time intelligence large-scale model; Step N2 includes using the constructed multimodal spatiotemporal training dataset as input, training a spatiotemporal intelligence large-scale model using a multidimensional spatiotemporal attention network model through a spatiotemporal multimodal training method, and capturing the time sequence relationships and spatial features annotated with 0-dimensional objects, 1-dimensional objects, 2-dimensional objects, and 3-dimensional objects in the data.
[0014] Optionally, the spatiotemporal multimodal training method refers to the spatiotemporal multimodal training of 0-dimensional objects, 1-dimensional objects, 2-dimensional objects, and 3-dimensional objects with 3-dimensional characterization + time dimension, in which a 3-dimensional characterization method is adopted in the foundation layer of the spatiotemporal intelligent large-scale model, and based on 3-dimensional spatial representation and time dimension, the time sequence and spatial features of the multimodal spatiotemporal multidimensional data are combined and modeled, so as to more accurately capture the complex relationships in time and space, reflect the spatiotemporal dynamic changes of the real world, and improve the expression in the spatiotemporal inference of the spatiotemporal intelligent large-scale model.
[0015] Optionally, when creating the multimodal spatiotemporal training dataset, data diversity and coverage are primarily obtained to ensure the robustness and generalization ability of the spatiotemporal intelligence large-scale model in various situations, including text data with various types of content, including but not limited to natural language descriptions, instructions, and dialogues, covering a wide range of semantic ranges and various levels of complexity; image data with but not limited to still and moving images from various perspectives, environments, and time periods; and geospatial data with but not limited to spatial data from various industries and formats with various precision, and to build spatial relationships related to the spatiotemporal data.
[0016] Optionally, the spatiotemporal intelligence capability of the spatiotemporal intelligence large-scale model inference service uses the topological association relationships between objects recorded by the spatiotemporal objects, including knowledge graphs and business flow logic related to the objects, to output in the inference process of the spatiotemporal intelligence large-scale model based on the spatial relationships of the spatiotemporal object representations, thereby reducing the computational complexity of inference by the spatiotemporal intelligence large-scale model and efficiently performing the inference task.
[0017] Optionally, said spatiotemporal relation representations comprise topological spatial relations, ordinal spatial relations, metric relations, and time marks or time stamps, said topological spatial relations comprise association, adjacency, containment, said ordinal spatial relations comprise orientation, and front-back, left-right, and said metric relations comprise length, and angle; The semantic understanding ability is to analyze and recognize spatiotemporal objects through vision information processing or digital input, and output spatiotemporal objects having 0-dimensional object, 1-dimensional object, 2-dimensional object, and 3-dimensional object characteristics, as well as spatial relationships between objects and between objects and the environment; The environmental perception ability is to understand and predict the relative relationship and change trend between spatiotemporal objects based on the spatiotemporal object representation, and provide support for further spatiotemporal relationship reasoning applications; The spatiotemporal reasoning ability is to understand the space in which objects exist in the real world, the temporal environment or the context scene of multimedia information, and the virtual environment based on the environmental perception ability, and to infer and analyze the relational relationships between objects to generate the tendency of spatial object features and relational relationships of future temporal features; The spatial decision-making ability is based on the semantic understanding, the environmental perception, and the spatiotemporal reasoning, and combines specific spatiotemporal scenes and workflows to realize the application of spatial decision-making in spatiotemporal scenes through spatiotemporal reasoning analysis or workflow learning. [Effects of the Invention]
[0018] The method for constructing a general-purpose artificial spatio-temporal intelligent large-scale model of the present invention first uses a multidimensional object abstraction representation method to process various objects and spatial relationships in the real world and establish a structured spatio-temporal object representation and data storage that can be collectively processed by the spatio-temporal intelligent large-scale model. The spatio-temporal objects are then processed using a multimodal spatio-temporal encoding algorithm, and the spatio-temporal data corresponding to the spatio-temporal objects is converted into coded data for subsequent training and application of the spatio-temporal intelligent large-scale model. A multimodal spatio-temporal training dataset is then created based on the multimodal spatio-temporal coded data, and a spatio-temporal multimodal training method with 3D characterization is used to train or fine-tune the spatio-temporal intelligent large-scale model. Finally, the trained spatio-temporal intelligent large-scale model is published as a spatio-temporal intelligent large-scale model inference service to provide spatio-temporal intelligence capabilities.
[0019] The present invention proposes a method for constructing a general-purpose artificial spatiotemporal intelligence model (which can be defined as Artificial General Spatiotemporal Intelligence, abbreviated as AGSTI). This construction method establishes a unified, structured topological representation and storage of various multimedia information in the real world of human society, such as various texts, graphics, images, audio, and videos, based on object representation methods and spatial relationship models for 0-dimensional, 1-dimensional, 2-dimensional, and 3-dimensional spatial objects. Using a multimodal large-scale model training method with three-dimensional characterization, an artificial spatiotemporal large-scale model is constructed that can collectively represent and process all objects in human society. By combining spatiotemporal information processing, spatiotemporal relationship analysis, knowledge graphs, and workflow logical inference, the model provides capabilities such as spatiotemporal relationship representation, semantic understanding, environmental perception, spatiotemporal reasoning, and spatial decision-making, thereby resolving the challenges posed by conventional large-scale language models, which have a single representation, a lack of topological relationships, and difficulty in dealing with time-space relationship problems in the 3D real world. The general-purpose artificial spatiotemporal intelligent large-scale model constructed according to the present invention can greatly improve the decision-making accuracy and adaptability of intelligent systems in complex scenes, and can be fully applied in application fields such as multimedia information processing, theoretical model derivation, autonomous driving, intelligent robots, and industrial engineering, providing new technical support for these application fields. [Brief explanation of the drawings]
[0020] Various other benefits and advantages will become apparent to those skilled in the art upon reading the following detailed description of the preferred embodiments. The drawings are for purposes of illustrating the preferred embodiments and are not intended to limit the invention. In addition, like elements are designated by like reference numerals throughout the drawings. [Figure 1] 1 is a flowchart of a method for constructing a general-purpose artificial spatiotemporal intelligence large-scale model according to an embodiment of the present invention; [Figure 2] FIG. 2 is a schematic diagram illustrating an example of a real-world object abstraction representation method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0021] In order to make the above-mentioned objects, features and advantages of the present invention clearer and easier to understand, the present invention will be described in more detail with reference to drawings and specific embodiments. It should be understood that the specific embodiments described in this specification are only used to explain the present invention, and are only a part of the embodiments of the present invention, not all of the embodiments of the present invention, and are not used to limit the present invention.
[0022] The method for constructing a general-purpose artificial spatio-temporal intelligent large-scale model according to the present invention includes the following steps 101 to 104, with reference to the flowchart of the method for constructing a general-purpose artificial spatio-temporal intelligent large-scale model shown in FIG.
[0023] Step 101: Use a multidimensional object abstraction representation method to process various objects and spatial relationships in the real world, and establish a structured spatiotemporal object representation and data storage that can be collectively processed by the spatiotemporal intelligence large-scale model, where multidimensional object abstraction refers to 0-dimensional objects, 1-dimensional objects, 2-dimensional objects, and 3-dimensional objects.
[0024] The present invention innovatively proposes a theory of multidimensional object abstraction, which includes 0-dimensional objects, 1-dimensional objects, 2-dimensional objects and 3-dimensional objects. In order to better understand the technical solution of the present invention, the theory of multidimensional object abstraction will first be described in detail.
[0025] The multidimensional object abstraction of this invention involves analyzing various real-world objects in human society, as well as data such as text, graphics, images, audio, and video obtained by recognizing various objects, and abstracting them into structured 0-dimensional, 1-dimensional, 2-dimensional, and 3-dimensional objects. The 3-dimensional feature model representation and training method is used to overcome the limitations of large-scale language models based on text representation at the data model representation and training level, and optimize the computational power requirements for model training and inference.
[0026] The so-called various objects in the real world include actual geospatial objects in the real world, virtual geospatial objects, and multimedia type objects of relative space specific to the multimedia context, which include each type of multimedia type object, such as text, graphics, image, audio, video, etc. To better understand the above theory, some concrete examples will be combined and described.
[0027] (1) 0-dimensional objects: Basic units of the real world, such as vector nodes, the number "1," the letter "A," the Chinese character "hair," and in the case of physical entities, such as "one tree, one rose, one person, one tiger, one image, one snapshot, auditory and olfactory information at one point in time, or one part or building component," as well as mathematical operation symbols such as "the sum symbol Π, the synthesis symbol Σ, or the sine symbol sin." In the general-purpose artificial spatiotemporal intelligence large-scale model of the present invention, these are all involved in the representation and training of the model in the form of 0-dimensional objects.
[0028] In combination with the above example, zero-dimensional objects can be understood as objectified modeling representations of basic units in the real world. Here, basic unit objects refer to various physical objects with spatial positions in the real world and multimedia information objects, including text, graphics, images, audio, and video. Spatial positions include the absolute and local coordinates of various objects, as well as the relative sequence number, page number, and relative coordinates of the scene and context in which the objects are located. Various physical objects are represented by their spatial positions, environment, and spatial relationships. Multimedia information objects are represented by the object's numerical characteristics, scene or context, and relative space, including, but not limited to, the sequence position of a character within a paragraph, the page number and correlation with other characters within a chapter, and the sequence position of an image frame within a video and correlation with other frames.
[0029] (2) One-dimensional objects: vector arch segments, a row of numbers or letters, a phrase or a sentence, physical entities such as "a row of trees, a row of roses, a row of people or a person walking or running, a row of tigers or a tiger walking or running, a segment of video, a segment of auditory or olfactory information," a subsystem formed by multiple parts or building components, a local calculation formula, etc. In the general-purpose artificial spatiotemporal intelligence large-scale model of the present invention, all of these participate in the representation and training of the model in the form of one-dimensional objects.
[0030] In combination with the above examples, a one-dimensional object can be understood as a representation of the change and evolution process of a real-world object in time, representing the process or trajectory of the object's continuous change in time and space as a one-dimensional object, or a data sequence of zero-dimensional objects, representing multiple zero-dimensional objects of the same type as one-dimensional objects, combining topological node features to record their flow process, logical relationships, and topological neighbor relationships in time and space. Data sequences include, but are not limited to, coordinate sequences, character sequences, character paragraphs, mathematical formulas, image sets, and video frames.
[0031] (3) Two-dimensional objects: vector polygons, polynomial or multi-segment data or strings, physical entities such as "a forest, a rose bush, a group of people, a herd of tigers, a complete film epic or a complete television video, a set of staged auditory and olfactory information," fully functional machines or buildings, complete computational or derivation processes (combinations of multiple equations), etc.
[0032] Combining the above examples, a 2D object is a modeling representation of real-world object adjacencies and object aggregations. Multiple 0D and 1D objects of the same type are represented as 2D objects, and their temporal and spatial adjacency relationships are recorded for use in broader data analysis and spatiotemporal object modeling. When processing real-world object associations and object aggregations, it involves fully recording the characteristics of the basic unit itself and complex derivation processes or the combination and association characteristics of multiple data according to the object's location, type, or change process in the time or space dimension.
[0033] (4) Three-dimensional objects: vector polyhedra, data sets or text sets, parks or communities, zoos, physical entities such as "a field, forest, or grassland, a film or television series," workshops or factories, buildings, a monograph on mathematics, physics, or chemistry, etc.
[0034] Combining the above examples, a 3D object is a modeling representation of a large dataset of real-world objects or the overall spatiotemporal evolution of a complex system, displaying a big dataset of multidimensional 0D, 1D, and 2D objects in a complex or composite process, recording their overall relationships or transitions in space and time. It involves combining objects and events from multiple spatiotemporal dimensions to generate an overall chronotope structure that supports comprehensive spatiotemporal analysis and prediction of complex systems.
[0035] The above multidimensional object abstraction theory can be more intuitively understood when combined with a schematic diagram of an abstract representation of an exemplary real-world object, as shown in Figure 2. Depending on their type, physical objects within the real world (such as the exemplary number "1" on the left side of Figure 2, a column of numbers, multiple segments of data, a data set, or a text set) can be represented as lumped spatiotemporal representations, i.e., 0-dimensional objects, 1-dimensional objects, 2-dimensional objects, and 3-dimensional objects, which correspond to vector nodes, arc segments, polygons, and polyhedra, have temporal attributes, and have multiple relationships such as spatial, topological, order, and metric relationships.
[0036] The theory of multidimensional object abstraction has been explained in detail above. To construct a general-purpose artificial spatiotemporal intelligence model (which can be defined as Artificial General Spatiotemporal Intelligence, abbreviated as AGSTI) according to the present invention, first, a multidimensional object abstraction representation method is used to process various objects and spatial relationships in the real world, and a structured spatiotemporal object representation and data storage that can be uniformly processed by a large-scale spatiotemporal intelligence big model is established. A preferred method includes the following:
[0037] First, data of various objects in the real world, including but not limited to multimedia information formats, is collected, then spatiotemporal information features of the various objects in the obtained multimedia information formats are recognized and analyzed, and the environment, scene or context in which the various objects exist, relative space, and knowledge graph associations and spatial relationships of the various objects are combined to generate multimodally structured spatiotemporal object representations of 0-dimensional objects, 1-dimensional objects, 2-dimensional objects, and 3-dimensional objects, where the multimedia information formats include text formats, graphics formats, image formats, audio formats, and video formats, and the spatiotemporal information features include position, form, state, and tense.
[0038] A spatiotemporal object representation simultaneously possesses location, tense, and topological relationship characteristics, thus ensuring the consistency and association of objects in the temporal and spatial dimensions. Spatiotemporal objects may be represented by topological relationships, directional or order relationships, and metric relationships, where topological relationships include association, adjacency, and containment, directional or order relationships include up / down, left / right, and east / west, north / south, and metric relationships include distance, angle, or perspective.
[0039] Step 102: The spatio-temporal object is processed by a multi-modal space-time encoding algorithm, and the spatio-temporal data corresponding to the spatio-temporal object is converted into multi-modal space-time encoded data for subsequent training and application of a spatio-temporal intelligence large-scale model.
[0040] After establishing a structured spatiotemporal object representation and data storage that can be collectively processed by the spatiotemporal intelligence large-scale model, the spatiotemporal object needs to be processed by a multimodal space-time encoding algorithm, thereby converting the spatiotemporal data corresponding to the spatiotemporal object into multimodal space-time encoded data for subsequent training and application of the spatiotemporal intelligence large-scale model. A preferred method includes the following steps:
[0041] Step T1: Based on the spatiotemporal objects, each basic object in the spatiotemporal data is regarded as a node of a spatiotemporal topology diagram, the spatiotemporal adjacency between each node is taken as the edge of the diagram, the spatiotemporal distance or correlation between objects is calculated, and the edge weight between the nodes is defined to construct a spatiotemporal topology diagram structure, where the correlation between objects is determined by similarity, proximity, or chronological order in the spatiotemporal data, and the spatiotemporal data refers to multi-modal spatiotemporal multidimensional data.
[0042] Step T2: Encode the data based on the graph neural network, construct a pre-training task by spatial association, and represent the spatio-temporal topology diagram structure in high dimension.
[0043] With regard to high-dimensional representation, a preferred method includes first updating the spatiotemporal features of each node by aggregating information of the node and neighboring nodes in a spatiotemporal topological diagram to capture the space-time associations among the multimodal spatiotemporal multidimensional data; and then deep learning the multimodal spatiotemporal multidimensional data by a multi-layer graph neural network to generate an encoded representation that strengthens the spatiotemporal associations for further model training.
[0044] Step 103: Based on the multimodal space-time encoding data, a multimodal space-time training dataset is generated, and a space-time multimodal training method based on three-dimensional characterization is used to train or fine-tune the dataset to form a space-time intelligent large-scale model.
[0045] After the spatiotemporal data is converted into multimodal spatiotemporal encoded data for subsequent training and application of a spatiotemporal intelligence large-scale model, a multimodal spatiotemporal training dataset may be created based on the multimodal spatiotemporal encoded data, and then a spatiotemporal multimodal training method with three-dimensional characterization may be used to train or fine-tune a spatiotemporal intelligence large-scale model based on the multimodal spatiotemporal training dataset.
[0046] A preferred method for training or fine-tuning to form a spatio-temporal intelligence large-scale model includes the following steps.
[0047] Step N1: Based on the multimodal spatiotemporal coding data, combine specific application scenarios and task requirements to collect or generate data containing spatiotemporal information, and create a multimodal spatiotemporal training dataset, which includes the labels and time-spatial attribute information of the multimodal data, and guides the adaptability and accuracy of the spatiotemporal intelligence large-scale model in specific tasks, and trains the spatiotemporal association ability of the spatiotemporal intelligence large-scale model.
[0048] Step N2: Using the constructed multimodal spatiotemporal training dataset as input, a multidimensional spatiotemporal attention network model is used to train a spatiotemporal intelligence large-scale model through the spatiotemporal multimodal training method, and the time sequence relationships and spatial features annotated with 0-dimensional objects, 1-dimensional objects, 2-dimensional objects, and 3-dimensional objects in the data are captured.
[0049] The spatiotemporal multimodal training method in step N2 refers to spatiotemporal multimodal training of 0D, 1D, 2D, and 3D objects with 3D characterization + time. Unlike traditional language model and multimodal language model training, the basic layer of traditional model training is usually based on sequence length and attention mechanisms, and uses tokenized 1D representations. The training concepts and architectures of the two are completely different.
[0050] The three-dimensional characterization method is adopted in the foundation layer of the spatiotemporal intelligent large-scale model according to the present invention, and based on the three-dimensional spatial representation and time dimension, the time sequence and spatial features of multi-modal spatiotemporal multidimensional data are combined and modeled to more accurately capture the complex relationships in time and space, reflect the spatiotemporal dynamic changes of the real world, and improve the representation in the spatiotemporal reasoning of the spatiotemporal intelligent large-scale model.
[0051] Meanwhile, when creating a multimodal spatiotemporal training dataset, data diversity and coverage are primarily required to ensure the robustness and generalization ability of spatiotemporal intelligence large-scale models in various situations. Text data includes various types of content, including but not limited to natural language descriptions, commands, and dialogues, covering a wide range of semantic ranges and levels of complexity. Image data includes but is not limited to still and moving images from various perspectives, environments, and time periods. Geospatial data includes but is not limited to spatial data from various industries and formats with various precisions. These geospatial data also establish spatial relationships related to the spatiotemporal data.
[0052] Step 104: The trained spatiotemporal intelligence large-scale model is issued as a spatiotemporal intelligence large-scale model inference service to provide spatiotemporal intelligence capabilities, including spatiotemporal relationship representation, semantic understanding, environmental perception, spatiotemporal reasoning, and spatial decision-making.
[0053] After training or fine-tuning to form a spatiotemporal intelligence large-scale model, i.e., after training the spatiotemporal intelligence large-scale model, the model can be issued as a spatiotemporal intelligence large-scale model inference service, thereby providing spatiotemporal intelligence capabilities, including spatiotemporal relationship representation, semantic understanding, environmental perception, spatiotemporal reasoning, and spatial decision-making.
[0054] The spatiotemporal intelligence capability of the spatiotemporal intelligence large-scale model inference service uses the inter-object topological relationships recorded by the spatiotemporal objects, including the knowledge graph and business flow logic related to the objects, to output in the inference process of the spatiotemporal intelligence large-scale model based on the spatial relationships of the spatiotemporal object representations, thereby reducing the computational complexity of inference by the spatiotemporal intelligence large-scale model and efficiently performing inference tasks.
[0055] So-called spatiotemporal relational representations include topological spatial relations, ordinal spatial relations, metric relations, and time marks or time stamps; topological spatial relations include association, adjacency, containment, etc.; ordinal spatial relations include orientation, front-back, left-right, etc.; metric relations include length, angle, etc.
[0056] The so-called semantic understanding ability is to analyze and recognize spatiotemporal objects through vision information processing or digitized input, and output spatiotemporal objects with 0-dimensional, 1-dimensional, 2-dimensional, and 3-dimensional object characteristics, as well as the spatial relationships between objects and between objects and the environment.
[0057] The so-called environmental perception ability is to understand and predict the relative relationships and change trends between spatiotemporal objects based on spatiotemporal object representation, and provide support for further spatiotemporal relationship reasoning applications.
[0058] The so-called spatiotemporal reasoning ability is to understand the space in which objects exist in the real world, the temporal environment or the context scene of multimedia information, and the virtual environment based on the aforementioned environmental perception ability, and to reason and analyze the relationships between objects to generate the spatial object features and relationship trends of future temporal features.
[0059] The so-called spatial decision-making ability is based on the aforementioned semantic understanding, environmental perception, and spatiotemporal reasoning, and combines specific spatiotemporal scenes and workflows to realize the application of spatial decision-making in spatiotemporal scenes through spatiotemporal reasoning analysis or workflow learning.
[0060] The general-purpose artificial spatiotemporal intelligence large-scale model according to the present invention has a wide range of specific applications and can be applied in all fields that require the above-mentioned spatiotemporal intelligence capabilities.
[0061] For example, in multimedia information processing, multimedia data is converted into a spatial representation of a spatiotemporal intelligent large-scale model to generate semantic understanding and inference results, resulting in context-consistent and highly relevant answers. In other words, in multimedia information processing applications, problems are understood and semantic analysis is performed based on spatiotemporal data, multimedia input is converted into a spatiotemporal object representation with spatiotemporal relationships, and the spatiotemporal intelligent large-scale model analyzes the time-spatial features of the problem and generates answers that match the spatiotemporal context. Multimodal data is combined to provide inference results with spatiotemporal accuracy in the question-answering process, thereby ensuring the temporal and spatial consistency of the answers. For time-dependent question-answering tasks, the spatiotemporal intelligent large-scale model can estimate possible future situations based on historical spatiotemporal data and provide predicted answers.
[0062] For autonomous driving scenarios, the system uses a large-scale spatiotemporal intelligence model to process and analyze spatiotemporal data for each object in the environment in real time. The large-scale spatiotemporal intelligence model predicts the movement trajectories of surrounding objects and combines the current spatiotemporal information to generate route planning and driving decisions. Multimodal data is used to comprehensively perceive the environment around the vehicle, analyzing driving scenes with spatiotemporal correlations to support autonomous driving decision-making in complex road conditions and dynamic environments. During the autonomous driving process, it can also predict potential hazards and adjust driving strategies in real time, improving driving safety and response capabilities. When applied to autonomous driving systems, it analyzes and processes spatiotemporal data in the environment, recognizing and predicting the movement trajectories and state changes of surrounding objects to support driving decision-making.
[0063] In intelligent robot scenarios, a spatiotemporal intelligence large-scale model can process multimodal spatiotemporal data from the robot's environment in real time, improving the robot's ability to perceive environmental changes. The spatiotemporal intelligence large-scale model performs path planning and autonomous decision-making, enabling the robot to accurately determine the spatiotemporal relationships of objects in complex environments and perform rational actions and reactions. Using the spatiotemporal intelligence large-scale model to predict behavior allows the robot to understand the spatiotemporal change trends of surrounding objects and support task execution and dynamic adjustments. When performing a task, the robot makes real-time adjustments through the spatiotemporal intelligence large-scale model, ensuring that its decision-making process adapts to the temporal and spatial characteristics of objects in the environment, improving its autonomy and task execution accuracy. Processing and analyzing multimodal environmental data in real time can improve the robot's perception and autonomous decision-making capabilities in complex spatiotemporal environments.
[0064] In industrial engineering applications, the spatiotemporal intelligence large-scale model can be used to intelligently process multidimensional engineering graphics and documents. In short, the general-purpose artificial spatiotemporal intelligence model of the present invention can be successfully applied to any field and scenario that requires the use of the above-mentioned spatiotemporal intelligence functions.
[0065] As described above, the method for constructing a general-purpose artificial spatio-temporal intelligent large-scale model according to the present invention uses object representations and spatial relationship models in the form of 0-dimensional objects, 1-dimensional objects, 2-dimensional objects, and 3-dimensional objects to establish a unified, structured topological representation and storage of various multimedia information in the real world of human society, such as text, graphics, images, audio, and video. A multimodal large-scale model training method using 3D characterization is used to construct an artificial spatio-temporal intelligent large-scale model that can collectively represent and process all objects in human society. By combining spatio-temporal information processing, spatio-temporal relationship analysis, knowledge graphs, and business flow logical inference, the model provides capabilities such as spatio-temporal relationship representation, semantic understanding, environmental perception, spatio-temporal reasoning, and spatial decision-making. This solves the challenges faced by traditional large-scale language models, which have a single representation, a lack of topological relationships, and difficulty in dealing with time-space relationship problems in the 3D real world. The model is widely applicable to fields such as multimedia information processing, theoretical model derivation, autonomous driving, intelligent robots, and industrial engineering, and is highly practical.
[0066] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they have acquired the basic creative concept. Therefore, it is intended that the appended claims be interpreted as including all changes and modifications that fall within the scope of the preferred embodiments and embodiments of the present invention.
[0067] It should be noted that, in this specification, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another and do not necessarily require or imply the existence of any actual relationship or order between those entities or operations. Furthermore, the terms "comprise," "include," or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or device that includes a set of elements includes not only those elements but also other elements not expressly listed, or elements inherent in such a process, method, article, or device. Unless further limited, an element defined by the phrase "comprises" does not exclude the presence of other identical elements in the process, method, article, or device that includes that element.
[0068] Although the embodiments of the present invention have been described above with reference to the drawings, the present invention is not limited to the above-mentioned specific embodiments, which are schematic and not limiting. Those skilled in the art can create many more forms under the teachings of the present invention without departing from the spirit of the present invention and the scope of the claims, and all of these fall within the scope of protection of the present invention.
Claims
1. A computer-implemented method for constructing a general-purpose artificial spatiotemporal intelligence large-scale model, comprising: A step of using a multidimensional object abstraction representation method to process various objects and spatial relationships in the real world, and establishing a structured spatiotemporal object representation and data storage that can be collectively processed by a spatiotemporal intelligence large-scale model, where the multidimensional object abstraction refers to a 0-dimensional object, a 1-dimensional object, a 2-dimensional object, and a 3-dimensional object; processing the spatio-temporal object with a multimodal space-time coding algorithm to convert the spatio-temporal data corresponding to the spatio-temporal object into multimodal space-time coded data for subsequent training and application of a spatio-temporal intelligence large-scale model; creating a multimodal spatio-temporal training dataset based on the multimodal spatio-temporal encoding data, and using a spatio-temporal multimodal training method with three-dimensional characterization to train or fine-tune the spatio-temporal intelligent large-scale model; publishing the trained spatio-temporal intelligence large-scale model as a spatio-temporal intelligence large-scale model inference service to provide spatio-temporal intelligence capabilities, wherein the spatio-temporal intelligence capabilities include spatio-temporal relationship representation, semantic understanding, environment perception, spatio-temporal reasoning, and spatial decision-making; The various real-world objects include actual geospatial objects in the real world, virtual geospatial objects, and multimedia type objects of relative space specific to a multimedia context, the multimedia type objects including types such as text, graphics, images, audio, and video; The zero-dimensional object is an object modeling representation of a basic unit of the real world, and the basic unit object refers to various physical objects having spatial positions in the real world and multimedia information objects, and the multimedia information objects include text, graphics, images, audio, and video, and the various physical objects are represented by the spatial positions, environments, and spatial relationships of the objects, and the multimedia information objects are represented by the numerical features, scenes or contexts, and relative spaces of the objects, and include the sequence positions of characters in a paragraph and the correlation degree relationships with other characters, graphics vector nodes, and the sequence positions of image frames in a video and the correlation degree relationships with other frames, The one-dimensional object is a modeling representation of the change and evolution process of a real-world object in the time dimension, and includes representing the process or trajectory of continuous change of an object in the time dimension and the space dimension as a one-dimensional object, or a data sequence of a zero-dimensional object, representing multiple zero-dimensional objects of the same type as a one-dimensional object, and recording their flow process, logical relationship, and topological adjacent relationship in time and space, and the data sequence includes a coordinate sequence, a character sequence, a character paragraph, a mathematical formula, a graphics vector arc segment, an image set, and a video frame; The two-dimensional object is a modeling representation of real-world object adjacencies and object aggregation, and one-dimensional objects of the same type are represented as two-dimensional objects, and their temporal and spatial adjacency relationships are recorded for use in more extensive data analysis and association modeling. When processing real-world object adjacencies and object aggregation, the characteristics of the basic unit itself and the complex derivation process or the combination and association characteristics of multiple data are completely recorded according to the position, type, or change process of the object in the time or space dimension; The three-dimensional object is a modeling representation of a large-scale dataset of real-world objects or the overall spatiotemporal evolution of a complex system, including displaying a big dataset of multidimensional zero-dimensional objects, one-dimensional objects, and two-dimensional objects in a complex or composite process, recording their overall relationship or transition process in space-time, and combining objects and events of multiple spatiotemporal dimensions to generate an overall chronotope structure that supports comprehensive spatiotemporal analysis and prediction of complex systems; processing the spatio-temporal data corresponding to the spatio-temporal object with a multi-modal space-time coding algorithm to convert the spatio-temporal data into multi-modal space-time coded data for subsequent training and application of a spatio-temporal intelligence large-scale model; Step T1 of constructing a spatiotemporal topology diagram structure based on the spatiotemporal objects by regarding each basic object in the spatiotemporal data as a node of a spatiotemporal topology diagram, regarding the spatiotemporal adjacency between each node as an edge of the diagram, calculating the spatiotemporal distance or correlation between objects, and defining the edge weight between nodes, wherein the correlation between objects is determined by similarity, proximity, or chronological order in the spatiotemporal data, and the spatiotemporal data refers to multimodal spatiotemporal multidimensional data; Step T2 of encoding data based on a graph neural network, constructing a pre-training task by spatial association, and expressing the spatio-temporal topological diagram structure in high dimensions; The step of training or fine-tuning the spatio-temporal intelligent large-scale model using a spatio-temporal multimodal training method with three-dimensional characterization includes: Step N1: based on the multimodal space-time coding data, by combining specific application scenarios and task requirements, collect or generate data containing space-time information, and create a multimodal space-time training dataset, the training dataset including labels and time-space attribute information of multimodal data, to guide the adaptability and accuracy of the space-time intelligence large-scale model in specific tasks, and train the space-time related ability of the space-time intelligence large-scale model; and step N2 of training a spatio-temporal intelligence large-scale model using a multidimensional spatio-temporal attention network model through a spatio-temporal multimodal training method with the constructed multimodal spatio-temporal training dataset as input, thereby capturing time sequence relationships and spatial features annotated with 0-dimensional objects, 1-dimensional objects, 2-dimensional objects, and 3-dimensional objects in the data.
2. processing various objects and spatial relationships in the real world using a multidimensional object abstraction representation scheme, The method includes a step of collecting, recognizing, and analyzing data of various objects in the real world, including a multimedia information format, obtaining spatiotemporal information features of the various objects in the multimedia information format, and combining the environment, scene, or context in which the various objects exist, relative space, and knowledge graphs and spatial relationships of the various objects to generate multimodally structured spatiotemporal object representations of 0-dimensional objects, 1-dimensional objects, 2-dimensional objects, and 3-dimensional objects, wherein the multimedia information format includes a text format, a graphics format, an image format, an audio format, and a video format, and the spatiotemporal information features include position, shape, state, and tense; 2. The construction method of claim 1, wherein the spatiotemporal object representation simultaneously has location, tense, and spatial relationship characteristics to ensure the consistency and association of objects in the temporal and spatial dimensions, and the spatiotemporal objects are represented by topological relationships, directional or order relationships, and metric relationships, wherein the topological relationships include association, adjacency, and containment, the directional or order relationships include up / down, left / right, east / west, north / south, and the metric relationships include distance, angle, or perspective.
3. The steps of encoding data based on a graph neural network, constructing a pre-training task by spatial association, and expressing the spatio-temporal topological diagram structure in a high dimension include: updating the spatiotemporal features of each node by aggregating information of the nodes and neighboring nodes in the spatiotemporal topology diagram, and capturing the space-time associations between the multimodal spatiotemporal multidimensional data; and deep learning the multi-modal spatio-temporal multi-dimensional data with a multi-layer graph neural network to generate a coded representation that reinforces spatio-temporal associations for further model training.
4. The construction method of claim 1, characterized in that the spatio-temporal multimodal training method refers to the spatio-temporal multimodal training of 0-dimensional objects, 1-dimensional objects, 2-dimensional objects, and 3-dimensional objects with 3-dimensional characterization + time dimension, and adopts a 3-dimensional characterization method in the foundation layer of the spatio-temporal intelligent large-scale model to combine and model the time sequence and spatial features of multimodal spatio-temporal multidimensional data based on 3-dimensional spatial representation and time dimension, so as to more accurately capture the complex relationships in time and space, reflect the spatio-temporal dynamic changes of the real world, and improve the expression in the spatio-temporal inference of the spatio-temporal intelligent large-scale model.
5. When creating the multimodal spatiotemporal training dataset, the main focus is on obtaining data diversity and coverage, so as to ensure the robustness and generalization ability of the spatiotemporal intelligence large-scale model in various situations; The construction method of claim 1, wherein the text data includes various types of content, including natural language descriptions, instructions, and dialogues, covering a wide range of semantic ranges and various levels of complexity; the image data includes still and moving images from various viewpoints, in various environments, and over various time periods; and the geospatial data includes spatial data from various industries and in various formats with various precision, and constructs spatial relationships related to spatiotemporal data.
6. The construction method of claim 1, characterized in that the spatiotemporal intelligence capability of the spatiotemporal intelligence large-scale model inference service uses and outputs the topological relationship between objects recorded by the spatiotemporal objects, including the knowledge graph and business flow logic related to the objects, in the inference process of the spatiotemporal intelligence large-scale model based on the spatial relationship of the spatiotemporal object representation, thereby reducing the calculation complexity of the inference by the spatiotemporal intelligence large-scale model and efficiently performing the inference task.
7. the spatiotemporal relationship representations include topological spatial relationships, ordinal spatial relationships, metric relationships, and time marks or time stamps, the topological spatial relationships include association, adjacency, and containment, the ordinal spatial relationships include orientation, and front-back, left-right, and the metric relationships include length and angle; The semantic understanding ability is to analyze and recognize spatiotemporal objects through vision information processing or digital input, and output spatiotemporal objects having 0-dimensional object, 1-dimensional object, 2-dimensional object, and 3-dimensional object characteristics, as well as spatial relationships between objects and between objects and the environment; The environmental perception ability is to understand and predict the relative relationship and change trend between spatiotemporal objects based on the spatiotemporal object representation, and provide support for further spatiotemporal relationship reasoning applications; The spatiotemporal reasoning ability is to understand the space in which objects exist in the real world, the temporal environment or the context scene of multimedia information, and the virtual environment based on the environmental perception ability, and to infer and analyze the relational relationships between objects to generate the tendency of spatial object features and relational relationships of future temporal features; The construction method of claim 6, characterized in that the spatial decision-making ability is to combine specific spatiotemporal scenes and business flows based on the semantic understanding, the environmental perception, and the spatiotemporal reasoning, and realize the application of spatial decision-making in spatiotemporal scenes through spatiotemporal reasoning analysis or business flow learning.
Citation Information
Patent Citations
Multi-modal knowledge graph construction method
CN112200317A
Scene-aware video dialogue
JP2023510430A
Multidimensional Deep Neural Networks
JP2023539954A
JPP7579022B