AI SYSTEM BASED ON PAIRS OF SPATIO-TEMPORAL INFORMATION
The AI system addresses the limitations of single-information processing by employing spatiotemporal data pairs for robust and accurate environmental perception, enabling comprehensive understanding and intelligent responses in complex environments.
Patent Information
- Application Number
- FR2024009069
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-23
- Filing Date
- 2024-08-23
- Publication Date
- 2025-11-28
AI Technical Summary
Current AI systems lack robustness and accuracy in environmental perception due to reliance on single-information processing methods, which fail to exploit spatiotemporal relationships and multimodal data, leading to inadequate performance in complex environments.
An AI system utilizing spatiotemporal information pairs, comprising data acquisition, storage, and processing terminals, which capture, store, and analyze multidimensional data from multiple sensors (visual, auditory, olfactory) to reconstruct three-dimensional models, synchronizing and fusing data for comprehensive environmental understanding.
Enhances perception accuracy and robustness by providing abundant complementary information, simplifying complex processing, and meeting high-accuracy needs in applications like humanoid robots and driverless vehicles.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: AI SYSTEM BASED ON SPATIO-TEMPORAL INFORMATION PAIRS FIELD OF INVENTION
[0001] The present invention relates to AI techniques, in particular to an AI system based on spatiotemporal information pairs. TECHNICAL BACKGROUND
[0002] Human society increasingly needs to perceive its environment and objects within it, and to process data from them, as AI technology develops rapidly. An efficiently and accurately established environmental perception model helps to further develop and apply humanoid robots, driverless vehicles, smart glasses, and various automated inspection equipment.
[0003] Currently, it is common practice to adopt a single-information processing method for establishing the perception model of the environment, but not to adopt a paired-information approach for analyzing and reconstructing three-dimensional (x,y,z) and four-dimensional (x,y,z,t) spatiotemporal relationships. A single data point can often only provide limited information, which can lead to problems such as insufficient robustness, a lack of complementary information, and processing difficulties for complex circumstances in the perception model. Taking the example of visual information processing, some methods rely only on a two-dimensional single image for object recognition and prediction, but do not construct stereoscopic image pairs for the same object in order to perform a fusion analysis on pairs of information.This not only imposes a limitation on the model to include the object entirely, but also leads to recognition errors. As the object's appearance is significantly different from different viewing angles, a single visual input cannot capture these changes, thus affecting the accuracy of perception.
[0004] Furthermore, the two-dimensional image lacking depth information leads to limiting the model to include 3D circumstances, which further reduces the accuracy of the model construction. Although some methods still exist that employ stereoscopic vision to capture images from different angles in order to obtain three-dimensional information about an object, they are used only for binocular rangefinding or holographic projection. Related methods do not fully exploit the potential of spatiotemporal information. from different angles, nor do they combine multimodal information such as sound and gas characteristics for auxiliary analysis. This single-method perception cannot meet the needs of current artificial intelligence in environmental perception models, especially in application circumstances that require high accuracy and robustness, such as humanoid robots, driverless vehicles, smart glasses, industrial inspections, and other complex environments, where the environmental perception model must be able to handle multiple types of objects and events, as well as adapt to a variety of environmental changes, such as lighting, weather, and barriers.A single-information input perception method is often difficult to apply in large series under these circumstances and cannot provide accurate, stable, and complete perception capabilities. Description of the invention
[0005] In view of the above problems and the new knowledge, the present invention proposes an AI system based on spatio-temporal information pairs.
[0006] The embodiment of the present invention presents an AI system based on a spatio-temporal information pair, comprising: a data acquisition terminal, a data storage terminal and a data processing terminal,
[0007] wherein the data acquisition terminal which covers the spatial object environment over 720° is constructed by a data acquisition device, which is used to capture and record multidimensional data in the spatial object environment over 720° in real time, forming a series of spatio-temporal information pairs, and includes a pair of visual acquisition devices, a pair of auditory acquisition devices and a pair of olfactory acquisition devices; among them, the pair of visual acquisition devices focuses in the same direction or in the same approximate direction, simulating the way human eyes observe a spatial object at the same time, so as to form a pair of stereoscopic images;
[0008] in which the data storage terminal is used to store pairs of spatio-temporal information collected at the same time in the form of pairs and according to time series, and simultaneously establish a temporal index, so that the pairs of spatio-temporal information collected at the same time form data forms which verify each other, add up and satisfy the processing of three-dimensional data;
[0009] wherein the data processing terminal is used to process and analyze all spatio-temporal information pairs within the data storage terminal, and to synchronously fuse spatio-temporal information pairs from different types of data acquisition equipment, in order to form a three-dimensional processing result or a video stream with depth of field, and to identify and understand a complex mode-and-event in the environment of spatial objects over 720°, which includes, but is not limited to, 'identification and tracking of spatial objects and their motion states in three-dimensional space, 'understanding and prediction of events, and 'complete analysis and simulation of environments.
[0010] Preferably, when the pair of visual acquisition devices is arranged, it is necessary to ensure that the fields of vision collected by the two have a sufficient field of vision overlap area; the pair of visual acquisition devices is used to record in real time a position, a shape and a state of motion of the same spatial object in the 720° spatial object environment or the 720° spatial object environment, in order to form a pair of visual spatio-temporal information.
[0011] The pair of auditory acquisition devices is used to record in real time a sound characteristic of the same spatial object in the spatial object environment over 720° or the spatial object environment over 720°, in order to form a pair of auditory spatio-temporal information.
[0012] The pair of olfactory acquisition devices is used to record in real time an odor characteristic of the same spatial object in the spatial object environment over 720° or the spatial object environment over 720°, in order to form a pair of olfactory spatio-temporal information;
[0013] in which the pair of visual spatio-temporal information, the pair of auditory spatio-temporal information and the pair of olfactory spatio-temporal information collected at the same time form forms of data which verify each other, add up and satisfy the processing of three-dimensional data, and are used to serve the data processing terminal for a three-dimensional data reconstruction processing.
[0014] The pair of visual acquisition devices includes, but is not limited to, a video camera and a lidar.
[0015] The pair of auditory acquisition devices includes, but is not limited to, a microphone.
[0016] The pair of olfactory acquisition devices includes, but is not limited to, a sensor of gas.
[0017] Preferably, the pair of visual acquisition devices focuses in the same direction or in the same approximate direction, in order to form a pair of stereoscopic images, and adjust an acquisition frequency according to different needs;
[0018] The pair of auditory acquisition devices captures sound waves from different directions by means of microphones arranged in different positions, in order to achieve complete localization of sound sources and extraction of sound characteristics;
[0019] The pair of olfactory acquisition devices monitors a distribution and change in gas concentration in the environment by means of gas sensors arranged in different positions, in order to identify and trace the source and diffusion path of specific gas.
[0020] Preferably, the spatio-temporal information pair includes a spatial relation.
[0021] The spatial relation refers to a topological spatial relation, a sequential spatial relation, and a metric spatial relation between spatial objects, in which the topological spatial relation refers to relations of association, contiguity, inclusion, intersection, superposition, and separation between spatial objects.
[0022] The sequential spatial relationship refers to an order of arrangement of spatial objects or events in space, including front and back, left and right, top and top, and orientations between east, west, south and north.
[0023] The metric spatial relationship refers to a distance or a far-close relationship between spatial objects.
[0024] Preferably, the spatio-temporal information pair further includes a clock attribute.
[0025] The clock attribute refers to a pair of spatiotemporal information collected at the same time, to which a timestamp is assigned at the same time. The means of assigning the timestamp includes, but is not limited to, committing the timestamp to each pair of spatiotemporal information, and the timestamp includes, but is not limited to, a year, a month, a day, an hour, a minute, a second, and a millisecond. The timestamp is used to record an exact moment of acquisition of multidimensional data and to provide a precise reference in a variety of temporal dimensions for the subsequent processing and analysis of the data.
[0026] Preferably, the spatio-temporal information pair further includes a 'label attribute.
[0027] The label attribute includes, but is not limited to, identification information for equipment acquiring multidimensional data, names or categories of spatial objects, behavior patterns, circumstance states, sound characteristics, and odor type information.
[0028] The label attribute provides deep-level semantic information to the spatio-temporal information pair, so that the data processing terminal can understand and analyze circumstantial information in detail.
[0029] Preferably, the data storage terminal is used to store pairs of spatiotemporal information collected simultaneously in pairs. The means of storing the spatiotemporal information pairs includes, but is not limited to, storage in two adjacent stacks based on a time series. Each timestamp contains a pair of data elements, and the specific storage architecture includes, but is not limited to, a distributed storage architecture in a plurality of nodes in which all the spatiotemporal information pairs are stored, and each node independently processes the spatiotemporal information pairs so as to achieve parallel processing and load balancing of the spatiotemporal information pairs.
[0030] Preferably, the data storage method for the spatio-temporal information pair includes, but is not limited to, the adoption of a data organization method based on a composite key-value pair, in which the composite key-value includes a geographic object identifier, an acquisition time identifier, and a plurality of characteristic labels for the spatio-temporal information pair. The characteristic label is generated by a label attribute contained in the spatio-temporal information pair; the composite key-value represents the corresponding spatio-temporal information pair and includes, but is not limited to, a multimodal data representation, contextual analysis, correlation analysis, and support for efficient multidimensional data querying and retrieval.
[0031] Preferably, the data processing terminal processes and analyzes all spatio-temporal information pairs in the data storage terminal by means of AI, and the processing and analysis method includes, but is not limited to, deep learning and computer vision.
[0032] The method by which the data processing terminal processes and analyzes all spatiotemporal information pairs in the data storage terminal using AI comprises the steps of:
[0033] a data preprocessing, which includes steps of: purifying pairs of collected multidimensional spatio-temporal information, in order to remove noise and irrelevant information, standardizing pairs of visual, auditory and olfactory spatio-temporal information, and rectifying pairs of stereoscopic images;
[0034] a spatio-temporal synchronization, which includes steps to: ensure that the data captured by the different types of data acquisition equipment are synchronized in time, and align the data from different types of data acquisition devices to ensure their spatial consistency;
[0035] a multimodal data fusion, which includes steps of: performing an analysis and reconstruction of three-dimensional (x,y,z) and four-dimensional (x,y,z,t) spatio-temporal relationships by pairs of information, extracting the features of the visual, auditory and olfactory spatio-temporal information pairs by means of a deep learning model, and combining feature information from the different modalities by means of a fusion algorithm to form a richer representation;
[0036] a three-dimensional reconstruction, which includes, but is not limited to, the steps of: extracting depth information from stereoscopic image pairs by means of a stereoscopic matching algorithm, and combining the depth information and visual data to reconstruct an object and a circumstance in a three-dimensional space in order to form a three-dimensional model or a video stream with a depth of field;
[0037] an object recognition and tracking, which includes a step of: tracking the movement state of the identified object by means of an object detection algorithm and a tracking algorithm.
[0038] an understanding and prediction of events, which includes a step of: understanding the occurrence and development of events by analyzing patterns of object behavior and changes in the environment, or predicting future events by means of a sequence prediction model;
[0039] an 'environment analysis and simulation, which includes steps of: fully analyzing multidimensional spatio-temporal information pairs to comprehensively analyze the environment, and simulating the environment and providing an interactive experience by means of simulation technology; and
[0040] decision support, which includes a step of providing decision support to the AI system based on the results of the processing and analysis.
[0041] Preferably, the data processing terminal is further used to trace historical data of spatial objects, match and analyze new and old data, and perform self-learning and optimization, in order to continuously learn from new data, and renew and optimize the algorithm and model of the data processing terminal;
[0042] The data processing terminal is further used to process and analyze all spatio-temporal information pairs in the data storage terminal, actively discover anomalies and errors in the spatio-temporal information pairs, and repair or report them, in order to ensure the quality and reliability of the data.
[0043] In the AI system based on a pair of spatio-temporal information presented by the present invention, the data acquisition terminal which covers the spatial object environment over 720° is constructed by a data acquisition device, which is used to capture and record multidimensional data in the spatial object environment over 720° in real time, forming a series of spatio-temporal information pairs.
[0044] The data storage terminal stores pairs of spatio-temporal information collected at the same time in pairs and according to time series, and simultaneously establishes a time index, so that pairs of spatio-temporal information collected at the same time form data forms which verify each other, add up and satisfy the data processing.
[0045] The data processing terminal processes and analyzes all spatio-temporal information pairs within the data storage terminal, and synchronously processes spatio-temporal information pairs from different types of data acquisition equipment in fusion, in order to form a three-dimensional processing result or a video stream with depth of field, and to identify and understand a complex mode-and-event in the environment of spatial objects over 720°, which includes, but is not limited to, identification and tracking of spatial objects and their motion states in three-dimensional space, understanding and prediction of events, and complete analysis and simulation of environments.
[0046] In view of the existing or traditional environmental perception system, which is based on the collection and processing of data without information pairs, resulting in the problem that it cannot provide an accurate, stable, and complete perceptual capacity, the present invention creatively proposes information pairs for the collection, storage, and processing of information in the human brain. Through extensive research, it can be seen that the human brain stores and processes information pairs that intrinsically possess the attribute or sign simultaneously, such as two eyes, two ears, two nostrils, upper and lower lips, and upper and lower teeth, rather than storing data only on a single eye, one ear, one nostril, one upper or lower lip, and upper or lower teeth.This is the key information base for humankind to generate intelligence and even wisdom.
[0047] The present invention designs the equipment and system based on new knowledge of information storage and processing in the human brain, so as to bring about a major change in AI technology. That is to say, the information collected, stored, and processed by the human brain comes in pairs, such as vision from a pair of eyes and hearing from a pair of ears. and the sense of smell from a pair of nostrils. These pairs of information include not only visual, auditory, and olfactory data collected at the same time, but also attribute characteristics or symbols such as time-specific data to these data.
[0048] The AI system of the present invention performs an analysis and reconstruction of three-dimensional (x,y,z) and four-dimensional (x,y,z,t) spatio-temporal relationships using pairs of information, providing abundant information, improving the robustness of perception models, increasing complementary information, simplifying the difficulty of processing complex circumstances, improving perception accuracy, and meeting the needs of environmental perception models for current AI technology. Particularly in application circumstances requiring a high degree of accuracy and robustness, it can be widely applied to provide precise, stable, and comprehensive perception capabilities. It can not only achieve effective understanding and intelligent response to the environment but also offer the potential for a major breakthrough in AI technology applications.It provides AI software and hardware system support for equipment such as humanoid robots, driverless cars, smart glasses, and inspection equipment, with excellent usability. BRIEF DESCRIPTION OF THE FIGURES
[0049] The present invention will be described in more detail below with the aid of examples of embodiment shown schematically in the figures, so that those skilled in the art may clearly understand a variety of other advantages and benefits. The drawings are used only to show the preferred embodiment and are not considered a limitation of the present invention. In the figures, the same elements, features, and components, which are functionally identical and have the same effect, are each designated by the same reference numerals, unless otherwise indicated. These illustrate:
[0050] [Fig.1] is a schematic illustration of the arrangement of the acquisition equipment of an AI device or system based on information pairs according to one embodiment of the present invention.
[0051] [Fig.2] is a flowchart of a pair-based AI device or system information according to one embodiment of the present invention.
[0052] [Fig.3] is a flowchart of a method for collecting information pairs visuals of objects for an AI device or system based on information pairs according to one embodiment of the present invention. DETAILED DESCRIPTION OF IMPLEMENTATION METHODS
[0053] In order to make the aforementioned objective, features, and advantages of the present invention more evident and easier to understand, we describe the present invention in more detail below in conjunction with the drawings and embodiments. It should be understood that the embodiments described herein are used only to explain the present invention and are only a part of the embodiments of the present invention, nor do they include all of the embodiments or limit the present invention.
[0054] The present invention proposes an AI system based on a pair of spatio-temporal information, based on new knowledge about the storage and processing of human brain information, comprising: a data acquisition terminal, a data storage terminal and a data processing terminal.
[0055] With regard to the data acquisition terminal, data acquisition equipment can be used to construct a data acquisition terminal that covers the environment of spatial objects over 720°, such as the way the head or human body turns. The 720° view refers to a 360° view in the horizontal direction plus a 360° view in the vertical direction; this is equivalent to covering the entire space focused on the data acquisition equipment.
[0056] The data acquisition equipment is used to capture and record multidimensional data in the environment of spatial objects from any view over 720° in real time, forming a series of spatio-temporal information pairs, and includes a pair of visual acquisition devices, a pair of auditory acquisition devices and a pair of olfactory acquisition devices; among them, the pair of visual acquisition devices focuses in the same direction or in the same approximate direction, simulating the way human eyes observe a spatial object at the same time, so as to form a pair of stereoscopic images.Of course, other pairs of acquisition devices can also be added depending on the actual needs to collect data from other dimensions, such as a magnetoelectric field acquisition device, a light perception acquisition device and a lidar.
[0057] In order to better understand the data acquisition equipment mentioned above, a schematic illustration of the arrangement of the preferred data acquisition equipment according to one embodiment of the present invention is shown with reference to [Fig. 1]. The acquisition device 201, such as the lidar, is arranged at the top to detect a distance; the visual information pair acquisition device 202, i.e., the pair of visual acquisition devices, must focus in the same direction or in the same approximate direction. simulating how human eyes observe a spatial object simultaneously, in order to form a pair of stereoscopic images.
[0058] The auditory information pair acquisition device 203, i.e., the pair of auditory acquisition devices, is positioned at both ends and protrudes outwards; the olfactory information pair acquisition device 204, i.e., the pair of olfactory acquisition devices, is positioned at both ends. It is understood that the schematic illustration in [Fig. 1] is only an exemplary structure, and that the arrangements of the acquisition device may vary. Therefore, which of the acquisition devices is selected according to the actual needs, it is not necessary to list them one by one.
[0059] Preferably, when the pair of visual acquisition devices is arranged, it is necessary to ensure that the fields of view collected by the two have a sufficient field-of-view overlap area; thus, the pair of visual acquisition devices can accurately record in real time the position, shape, and motion state of the same spatial object in the environment of any 720° spatial object field of view, or the environment of any 720° spatial object field of view, in order to form a pair of visual spatiotemporal information and adjust the acquisition frequency according to different needs. The pair of visual acquisition devices includes, but is not limited to, a video camera and a lidar.
[0060] The pair of auditory acquisition devices can record in real time a sound characteristic of the same spatial object in the environment of any spatial object within a 720° field of view, or in the environment of any spatial object within a 720° field of view, in order to form a pair of auditory spatiotemporal information; the pair of auditory acquisition devices includes, but is not limited to, a microphone. The pair of auditory acquisition devices can capture sound waves from different directions by means of microphones arranged in different positions, in order to achieve complete localization of sound sources and extraction of sound characteristics.
[0061] The pair of olfactory acquisition devices can record in real time an odor characteristic of the same spatial object in the environment of any spatial object within a 720° field of view, in order to form a pair of spatiotemporal olfactory information; the pair of olfactory acquisition devices includes, but is not limited to, a gas sensor. The pair of olfactory acquisition devices can monitor the distribution and change in concentration of gas in the environment by means of gas sensors arranged at different positions, in order to identify and trace the source and diffusion path of a specific gas.
[0062] Among them, the pair of visual spatio-temporal information, the pair of auditory spatio-temporal information and the pair of spatio- Temporal olfactory data collected simultaneously form data forms that verify, add up and satisfy the three-dimensional data processing, and are used to serve the data processing terminal for three-dimensional data reconstruction processing.
[0063] The spatio-temporal information pair must include a spatial relationship, a clock attribute, and a label attribute to facilitate subsequent data management, storage, processing, and analysis. The spatial relationship to the spatio-temporal information pair refers to a topological spatial relationship, a sequential spatial relationship, and a metric spatial relationship between the spatial objects, where the topological spatial relationship refers to relationships of association, contiguity, inclusion, intersection, overlap, and separation between the spatial objects.
[0064] The sequential spatial relation refers to an order of arrangement of spatial objects or events in space, including front and back, left and right, top and top, and orientations between east, west, south and north; the metric spatial relation refers to a distance or a far-and-close relationship between spatial objects.
[0065] The clock attribute to the spatiotemporal information pair refers to a pair of spatiotemporal information collected at the same time, to which a timestamp is assigned at the same time. The means of assigning the timestamp includes, but is not limited to, assigning the timestamp to each pair of spatiotemporal information, and the timestamp includes, but is not limited to, a year, a month, a day, an hour, a minute, a second, and a millisecond. The timestamp is used to record an exact moment of acquisition of multidimensional data and to provide a precise reference in a variety of temporal dimensions for the subsequent processing and analysis of the data.
[0066] The label attribute to the spatio-temporal information pair includes, but is not limited to, identification information for equipment acquiring multidimensional data, names or categories of spatial objects, behavior patterns (e.g., patterns of human behavior in a spatial circumstance and behavior patterns of various equipment such as robots), circumstance states, sound characteristics, and odor type information (e.g., hazardous gases such as sulfur gas and sulfur dioxide belong to one odor type, and oxygen belongs to another); the label attribute provides deep-level semantic information to the spatio-temporal information pair, so that the data processing terminal can understand and analyze circumstance information in detail.
[0067] The data storage terminal is used to store pairs of spatio-temporal information collected at the same time in pairs and according to time series, and simultaneously establish a time index, so that the pairs of spatio-temporal information collected at the same time form data forms which verify each other, add up and satisfy the processing of three-dimensional data.
[0068] Preferably, the data storage terminal is used to store pairs of spatio-temporal information collected at the same time as pairs in two adjacent stacks based on a time series. Each timestamp contains a pair of data elements, and the specific storage architecture includes, but is not limited to, a distributed storage architecture, in a plurality of nodes of which all the pairs of spatio-temporal information are stored, and each node independently processes the pairs of spatio-temporal information, so as to achieve parallel processing and load balancing of the pairs of spatio-temporal information.
[0069] In terms of specific storage, it is preferable to apply the data organization method based on the composite key-value pair to the storage of data of spatio-temporal information pairs, in which the composite key-value includes a sign of geographic object, a sign of acquisition time and a plurality of characteristic labels of the spatio-temporal information pair.The feature label is generated by a label attribute contained in the spatio-temporal information pair; the composite key-value represents the corresponding spatio-temporal information pair, and includes, but is not limited to, a multimodal data representation (i.e., a record of combinations of spatio-temporal information pairs such as visual, auditory, and olfactory), a contextual analysis, and a correlation analysis (a record of a correlation between the data), thus supporting efficient multidimensional data querying and retrieval.
[0070] The data processing terminal is used to process and analyze all spatio-temporal information pairs within the data storage terminal, and to synchronously fuse spatio-temporal information pairs from different types of data acquisition equipment, in order to form a three-dimensional processing result or a video stream with depth of field, and to identify and understand a complex mode-and-event in the environment of spatial objects from any view over 720°, which includes, but is not limited to, identification and tracking of spatial objects and their motion states in three-dimensional space, understanding and prediction of events, and complete analysis and simulation of environments.
[0071] Preferably, the data processing terminal processes and analyzes all spatiotemporal information pairs in the data storage terminal using AI, and the processing and analysis method includes, but is not limited to, deep learning and computer vision. The method by which the data processing terminal processes and analyzes all spatiotemporal information pairs in the data storage terminal using AI includes the following steps:
[0072] a data preprocessing, which includes steps of: purifying collected multidimensional spatio-temporal information pairs, in order to remove noise and irrelevant information, standardizing visual, auditory and olfactory spatio-temporal information pairs, and rectifying stereoscopic image pairs;
[0073] a spatio-temporal synchronization, which includes steps of: ensuring that the data captured by the different types of data acquisition equipment are synchronized in time, and aligning the data from the different types of data acquisition devices to ensure their spatial consistency;
[0074] a multimodal data fusion, which includes steps of: performing an analysis and reconstruction of three-dimensional (x,y,z) and four-dimensional (x,y,z,t) spatio-temporal relationships by information pairs, extracting features from visual, auditory and olfactory spatio-temporal information pairs by means of a deep learning model (e.g. convolutional neural network or recurrent neural network), and combining feature information from the different modalities by means of a fusion algorithm (e.g. weighted average, decision-level fusion or feature-level fusion) to form a richer representation;
[0075] a three-dimensional reconstruction, which includes, but is not limited to, steps of: extracting depth information from stereoscopic image pairs by means of a stereoscopic matching algorithm (e.g., block matching or a stereo matching network based on deep learning), and combining the depth information and visual data to reconstruct an object and circumstance in a three-dimensional space in order to form a three-dimensional model or video stream with a depth of field;
[0076] object recognition and tracking, which includes a step of: tracking the movement state of the identified object by means of an object detection algorithm (for example, YOLO or SSD) and a tracking algorithm (for example, a Kelman Filter or a deep learning plotter);
[0077] an understanding and prediction of events, which includes a step of: understanding the occurrence and development of events by analyzing patterns of behavior of objects and changes in environment, or predict future events using a sequence prediction model (e.g., a long- and short-term memory network or a Transformer model);
[0078] an 'environment analysis and simulation, which includes steps of: fully analyzing multidimensional spatio-temporal information pairs to comprehensively analyze the environment (this includes analysis of factors such as light, sound, and smell), and simulating the environment and providing an interactive experience by means of simulation technology (e.g., virtual reality or augmented reality); and
[0079] decision support, which includes a step of providing decision support (e.g., route planning, anomaly detection or resource allocation) to the AI system based on the results of processing and analysis.
[0080] Furthermore, the data processing terminal is also used to trace historical data of spatial objects, match and analyze new and old data, and perform self-learning and optimization in order to continuously learn from new data and renew and optimize the algorithm and model of the data processing terminal. The data processing terminal is also used to process and analyze all spatiotemporal information pairs in the data storage terminal, actively discover anomalies and errors in the spatiotemporal information pairs, and repair or report them in order to ensure data quality and reliability. This achieves a self-learning and error-correction capability for the entire AI system and further improves the intelligence of the AI system.
[0081] The entire procedure of the aforementioned AI system based on spatio-temporal information pairs, which includes the steps from construction to data processing and analysis, can be summarized as follows in [Fig.2].
[0082] SI: Construct data acquisition equipment in all directions.
[0083] That is to say, the data acquisition terminal which can cover the spatial object environment from any view over 720° is constructed by means of visual, auditory and olfactory data acquisition equipment, in which the arrangement mode of each data acquisition device can be adjusted according to requirements and is not limited to the layout mode of [Fig.1].
[0084] S2: Collect continuous spatio-temporal information pairs multidimensional in a circumstance.
[0085] After the data acquisition equipment in all directions has been built, the multidimensional data is captured and recorded in the circumstance over 720° in real time by means of the data acquisition equipment, in order to form a series of spatio-temporal information pairs, in In these devices, the pair of visual acquisition units focuses in the same or approximately the same direction, simulating how human eyes observe a spatial object simultaneously, thus forming a pair of stereoscopic images. For example, a flowchart of a method for acquiring visual information from the pair of acquisition units is shown with reference to [Fig. 3]. The visual information pair acquisition device (i.e., the pair of visual acquisition units) performs an acquisition on the observation object A and obtains a pair of stereo images by acquiring in the visual information pair overlap area A1 and the visual information pair overlap area A2 using the visual information. Its spatial coordinate system is illustrated in [Fig. 3].
[0086] S3: Give the spatial relation, the clock attribute and the label attribute to the spatio-temporal information pair.
[0087] It is necessary to assign a spatial relationship, a clock attribute, and a label attribute to the resulting spatiotemporal information pair. The spatial relationship includes a topological spatial relationship, a sequential spatial relationship, and a metric spatial relationship. The clock attribute refers to a pair of visual, auditory, and olfactory data collected at the same time, to which a timestamp is assigned at the same time. The label attribute includes not only the identifying information of the equipment to which a data source belongs, but also, but not limited to, information such as names and categories of geographic objects, human behavior patterns, and equipment behavior patterns.
[0088] S4: Establish a mechanism for storing spatio-temporal information.
[0089] After the first three steps have been completed, the data storage terminal performs data storage. In terms of data storage, the spatiotemporal information collected simultaneously is organized and stored in pairs, so that the pairs of information can form data forms that verify, add up, and satisfy the three-dimensional data processing requirements.
[0090] S5: Perform fusion analysis and processing on the spatiotemporal information pair to form a three-dimensional processing result or a video stream with a depth of field and a sign of attributes. After the spatiotemporal information pair has been processed and stored, the data processing terminal finally completes the data analysis and processing. In terms of data processing, data processing and analysis methods such as AI, like deep learning and computer vision, based on multidimensional spatiotemporal information pairs, are used to process synchronously fusing spatio-temporal information pairs from different types of data acquisition equipment, in order to form a three-dimensional processing result or video stream with depth of field, and identifying and understanding a complex mode-and-event in a circumstance, which includes, but is not limited to, identification and tracking of spatial objects and their motion states in three-dimensional space, understanding and predicting events, and comprehensive analysis and simulation of environments.
[0091] The 5 steps summarized above are clearly understood in combination with the preceding, it is not necessary to detail them.
[0092] In summary, in the AI system based on a pair of spatio-temporal information presented by the present invention, the data acquisition terminal which covers the spatial object environment from any view over 720° is constructed by a data acquisition device, which is used to capture and record multidimensional data in the spatial object environment from any view over 720° in real time, forming a series of spatio-temporal information pairs.
[0093] The data storage terminal stores the spatio-temporal information pairs collected at the same time in the form of pairs and according to time series, and simultaneously establishes a time index, so that the spatio-temporal information pairs collected at the same time form data forms which verify each other, add up and satisfy the data processing.
[0094] The data processing terminal processes and analyzes all spatio-temporal information pairs within the data storage terminal, and synchronously fuses spatio-temporal information pairs from different types of data acquisition equipment, in order to form a three-dimensional processing result or a video stream with depth of field, and to identify and understand a complex mode-and-event in the environment of spatial objects from any view over 720°, which includes, but is not limited to, identification and tracking of spatial objects and their motion states in three-dimensional space, understanding and prediction of events, and complete analysis and simulation of environments.
[0095] In view of the existing or traditional environmental perception system, which adopts a single-information input perception method often difficult to apply in large series under these circumstances and cannot provide accurate, stable, and complete perception capabilities, the present invention creatively proposes information pairs for the collection, storage, and processing of information by the human brain; that is, the information collected, stored, and processed by the human brain comes in pairs, such as vision from a pair of eyes, hearing from a pair of ears, and smell from a pair of nostrils. These information pairs include not only visual, auditory, and olfactory data collected at the same time, but also assign attributes or symbols such as specific time to this data.
[0096] The AI system of the present invention performs an analysis and reconstruction of three-dimensional (x,y,z) and four-dimensional (x,y,z,t) spatio-temporal relationships using pairs of information, providing abundant information, improving the robustness of perception models, increasing complementary information, simplifying the difficulty of processing complex circumstances, improving perception accuracy, and meeting the needs of environmental perception models for current AI technology. Particularly in application circumstances requiring a high degree of accuracy and robustness, it can be widely applied to provide precise, stable, and comprehensive perception capabilities. It can not only achieve effective understanding and intelligent response to the environment but also offer the potential for a major breakthrough in AI technology applications.It provides AI software and hardware system support for equipment such as humanoid robots, driverless cars, smart glasses, and inspection equipment, with excellent usability.
[0097] In the preceding detailed description, various features aimed at improving the rigor of the presentation were grouped into one or more examples. It should be clarified, however, that the above description is purely illustrative and is in no way intended to be restrictive. It serves to cover all variants, modifications, and equivalents of the different features and implementation examples. Those skilled in the art, by virtue of their technical knowledge, will clearly see many other examples immediately and directly arising from the above description.
[0098] Finally, it should also be noted that the herein-included relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, but they neither require nor necessarily suggest any such actual relation or sequence between these entities or operations. Furthermore, the terms "comprising" and "possessing," or any variant thereof, are intended to cover non-exclusive inclusion, such that a process, method, object, or terminal device including a series of elements includes not only those elements but also other elements not expressly listed, or a process, method, object, or terminal device further includes elements that are inherent in it. In the absence of other restrictions, the elements defined by the phrase "including one..." do not exclude the existence of others. identical elements in the process, method, object or terminal equipment which comprises the elements.
[0099] Embodiments of the present invention have been described above in conjunction with the drawings, but the present invention is not limited to the specific embodiments described above, and the specific embodiments are merely indicative and not restrictive. Inspired by the present invention, those skilled in the art may also implement numerous forms without departing from the object of the present invention and the scope of protection of the claims, all of which fall within the protection of the present invention.
Claims
1. Demands AI system based on a pair of spatio-temporal information, characterized in that it provides AI software and hardware system support to a humanoid robot, a driverless car, smart glasses and inspection equipment, and it includes a data acquisition terminal, a data storage terminal and a data processing terminal, in which said data acquisition terminal which covers a 720° spatial object environment is constructed by data acquisition equipment, which is used to capture and record multidimensional data in said 720° spatial object environment in real time, forming a series of spatio-temporal information pairs, and includes a pair of visual acquisition devices (202), a pair of auditory acquisition devices (203) and a pair of olfactory acquisition devices (204);among them, said pair of visual acquisition devices (202) focuses in the same direction or in the same approximate direction, simulating the way human eyes observe a spatial object at the same time, so as to form a pair of stereoscopic images; in which said data storage terminal is used to store pairs of spatio-temporal information collected at the same time in pairs and according to time series, and simultaneously establish a temporal index, so that pairs of spatio-temporal information collected at the same time form data forms which verify each other, add up and satisfy the processing of three-dimensional data; wherein said data processing terminal is used to process and analyze all spatiotemporal information pairs within said data storage terminal, and to synchronously fuse spatiotemporal information pairs from different types of data acquisition equipment in order to form a three-dimensional processing result or a video stream with depth of field, and to identify and understand a complex mode-and-event in said 720° spatial object environment, which includes one or more spatial object identification and tracking operations and their states. of movement in three-dimensional space, of understanding and predicting events, and of complete analysis and simulation of environments; The method by which said data processing terminal processes and analyzes all spatiotemporal information pairs in said data storage terminal using AI comprises the following steps: a preprocessing of data, which includes steps of: purifying pairs of collected multidimensional spatio-temporal information, in order to remove noise and irrelevant information, standardizing pairs of visual, auditory and olfactory spatio-temporal information, and rectifying pairs of stereoscopic images; spatio-temporal synchronization, which includes steps of: ensuring that data captured by different types of data acquisition equipment are synchronized in time, and aligning data from different types of data acquisition devices to ensure their spatial consistency; multimodal data fusion, which includes steps of: performing an analysis and reconstruction of three-dimensional (x,y,z) and four-dimensional (x,y,z,t) spatio-temporal relationships by pairs of information, extracting features from pairs of visual, auditory, and olfactory spatio-temporal information using a deep learning model, and combining feature information from different modalities using a fusion algorithm to form a richer representation; a three-dimensional reconstruction, which includes steps of: extracting depth information from stereoscopic image pairs using a stereoscopic matching algorithm, and combining the depth information and visual data to reconstruct an object and circumstance in a three-dimensional space in order to form a three-dimensional model or video stream with depth of field; object recognition and tracking, which includes a step of: tracking the movement state of the identified object using an object detection algorithm and a tracking algorithm;
2. an understanding and prediction of events, which includes a step of: understanding the occurrence and development of events by analyzing patterns of object behavior and environmental changes, or predicting future events using a sequence prediction model; an 'environment analysis and simulation, which includes steps of: fully analyzing multidimensional spatio-temporal information pairs to comprehensively analyze the environment, and simulating the environment and providing an interactive experience using simulation technology; and decision support, which includes a step of providing decision support to the AI system based on the results of the processing and analysis. AI system according to claim 1, characterized in that when said pair of visual acquisition devices (202) is arranged, it is necessary to ensure that the fields of vision collected by the two have a sufficient field of vision overlap area; said pair of visual acquisition devices (202) is used to record in real time a position, shape and state of motion of the same spatial object in said 720° spatial object environment or said 720° spatial object environment, in order to form a pair of visual spatio-temporal information; said pair of auditory acquisition devices (203) is used to record in real time a sound characteristic of the same spatial object in said spatial object environment over 720° or said spatial object environment over 720°, in order to form a pair of auditory spatio-temporal information; said pair of olfactory acquisition devices (204) is used to record in real time an odor characteristic of the same spatial object in said spatial object environment over 720° or said spatial object environment over 720°, in order to form a pair of spatio-temporal olfactory information; in which the pair of visual spatio-temporal information, the pair of auditory spatio-temporal information, and the pair of olfactory spatio-temporal information collected simultaneously form data forms that verify each other, add up, and satisfy the three-dimensional data processing requirements, and are used to serve said data processing terminal for three-dimensional data reconstruction processing; said pair of visual acquisition devices (202) includes one or more of a video camera and a lidar; said pair of auditory acquisition devices (203) includes a microphone; said pair of olfactory acquisition devices (204) includes a gas sensor.
3. AI system according to claim 2, characterized in that the pair of visual acquisition devices (202) focuses in the same direction or in the same approximate direction, in order to form a pair of stereoscopic images, and adjust an acquisition frequency according to different needs; said pair of auditory acquisition devices (203) captures sound waves from different directions by means of microphones arranged in different positions, in order to achieve complete localization of sound sources and extraction of sound features; said pair of olfactory acquisition devices (204) monitors a distribution and change in concentration of gas in an environment by means of gas sensors arranged in different positions, in order to identify and trace a specific gas source and diffusion path.
4. AI system according to claim 1, characterized in that said spatio-temporal information pair includes a spatial relation; said spatial relation refers to a topological spatial relation, a sequential spatial relation, and a metric spatial relation between spatial objects, in which said topological spatial relation refers to relations of association, contiguity, inclusion, intersection, superposition, and separation between spatial objects; said sequential spatial relation refers to an order of arrangement of spatial objects or events in space, and includes front and back, left and right, top and top orientations, and between east, west, south, and north; said metric spatial relation refers to a distance or a far-and-close relation between spatial objects.
5. AI system according to claim 1, characterized in that said spatio-temporal information pair further includes a clock attribute; said clock attribute refers to a spatio-temporal information pair collected at the same time, to which a timestamp is given at the same time, the means of giving the timestamp includes an engagement of the timestamp in each spatio-temporal information pair, and the timestamp includes one or more of a year, month, day, hour, minute, second, and millisecond; said timestamp is used to record an exact time of multidimensional data acquisition and to provide an accurate reference in a variety of temporal dimensions for subsequent data processing and analysis.
6. AI system according to claim 1, characterized in that said spatio-temporal information pair further includes a 'label attribute'; said label attribute includes one or more of the identification information of the equipment acquiring the multidimensional data, names or categories of spatial objects, behavioral patterns, circumstance states, sound characteristics and odor type information; said label attribute provides deep-level semantic information to said spatio-temporal information pair, so that said data processing terminal can understand and analyze circumstance information in detail.
7. AI system according to claim 1, characterized in that said data storage terminal is used to store pairs of spatio-temporal information collected at the same time in the form of pairs, the means of storing the spatio-temporal information pairs includes storage in two adjacent stacks based on a time series, each timestamp contains a pair of data elements, and a specific storage architecture includes a distributed storage architecture, in a plurality of nodes of which, all the spatio-temporal information pairs are stored, and each node independently processes the spatio-temporal information pairs, so as to achieve parallel processing and load balancing of the spatio-temporal information pairs.
8. AI system according to claim 1, characterized in that the data storage method of said spatio-temporal information pair includes an adoption of a data organization method based on a composite key-value pair, wherein said composite key-value includes a geographic object sign, an acquisition time sign and a plurality of feature labels, said feature label is generated by a label attribute contained in said spatio-temporal information pair; said composite key-value represents a spatio-temporal information pair corresponding to that, and includes one or more of a multimodal data representation, a contextual analysis, a correlation analysis, and support for efficient multidimensional data querying and retrieval.
9. AI system according to claim 1, characterized in that said data processing terminal processes and analyzes all spatio-temporal information pairs in said data storage terminal by means of AI, and the processing and analysis method includes one or more of a deep learning and a computer vision.
10. AI system according to claim 1, characterized in that said data processing terminal is further used to trace historical data of spatial objects, match and analyze new and old data, and perform self-learning and optimization, in order to continuously learn new data, and renew and optimize the algorithm and model of said data processing terminal; said data processing terminal is further used to process and analyze all spatio-temporal information pairs in said data storage terminal, actively discover anomalies and errors in spatio-temporal information pairs, and repair or report them, in order to ensure data quality and reliability.