Artificial intelligence systems based on spatiotemporal information pairs

The AI system addresses the limitations of single-data processing by using spatiotemporal information pairs for 3D reconstruction, improving perception accuracy and robustness in complex environments through multi-modal data fusion and deep learning analysis.

JP7799289B2Active Publication Date: 2026-01-15PEKING UNIV +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024139712
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2024-05-23
Filing Date
2024-08-21
Publication Date
2026-01-15
Estimated Expiration
2044-08-21

AI Technical Summary

Technical Problem

Current environmental perception models rely on single-data information processing, failing to analyze and reconstruct spatiotemporal relationships, leading to limited information, lack of robustness, and difficulty in processing complex scenes, especially in applications requiring high accuracy and robustness like humanoid robots and autonomous driving.

Method used

An artificial intelligence system that collects, stores, and processes spatiotemporal information pairs using multiple data collection devices (visual, auditory, olfactory) to form stereoscopic images and audio-olfactory pairs, synchronizes and fuses this data for 3D reconstruction, and uses deep learning for analysis.

Benefits of technology

Enhances perception accuracy and robustness by providing abundant complementary information, enabling efficient understanding and intelligent response to complex environments, suitable for large-scale applications in scenarios requiring high precision and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007799289000001
    Figure 0007799289000001
  • Figure 0007799289000002
    Figure 0007799289000002
  • Figure 0007799289000003
    Figure 0007799289000003
Patent Text Reader

Abstract

To provide an artificial intelligence system based on spatial-temporal information pairs in which: by integrally deploying paired visual, auditory and olfactory collection devices, multi-dimensional continuous spatial-temporal information pairs such as positions, morphologies, motion states, sounds and odors in the ambient environment of the device or of the same spatial object in the environment are recorded in real time.SOLUTION: In an artificial intelligence system, spatial-temporal information pairs at the same time are given with not only spatial relationships and a clock attribute, but also rich tag attributes such as categories and behavior patterns, the spatial-temporal information pairs are organized and stored in a form of pairs, synchronized for being analyzed by fusion processing, thereby achieving cross-validation and forming a 3D processing result or video stream with depth of field and attribute identifiers, thereby recognizing and understanding spatial-temporal relationships and events in scenes, and providing artificial intelligence software and hardware system support for humanoid robots, unmanned vehicles, smart glasses, patrol devices, etc.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the technical field of artificial intelligence, in particular to an artificial intelligence system based on spatio-temporal information pairs. [Background technology]

[0002] With the rapid development of artificial intelligence-related technologies, human society is increasingly required to perceive and process data about the surrounding environment and target objects within it. Building an efficient and accurate environmental perception model will contribute to the further development and application of humanoid robots, autonomous driving, smart glasses, various automatic patrol devices, and more.

[0003] Current environmental perception model construction typically uses single-data information processing methods to construct the model, but does not analyze and reconstruct the three-dimensional (x, y, z) and four-dimensional (x, y, z, t) spatiotemporal relationships in the form of information pairs. In many cases, single data can only provide limited information, which can lead to problems such as a lack of robustness in the perception model, a lack of complementary information, and difficulty in processing complex scenes.

[0004] Taking visual information processing as an example, some methods rely only on a single 2D image for target recognition and prediction, without building a stereoscopic image pair for the same object and performing fusion analysis of the information pair. This not only limits the model's comprehensive understanding of the target object, but also may lead to recognition errors. Objects appear very different at different visual angles, and a single visual information cannot capture these changes, affecting the accuracy of perception.

[0005] Furthermore, the lack of depth information in 2D images limits the model's ability to understand the 3D aspects of a scene, further reducing the accuracy of model construction. While some methods use stereoscopic techniques to capture images from various angles and obtain 3D information about objects, these methods are limited to binocular ranging or holographic projection. These methods do not fully utilize the potential of spatiotemporal information from various viewing angles, nor do they combine it with multimodal information such as sound or gas signatures for auxiliary analysis. This single perception method cannot meet current artificial intelligence needs for environmental perception models. In particular, in applications requiring high accuracy and robustness, such as complex environments for humanoid robots, autonomous driving, smart glasses, and industrial patrols, environmental perception models must be able to handle multiple types of objects and events and adapt to various environmental changes, such as lighting, weather, and occlusion. Perception methods that rely on a single input often have difficulty being applied on a large scale in these scenes and, moreover, are unable to provide accurate, stable, and comprehensive perception functions. Summary of the Invention [Problem to be solved by the invention]

[0006] In view of the above problems and completely new findings, the present invention proposes an artificial intelligence system based on spatio-temporal information pairs. [Means for solving the problem]

[0007] An embodiment of the present invention comprises: Including a data collection side, a data storage side, and a data processing side; A data collection device is used to establish a data collection terminal covering a 720° spatial object environment, the data collection device is used to capture and record multi-dimensional data in the 720° spatial object environment in real time, and form a series of spatiotemporal information pairs, including a pair of visual collection devices, a pair of auditory collection devices, and a pair of olfactory collection devices, and the pair of visual collection devices focus in the same or approximately the same direction, simulating the way in which both human eyes simultaneously observe any spatial object, and form a stereoscopic image pair; the data storage side is used to store the spatiotemporal information pairs collected at the same time in a paired form in chronological order, so that the spatiotemporal information pairs collected at the same time are mutually verified and complemented, and are in a data format that satisfies three-dimensional data processing, and to establish a time-based index; The data processing side processes and analyzes all spatiotemporal information pairs in the data storage side, synchronizes and fuses the spatiotemporal information pairs from various types of data collection devices, and forms a three-dimensional processing result or video stream with depth of field, which is used to recognize and understand complex patterns and events in the 720° spatial object environment, including, but not limited to, recognizing and tracking spatial objects and their motion states in three-dimensional space, understanding and predicting events, and comprehensively analyzing and simulating the environment, to provide an artificial intelligence system.

[0008] Optionally, the paired vision collecting devices need to ensure that the visual fields collected by both of them have a sufficient visual field overlap area when positioned, and the paired vision collecting devices are used to record the position, shape, and motion state of the 720° spatial object environment or the same spatial object in the 720° spatial object environment in real time, and form a visual-spatiotemporal information pair; The paired auditory collection devices are used to record the audio characteristics of the 720° spatial object environment or the same spatial object in the 720° spatial object environment in real time, and form an audio-spatiotemporal information pair; The olfactory collection device is used to record the odor characteristics of the same spatial object in the 720° spatial object environment or the 720° full-angle spatial object environment in real time, and obtain olfactory spatiotemporal information pairs; The visual spatiotemporal information pair, the auditory spatiotemporal information pair, and the olfactory spatiotemporal information pair at the same time are mutually verified and complemented to form a data format that satisfies data processing, and are used for three-dimensional data reconstruction processing by the data processing side; The paired vision gathering devices include, but are not limited to, cameras, laser radars, The paired hearing collection device includes, but is not limited to, a microphone; The paired olfactory collecting device includes, but is not limited to, a gas sensor.

[0009] Optionally, said pair of vision collection devices focus in the same or nearly the same direction to form a stereoscopic image pair, and collection frequency is adjusted according to demand; The paired auditory collection devices use microphones distributed at various positions to capture sound waves from various directions, thereby realizing omnidirectional sound source location and voice feature extraction; The paired olfactory collecting device monitors the gas distribution and concentration changes in the environment through gas sensors placed at various locations, and recognizes and tracks the source and diffusion path of specific gases.

[0010] Optionally, said spatiotemporal information pair comprises a spatial relationship; The spatial relationships refer to topological, ordinal, and metric spatial relationships between spatial objects, and the topological spatial relationships refer to association, adjacency, containment, intersection, overlap, and separation relationships between spatial objects; The spatial order relationship refers to the spatial arrangement order of spatial objects or events, including front-back, left-right, up-down, and east-west-north-south directional relationships; The metric spatial relationship refers to the distance or perspective relationship between spatial objects.

[0011] Optionally, said spatiotemporal information pair further comprises a clock attribute; The clock attribute refers to giving the same time identifier to pairs of spatiotemporal information collected at the same time, and a specific method includes embedding a timestamp into each pair of spatiotemporal information, the timestamp including but not limited to year, month, day, hour, minute, second, and millisecond, for recording the exact time of collecting multidimensional data and providing accurate references on various time dimensions for subsequent data processing and analysis.

[0012] Optionally, said spatiotemporal information pair further comprises a tag attribute; The tag attributes include, but are not limited to, collecting identifier information of the device to which the multidimensional data belongs, spatial object name or category, behavior pattern, scene state, sound feature, and smell type information; The tag attributes provide high-level semantic information to the spatiotemporal information pairs so that the data processing side can understand and analyze the scene information in a fine-grained manner.

[0013] Optionally, the data storage side is used to store pairs of spatiotemporal information collected at the same time in a paired manner, a specific storage manner includes, but is not limited to, storing in two adjacent stacks in chronological order, each time identifier includes one paired data item, and a specific storage architecture includes, but is not limited to, a distributed storage architecture; All spatiotemporal information pairs are stored in a distributed manner on multiple nodes of the distributed storage architecture, and each node processes each spatiotemporal information pair independently, thereby achieving parallel processing and load balancing of the spatiotemporal information pairs.

[0014] Optionally, the data storage method for the spatiotemporal information pairs includes, but is not limited to, adopting a data organization method based on composite key-value pairs, where a composite key includes a spatial object identifier of the spatiotemporal information pair, a collection time identifier, and several feature tags generated by tag attributes included in the spatiotemporal information pair, and a composite value represents a corresponding spatiotemporal information pair, supporting multi-dimensional efficient data query and search, including, but not limited to, multimodal data representation, context analysis, and relevance analysis.

[0015] Optionally, the data processing side utilizes artificial intelligence to process and analyze all spatiotemporal information pairs in the data storage side, and the processing and analysis methods include, but are not limited to, deep learning and machine vision; The method in which the data processing side processes and analyzes all spatiotemporal information pairs in the data storage side using artificial intelligence is as follows: data preprocessing, including cleaning the collected multidimensional spatiotemporal information pairs, removing noise and irrelevant information, standardizing the visual, auditory, and olfactory spatiotemporal information pairs, and correcting the stereo image pairs; Spatiotemporal synchronization, which includes ensuring temporal synchronization of data captured by different types of data collection devices and aligning and spatially coherent data from the different types of data collection devices; Multimodal data fusion involves analyzing and reconstructing 3D (x,y,z) and 4D (x,y,z,t) spatiotemporal relationships in the form of information pairs, using deep learning models to perform feature extraction on visual, auditory, and olfactory spatiotemporal information pairs, and combining feature information from various modalities through a fusion algorithm to form a richer representation. 3D reconstruction, including but not limited to using a stereo matching algorithm to extract depth information from stereo image pairs and combining the depth information with visual data to reconstruct objects and scenes in 3D space to form a 3D model or video stream with depth of field; object recognition and tracking, including tracking the motion state of the recognized object using a target detection algorithm and a tracking algorithm; Event understanding and prediction, including understanding the occurrence and progression of events by analyzing object behavior patterns and environmental changes, or predicting future events using sequence prediction models; Analysis and simulation of an environment, including integrating and analyzing multidimensional space-time information pairs, conducting a comprehensive analysis of the environment, and simulating the environment using simulation technology to provide an interactive experience; and decision support, providing decision support to the artificial intelligence system based on the results of the processing and analysis.

[0016] Optionally, the data processing side is further used to perform retrospective analysis of historical data of spatial objects, matching analysis of old and new data, self-learning and optimization, so as to constantly learn from new data and update and optimize the algorithms and models of the data processing side; The data processing side is further used to actively discover abnormalities and errors in the spatiotemporal information pairs based on processing and analysis of all spatiotemporal information pairs in the data storage side, and repair or report them to ensure the quality and reliability of the data. [Effects of the Invention]

[0017] The artificial intelligence system based on spatiotemporal information pairs according to the present invention uses a data collection device to establish a data collection terminal covering a 720° spatial object environment, and the data collection device is used to capture and record multi-dimensional data in the 720° spatial object environment in real time to form a series of spatiotemporal information pairs.

[0018] The data storage side stores pairs of spatiotemporal information collected at the same time in a chronological order in pairs so that the pairs of spatiotemporal information collected at the same time are mutually verified, mutually complemented, and in a data format that satisfies data processing, and establishes a time-based index.

[0019] The data processing side processes and analyzes all spatiotemporal information pairs in the data storage side, synchronizes and fuses the spatiotemporal information pairs from various types of data collection devices, forms a three-dimensional processing result or video stream with depth of field, and recognizes and understands complex patterns and events in the 720° spatial object environment, including but not limited to recognizing and tracking spatial objects and their motion states in three-dimensional space, understanding and predicting events, and comprehensively analyzing and simulating the environment.

[0020] In response to the problem that existing or conventional environmental perception systems are built on the basis of collecting and processing non-paired information and data, and as a result are unable to provide accurate, stable, and comprehensive perception functions, the present invention innovatively proposes an information pair method for the human brain to collect, store, and process information. Through careful investigation, it has been found that the human brain stores and processes information pairs with attributes or identifiers at the same time, such as sensory data from two eyes, two ears, two nostrils, upper and lower lips, and upper and lower teeth, rather than data from a single eye, a single ear, a single nostril, upper or lower lip, or upper or lower teeth, which is an important information foundation for humans to create intelligence and even wisdom.

[0021] This invention aims to revolutionize artificial intelligence technology by designing devices and systems based on a completely new understanding of the information storage process of the human brain. The information collected, stored, and processed by the human brain is generated in pairs, such as binocular vision, binaural hearing, and binasal olfactory information. These information pairs not only contain visual, auditory, and olfactory data collected at the same time, but also assign attributes and identifiers, such as specific times, to these data.

[0022] The artificial intelligence system of the present invention analyzes and reconstructs three-dimensional (x,y,z) and four-dimensional (x,y,z,t) spatiotemporal relationships in the form of spatiotemporal information pairs, providing abundant information, improving the robustness of the perception model, increasing complementary information, reducing the difficulty of processing complex scenes, and improving perception accuracy. This meets the current needs for artificial intelligence environmental perception models, and is capable of large-scale applications, especially in application scenarios requiring high precision and high robustness, providing accurate, stable, and comprehensive perception functions. This not only enables efficient understanding and intelligent response to the environment, but also brings about great advances in the application of artificial intelligence technology, and can provide artificial intelligence software and hardware system support for humanoid robots, unmanned vehicles, smart glasses, patrol devices, etc., and is very practical. [Brief explanation of the drawings]

[0023] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of the preferred embodiments. The drawings are used only for purposes of illustrating the preferred embodiments and are not intended to limit the invention. Furthermore, like reference numerals represent like parts throughout the drawings. [Figure 1] 1 is a schematic diagram of the arrangement of a collection device of an artificial intelligence device or system based on information pairs according to an embodiment of the present invention; [Figure 2] 1 is a flowchart of an information pair-based artificial intelligence device or system according to an embodiment of the present invention. [Figure 3] 1 is a schematic diagram of a method for visual information pair collection target of an artificial intelligence device or system based on information pair in an embodiment of the present invention; DETAILED DESCRIPTION OF THE INVENTION

[0024] In order to make the above-mentioned objects, features and advantages of the present invention more clearly understandable, the present invention will be described in more detail below with reference to the drawings and specific embodiments. It should be understood that the specific examples described in this specification are only used to explain the present invention, and are only some examples of the present invention, not all examples of the present invention, and do not limit the present invention.

[0025] Based on a completely new understanding of the information storage process by the human brain, the present invention proposes an artificial intelligence system based on a spatiotemporal information pair, including a data collection side, a data storage side, and a data processing side.

[0026] On the data collection side, it is possible to use a data collection device to build a data collection side that covers a 720° spatial object environment, such as the rotation of a human head or body, where the so-called 720° full angle means 360° horizontally and 360° vertically, thus covering the entire spatial range centered on the data collection device.

[0027] The data collection device is used to capture and record multi-dimensional data in a 720° full-angle spatial object environment in real time, and form a series of spatiotemporal information pairs, including a pair of visual collection devices, a pair of auditory collection devices, and a pair of olfactory collection devices, and the pair of visual collection devices focus in the same or approximately the same direction, simulating the way that both human eyes simultaneously observe any spatial object, and form a stereoscopic image pair. Of course, according to actual needs, it is also possible to add other paired collection devices, such as magnetic field collection devices, optical induction collection devices, and laser radar, to collect data in other dimensions.

[0028] In order to better understand the above data collection device, please refer to Figure 1, which shows a structural schematic diagram of a preferred data collection device in an embodiment of the present invention, where a collection device 201 such as a laser radar is provided on top to detect distance. The visual information pair collection device 202 is a pair of visual collection devices that need to focus in the same direction or approximately the same direction, simulating the way both human eyes simultaneously observe any spatial object, and form a stereoscopic image pair.

[0029] The paired auditory information collecting devices 203 are paired auditory collecting devices installed at both ends and protruding outward. The paired olfactory information collecting devices 204 are paired olfactory collecting devices installed at both ends. Note that the structural diagram shown in FIG. 1 is only an example of the structure, and there are various arrangements of the collecting devices. The specific type of collecting device to be installed is determined according to actual needs, but is not listed one by one.

[0030] Preferably, the paired vision collecting devices should be positioned to ensure that the fields of view collected by both devices have a sufficient overlapping area, so that the paired vision collecting devices can accurately record the position, shape, and motion status of the same spatial object in the 720° full-angle spatial object environment or the 720° full-angle spatial object environment in real time, and form a visual-spatiotemporal information pair, the collection frequency of which can be adjusted according to demand. The paired vision collecting devices include, but are not limited to, video cameras, laser radars, etc.

[0031] The paired auditory collecting devices can record the audio features of a 720° full-angle spatial object environment or the same spatial object in the 720° full-angle spatial object environment in real time to form an audio spatiotemporal information pair. The paired auditory collecting devices can include, but are not limited to, microphones, etc. The paired auditory collecting devices can achieve omnidirectional sound source location and audio feature extraction by capturing sound waves from various directions using microphones distributed at various positions.

[0032] The paired olfactory collecting devices can record the odor characteristics of a 720° full-angle spatial object environment or the same spatial object within the 720° full-angle spatial object environment in real time to obtain a pair of olfactory spatiotemporal information. The paired olfactory collecting devices can include, but are not limited to, gas sensors. The paired olfactory collecting devices can monitor gas distribution and concentration changes in the environment using gas sensors arranged at various positions, and can recognize and track the source and diffusion path of a specific gas.

[0033] Here, the visual spatiotemporal information pair, auditory spatiotemporal information pair, and olfactory spatiotemporal information pair from the same time are mutually verified and complemented to form a data format that satisfies data processing, and are used by the data processing side to reconstruct 3D data.

[0034] A spatiotemporal information pair needs to have information such as spatial relationship, clock attribute, tag attribute, etc., which facilitates storage management and subsequent data processing and analysis. The spatial relationship in a spatiotemporal information pair refers to the topological spatial relationship, ordinal spatial relationship, and metric spatial relationship between spatial objects, and the topological spatial relationship refers to the association, adjacency, and containment relationships between spatial objects, including the intersection, overlap, and separation relationships between spatial objects.

[0035] An ordinal spatial relationship refers to the spatial arrangement order of spatial objects or events, including front-back, left-right, up-down, and east-west-north-south directional relationships, and a metric spatial relationship refers to relationships such as distance or perspective between spatial objects.

[0036] The clock attribute in a spatiotemporal information pair refers to giving the same time identifier to pairs of spatiotemporal information collected at the same time, and a specific method includes embedding a timestamp into each spatiotemporal information pair, where the timestamp includes but is not limited to year, month, day, hour, minute, second, and millisecond, and the timestamp is used to record the exact time of collecting multidimensional data and provide accurate references on various time dimensions for subsequent data processing and analysis.

[0037] The tag attributes in the spatiotemporal information pair include, but are not limited to, collecting information such as identifier information of the device to which the multidimensional data belongs, spatial object name or category, behavioral patterns (e.g., behavioral patterns of people in a spatial scene, behavioral patterns of various devices, such as robots, etc.), scene state, audio features, and odor type (e.g., gas, dangerous gases such as sulfur dioxide are one odor type, oxygen is another odor type, etc.). The tag attributes provide high-level semantic information to the spatiotemporal information pair, allowing the data processing side to understand and analyze the scene information in detail.

[0038] On the other hand, the data storage side is used to store pairs of spatiotemporal information collected at the same time in chronological order in pairs so that the pairs of spatiotemporal information collected at the same time are mutually verified, mutually complemented, and form a data format that satisfies data processing, and to establish a time-based index.

[0039] Preferably, the data storage side stores the spatiotemporal information pairs collected at the same time in pairs in two adjacent stacks in chronological order, each time identifier includes one paired data item, and the specific storage architecture includes, but is not limited to, a distributed storage architecture. All spatiotemporal information pairs are distributed and stored on multiple nodes of the distributed storage architecture, and each node processes the spatiotemporal information pair individually, thereby realizing parallel processing and load balancing of the spatiotemporal information pairs.

[0040] In specific storage, as a preferred form, data storage of spatiotemporal information pairs adopts a data organization method based on composite key-value pairs, in which the composite key includes the spatial object identifier of the spatiotemporal information pair, the collection time identifier, and some feature tags generated by the tag attributes included in the spatiotemporal information pair, and the composite value represents the corresponding spatiotemporal information pair, including but not limited to multimodal data representation (i.e., recording the combination of spatiotemporal information pairs such as visual, auditory, olfactory, etc.), context analysis and relevance analysis (recording the relationship between data), etc., thereby ensuring support for efficient multidimensional data query and search.

[0041] The data processing side processes and analyzes all spatiotemporal information pairs in the data storage side, synchronizes and fuses the spatiotemporal information pairs from various types of data collection devices, and forms a three-dimensional processing result or video stream with depth of field, which is used to recognize and understand complex patterns and events in a 720° full-angle spatial object environment, including, but not limited to, recognizing and tracking spatial objects and their motion states in three-dimensional space, understanding and predicting events, and comprehensively analyzing and simulating the environment.

[0042] Preferably, the data processing side uses artificial intelligence to process and analyze all spatiotemporal information pairs in the data storage side, and the processing and analysis method includes, but is not limited to, deep learning and machine vision. The data processing side uses artificial intelligence to process and analyze all spatiotemporal information pairs in the data storage side, First, performing data preprocessing, including cleaning the collected multidimensional spatiotemporal information pairs, removing noise and irrelevant information, standardizing the visual, auditory, and olfactory spatiotemporal information pairs, and correcting the stereo image pairs; performing spatiotemporal synchronization after data processing, including ensuring temporal synchronization of the data captured by the different types of data collection devices and aligning the data of the different types of data collection devices to ensure their spatial consistency; After the spatiotemporal synchronization, multimodal data fusion is performed to analyze and reconstruct three-dimensional (x, y, z) and four-dimensional (x, y, z, t) spatiotemporal relationships in the form of information pairs, including using deep learning models such as convolutional neural networks and recurrent neural networks to perform feature extraction on the visual, auditory, and olfactory spatiotemporal information pairs, and combining the feature information of different modalities to form a richer representation through fusion algorithms such as weighted average, decision layer fusion, or feature layer fusion; performing 3D reconstruction after multimodal data fusion, including extracting depth information from the stereo image pairs using a stereo matching algorithm, such as block matching or a deep learning-based stereo matching network, and combining the depth information with the visual data to reconstruct objects and scenes in 3D space to form a 3D model or video stream with depth of field; After the 3D reconstruction is completed, performing object recognition and tracking includes using a target detection algorithm such as YOLO or SSD to recognize an object in space, and using a tracking algorithm such as Kalman filter or deep learning tracker to track the motion state of the recognized object; After the object is recognized and tracked, understanding and predicting events includes understanding the occurrence and progression of events by analyzing the object's behavioral patterns and environmental changes, or predicting future events using a sequence prediction model such as a long short-term memory network or a Transformer model; After understanding and predicting the event, analyzing and simulating the environment, including integrating and analyzing multidimensional space-time information pairs, conducting a comprehensive analysis of the environment including analyzing elements such as light, sound, and smell, and utilizing simulation technologies such as virtual reality and augmented reality to simulate the environment and provide an interactive experience; and finally, providing decision support based on the results of the processing and analysis, including providing decision support such as route planning, anomaly detection, or resource allocation to the entire artificial intelligence system.

[0043] Furthermore, the data processing side is further used to review the historical data of the spatial object, perform matching analysis of old and new data, self-learning and optimization, so as to constantly learn from new data and update and optimize the algorithms and models of the data processing side, and the data processing side is further used to actively find abnormalities and errors in the spatiotemporal information pairs based on processing and analysis of all spatiotemporal information pairs in the data storage side, and repair or report them to ensure the quality and reliability of the data, thereby realizing the self-learning and error correction functions of the entire artificial intelligence system and further enhancing the intelligence of the artificial intelligence system.

[0044] The entire process from building an artificial intelligence system based on the above spatiotemporal information pair to data processing and analysis is outlined using Figure 2.

[0045] Step 1: Build an omnidirectional data collection device.

[0046] In other words, an environmental data collection system is constructed that can collect data covering a full 720° angle using visual, auditory, and olfactory data collection devices. The arrangement of each collection device can be adjusted according to needs and is not limited to the arrangement shown in Figure 1.

[0047] Step 2: Collect multidimensional continuous spatiotemporal information pair data in the scene.

[0048] After the construction of the data collection device is completed, the data collection device is used to capture and record multidimensional data within the 720° scene in real time to form a series of spatiotemporal information pairs, where the paired visual collection devices need to focus in the same or approximately the same direction, simulating the way in which both human eyes simultaneously observe any spatial object, and form a stereoscopic image pair. For example, see the structural schematic diagram of the collection target collection by an example of a paired visual collection device shown in Figure 3. The visual information pair collection device (i.e., the paired visual collection device) collects the observation target A and obtains a stereoscopic image pair through the visual information pair overlap area A1 and the visual information pair overlap area A2. The spatial coordinate system is shown in Figure 3.

[0049] Step 3: The space-time information pair is given a spatial relationship, a clock attribute, and a tag attribute.

[0050] After obtaining the spatiotemporal information pair, it is necessary to assign information such as spatial relationship, clock attribute, and tag attribute to the spatiotemporal information pair. Spatial relationship includes topological spatial relationship, ordinal spatial relationship, and metric spatial relationship. Clock attribute refers to assigning the same time identifier to data pairs such as visual, auditory, and olfactory data collected at the same time. Tag attribute includes, but is not limited to, not only identifier information of the device to which the data source belongs, but also information such as spatial object name, category, behavior pattern, human behavior, and device behavior.

[0051] Step 4: Establish a spatiotemporal information pair storage mechanism.

[0052] After the first three steps are completed, data storage is performed on the data storage side. Regarding data storage, the spatiotemporal information pairs collected at the same time are organized and stored in pairs so that the information pairs can be mutually verified and complemented, and the function of three-dimensional data processing can be achieved.

[0053] Step 5: The spatiotemporal information pair data is fused and analyzed to form a 3D processing result or video stream with depth of field and attribute identifiers.

[0054] After the spatiotemporal information pairs are processed and stored, the final data processing and analysis are performed on the data processing side. Regarding data processing, based on the multidimensional spatiotemporal information pair data, data processing and analysis methods such as deep learning and machine vision artificial intelligence are used to synchronize and fuse the spatiotemporal information pairs from various types of devices to form a three-dimensional processing result or video stream with depth of field, and recognize and understand complex patterns and events in the scene, including, but not limited to, recognizing and tracking spatial objects and their motion states in three-dimensional space, understanding and predicting events, and comprehensively analyzing and simulating the environment.

[0055] The five steps outlined above will not be explained in detail as they can be clearly understood in conjunction with the preceding content.

[0056] Based on the above, the artificial intelligence system based on spatiotemporal information pairs of the present invention uses a data collection device to establish a data collection side covering a 720° full-angle spatial object environment, and this data collection device is used to capture and record multi-dimensional data within the 720° full-angle spatial object environment in real time to form a series of spatiotemporal information pairs.

[0057] The data storage side is used to store pairs of spatiotemporal information collected at the same time in a chronological order in pairs so that the pairs of spatiotemporal information collected at the same time are mutually verified, mutually complemented, and have a data format that satisfies data processing, and to establish a time-based index.

[0058] The data processing side processes and analyzes all spatiotemporal information pairs in the data storage side, synchronizes and fuses the spatiotemporal information pairs from various types of data collection devices, forms a three-dimensional processing result or video stream with depth of field, and recognizes and understands complex patterns and events in the 720° full-angle spatial object environment, including but not limited to recognizing and tracking spatial objects and their motion states in three-dimensional space, understanding and predicting events, and comprehensively analyzing and simulating the environment.

[0059] In response to the problem that the conventional environmental perception model construction uses a single-information input perception method, which is difficult to apply on a large scale in these scenes and cannot provide accurate, stable and comprehensive perception functions, the present invention innovatively combines the information collection, storage and processing methods of the human brain, that is, the information collected, stored and processed by the human brain is generated in pairs, such as information from binocular vision, binaural hearing, binaural olfactory sense, etc. These information pairs not only include visual, auditory and olfactory data collected at the same time, but also can be given specific attributes and identifiers.

[0060] The artificial intelligence system of the present invention analyzes and reconstructs three-dimensional (x,y,z) and four-dimensional (x,y,z,t) spatiotemporal relationships in the form of spatiotemporal information pairs, providing abundant information, improving the robustness of the perception model, increasing complementary information, reducing the difficulty of processing complex scenes, and improving perception accuracy. This meets the current needs for artificial intelligence environmental perception models, and is capable of large-scale applications, especially in application scenarios requiring high precision and high robustness, providing accurate, stable, and comprehensive perception functions. This not only enables efficient understanding and intelligent response to the environment, but also brings about great advances in the application of artificial intelligence technology, and can provide artificial intelligence software and hardware system support for humanoid robots, unmanned vehicles, patrol devices, etc., and is very practical.

[0061] Although preferred embodiments of the present invention have been described, additional variations and modifications may be made to these embodiments by those skilled in the art once they have acquired the basic inventive concept. It is therefore intended that the appended claims be interpreted as including the preferred embodiments and all variations and modifications that are within the scope of the embodiments of the present invention.

[0062] It should be noted that, in this specification, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another and do not necessarily require or imply that such an actual relationship or order exists between those entities or operations. Furthermore, terms such as "comprise," "comprises," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or terminal device that includes a set of elements includes not only those elements but also other elements not expressly listed or elements inherent in such process, method, article, or terminal device. Absent further limitations, an element defined by the phrase "comprises one ..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes that element.

[0063] Although the examples of the present invention have been described with reference to the drawings, the present invention is not limited to the specific embodiments described above, and the specific embodiments described above are not intended to be limiting but are merely illustrative. Those skilled in the art can create many forms under the guidance of the present invention without departing from the purpose of the present invention and the scope protected by the claims, and all of these are subject to the protection of the present invention.

Claims

1. An artificial intelligence system based on spatiotemporal information pairs, comprising: It is used to provide artificial intelligence software and hardware system support for humanoid robots, unmanned vehicles, smart glasses, and patrol devices, including data collection, data storage, and data processing. A data collection terminal is constructed using data collection devices to cover a 720° spatial object environment, the data collection devices are used to capture and record multi-dimensional data in the 720° spatial object environment in real time, and form a series of spatiotemporal information pairs, including a pair of visual collection devices, a pair of auditory collection devices, and a pair of olfactory collection devices, and the pair of visual collection devices focus in the same or approximately the same direction, simulating the way in which both human eyes simultaneously observe any spatial object, and form a stereoscopic image pair; the data storage side is used to store the spatiotemporal information pairs collected at the same time in a paired form in chronological order, so that the spatiotemporal information pairs collected at the same time are mutually verified and complemented, and are in a data format that satisfies three-dimensional data processing, and to establish a time-based index; The data processing side processes and analyzes all spatiotemporal information pairs in the data storage side, synchronizes and fuses the spatiotemporal information pairs from various types of data collection devices, and forms a three-dimensional processing result or video stream with depth of field, which is used to recognize and understand complex patterns and events in the 720° spatial object environment, wherein the complex patterns and events include one or more of: recognition and tracking of spatial objects and their motion states in three-dimensional space, understanding and prediction of events, and comprehensive analysis and simulation of the environment.

2. The pair of visual collection devices must ensure that the visual fields collected by both devices have a sufficient overlapping area when positioned, and the pair of visual collection devices are used to record the position, shape, and movement state of the 720° spatial object environment or the same spatial object in the 720° spatial object environment in real time, and form a visual-spatiotemporal information pair; The paired auditory collecting devices are used to record the audio characteristics of the 720° spatial object environment or the same spatial object in the 720° spatial object environment in real time, and form an audio-spatiotemporal information pair; The olfactory collection device is used to record the odor characteristics of the 720° spatial object environment or the same spatial object in the 720° spatial object environment in real time, and obtain olfactory spatiotemporal information pairs; The visual spatiotemporal information pair, the auditory spatiotemporal information pair, and the olfactory spatiotemporal information pair at the same time are mutually verified and complemented to form a data format that satisfies data processing, and are used for three-dimensional data reconstruction processing by the data processing side; the pair of vision gathering devices includes one or more of a camera, a laser radar; the pair of hearing collection devices includes a microphone; The artificial intelligence system of claim 1 , wherein the paired olfactory collection device includes a gas sensor.

3. The pair of vision acquisition devices are focused in the same or nearly the same direction to form a stereoscopic image pair, and the acquisition frequency is adjusted according to demand; The paired auditory collection devices use microphones distributed at various positions to capture sound waves from various directions, thereby realizing omnidirectional sound source location and voice feature extraction; The artificial intelligence system of claim 2, wherein the paired olfactory collection device monitors changes in gas distribution and concentration in the environment using gas sensors placed at various locations to recognize and track the source and diffusion path of specific gases.

4. the spatiotemporal information pair includes a spatial relationship; The spatial relationships refer to topological, ordinal, and metric spatial relationships between spatial objects, and the topological spatial relationships refer to association, adjacency, containment, intersection, overlap, and separation relationships between spatial objects; The spatial order relationship refers to the spatial arrangement order of spatial objects or events, including front-back, left-right, up-down, and east-west-north-south directional relationships; The artificial intelligence system of claim 1 , wherein the metric spatial relationship refers to a distance or perspective relationship between spatial objects.

5. the spatiotemporal information pair further includes a clock attribute; The clock attribute refers to pairs of spatiotemporal information collected at the same time being given the same time identifier, and a specific method includes embedding a timestamp into each pair of spatiotemporal information, the timestamp including one or more of year, month, day, hour, minute, second, and millisecond, and the timestamp is used to record the exact time of collecting multidimensional data and provide accurate references on various time dimensions for subsequent data processing and analysis.

6. The spatiotemporal information pair further includes a tag attribute; The tag attributes include collecting one or more of identifier information of a device to which the multidimensional data belongs, a spatial object name or category, a behavior pattern, a scene state, a sound feature, and an odor type information; The artificial intelligence system of claim 1 , wherein the tag attributes provide high-level semantic information for spatiotemporal information pairs, so that the data processing side can finely understand and analyze scene information.

7. The data storage side is used to store pairs of spatiotemporal information collected at the same time in a paired manner, and the paired storage manner includes storing the pairs in two adjacent stacks in chronological order, and each time identifier includes one paired data item; the specific storage architecture includes a distributed storage architecture; 2. The artificial intelligence system of claim 1, wherein all spatiotemporal information pairs are stored in a distributed manner on multiple nodes of the distributed storage architecture, and each node processes the spatiotemporal information pairs individually, thereby achieving parallel processing and load balancing of the spatiotemporal information pairs.

8. 2. The artificial intelligence system of claim 1, wherein the data storage method for the spatiotemporal information pairs includes adopting a data organization method based on composite key-value pairs, wherein the composite key includes a spatial object identifier of the spatiotemporal information pair, a collection time identifier, and several feature tags generated by tag attributes included in the spatiotemporal information pair, and the composite value represents the corresponding spatiotemporal information pair, and includes one or more of support for multimodal data representation, context analysis and relevance analysis, and multidimensional efficient data query and search.

9. The data processing side uses artificial intelligence to process and analyze all spatiotemporal information pairs in the data storage side, and the processing and analysis method includes one or more of deep learning and machine vision; The method in which the data processing side processes and analyzes all spatiotemporal information pairs in the data storage side using artificial intelligence is as follows: data pre-processing, including cleaning the collected multidimensional spatiotemporal information pairs, removing noise and irrelevant information, standardizing the visual, auditory, and olfactory spatiotemporal information pairs, and correcting and processing the stereo image pairs; Spatiotemporal synchronization includes ensuring temporal synchronization of data captured by different types of data collection devices, aligning the data of the different types of data collection devices, and ensuring consistency in their spatiotemporal relationships; Multimodal data fusion that analyzes and reconstructs 3D (x, y, z) and 4D (x, y, z, t) spatiotemporal relationships in the form of information pairs, including using deep learning models to perform feature extraction on visual, auditory, and olfactory spatiotemporal information pairs, and combining feature information from various modalities through a fusion algorithm to form a richer representation; and 3D reconstruction, which involves extracting depth information from the stereo image pairs using a stereo matching algorithm and combining the depth information with the visual data to reconstruct objects and scenes in 3D space to form a 3D model or video stream with depth of field; object recognition and tracking, including tracking the motion state of the recognized object using a target detection algorithm and a tracking algorithm; Event understanding and prediction, including understanding the occurrence and progression of events by analyzing object behavior patterns and environmental changes, or predicting future events using sequence prediction models; Analysis and simulation of an environment, including integrating and analyzing multidimensional space-time information pairs, conducting a comprehensive analysis of the environment, and simulating the environment using simulation technology to provide an interactive experience; and decision support that provides decision support to the artificial intelligence system based on the results of the processing and analysis.

10. The data processing side is further used to perform retrospective analysis of the spatial object's historical data, matching analysis of new and old data, self-learning and optimization, so as to constantly learn from new data and update and optimize the algorithm and model of the data processing side; The artificial intelligence system of claim 1, characterized in that the data processing side is further used to actively discover abnormalities and errors in the spatiotemporal information pairs based on processing and analysis of all spatiotemporal information pairs in the data storage side, and repair or report them to ensure data quality and reliability.

Citation Information

Patent Citations

  • Position identification assistance system and position identification assistance method

    WO2023042412A1