Spatial representation method and system based on four-dimensional multi-sensing spatio-temporal data framework
By acquiring multimodal perception data in three-dimensional space, performing active object detection and classification, and combining multimodal fusion to generate spatiotemporal context data, the problem of insufficient accuracy and realism of existing spatial representation methods is solved, achieving comprehensive and realistic spatial representation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHIJING SPACETIME (SHENZHEN) TECHNOLOGY CO LTD
- Filing Date
- 2026-02-12
- Publication Date
- 2026-05-15
AI Technical Summary
Existing spatial representation methods cannot fully and realistically reflect the situation in three-dimensional space, resulting in insufficient accuracy and realism in spatial representation.
A method based on a four-dimensional multi-sensory spatiotemporal data framework is adopted. By acquiring data from multiple different modal perception dimensions in the target three-dimensional space, active object detection and classification are performed, multimodal fusion is carried out, and finally the data is stored as spatiotemporal context data.
It achieves the organic integration of multi-dimensional and multi-type data in three-dimensional space, generating spatial representation data that can comprehensively and realistically reflect the actual situation of space, and meeting the data needs of intelligent space applications such as enterprise offices and commercial exhibitions.
Smart Images

Figure CN122045467A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, and in particular to a spatial representation method and system based on a four-dimensional multi-sensory spatiotemporal data framework. Background Technology
[0002] Spatial representation integrates data within a space to provide comprehensive data support for intelligent space applications across various scenarios, including enterprise offices, commercial exhibitions, cultural and tourism displays, product sales, education and training, and family and clan heritage. However, current spatial representation data is incomplete and fails to fully and accurately reflect the situation within three-dimensional space. Consequently, current spatial representation methods often fail to provide the representation data needed for actual intelligent space applications, reducing the accuracy and realism of spatial representation. Summary of the Invention
[0003] The main objective of this disclosure is to propose a spatial representation method and system based on a four-dimensional multi-sensory spatiotemporal data framework, which can improve the accuracy and authenticity of spatial representation.
[0004] To achieve the above objectives, a first aspect of this disclosure proposes a spatial representation method based on a four-dimensional multi-sensory spatiotemporal data framework, comprising: Data acquired by multiple acquisition devices with different modal sensing dimensions along multiple different time periods within the target three-dimensional space are obtained to obtain four-dimensional multi-sensory data along the time dimension within the target three-dimensional space. Based on the four-dimensional multi-sensory data, active objects in the target three-dimensional space are detected and classified to obtain the classification results of the active objects; Based on the classification results of the activity object and the four-dimensional multi-sensory data, multimodal fusion is performed to obtain the spatiotemporal context data of the target in three-dimensional space; The spatiotemporal context data within the target three-dimensional space is stored to obtain the spatial representation data of the target three-dimensional space.
[0005] In some embodiments, acquiring data from multiple acquisition devices with different modal sensing dimensions to obtain four-dimensional multi-sensory data along the time dimension in the target three-dimensional space includes: Data is acquired from at least two of the following dimensions: visual, auditory, tactile, olfactory, gustatory, EEG, ECG, EMG, and EEG, to obtain four-dimensional multi-sensory data along the time dimension in the target's three-dimensional space.
[0006] In some embodiments, when the four-dimensional multi-sensory data includes visual, auditory, tactile, olfactory, and gustatory dimensions, acquiring data from at least two of the following dimensions—visual, auditory, tactile, olfactory, gustatory, EEG, ECG, EMG, and EEG—to obtain four-dimensional multi-sensory data along the time dimension within the target three-dimensional space includes: The visual data in the target three-dimensional space is obtained by acquiring at least one of the following: screen content, spatial structure, spatial objects, lighting, character objects, eye movement information, and posture estimation information, along the time dimension: The auditory acquisition device collects at least one of the following in the target three-dimensional space: sound source location and orientation, audio content, ambient sound, audio tone, and audio quality, to obtain auditory data in the target three-dimensional space along the time dimension. The tactile data is obtained by collecting at least one of the following in the target three-dimensional space: temperature, physical contact, air quality, vibration information, and object texture, along the time dimension: tactile data is collected by the tactile acquisition device in the target three-dimensional space. The olfactory acquisition device collects at least one of the odor, air composition, and pleasure information in the target three-dimensional space to obtain olfactory data in the target three-dimensional space along the time dimension. The taste sensor collects at least one of the taste, taste description, and food texture in the target three-dimensional space to obtain taste data in the target three-dimensional space along the time dimension. Based on the visual data, auditory data, tactile data, olfactory data, and gustatory data, four-dimensional five-sense data are obtained along the time dimension in the three-dimensional space of the target.
[0007] In some embodiments, the detection and classification of active objects in the target three-dimensional space based on the four-dimensional multi-sensory data to obtain the classification result of the active objects includes: Based on multiple sets of four-dimensional multi-sensory data, the visual features, motion features, interaction features, and thermal features of each active object in the target three-dimensional space are extracted; Scoring is performed based on the visual features, motion features, interaction features, and thermal features corresponding to each of the aforementioned activity objects, and a weighted sum is calculated based on the scoring results to obtain the multimodal score of each of the aforementioned activity objects; The classification result of each activity object is determined based on the multimodal score of each activity object.
[0008] In some embodiments, the step of extracting visual features, motion features, interaction features, and thermal features of each active object in the target three-dimensional space based on multiple sets of four-dimensional multi-sensory data includes: Based on multiple sets of four-dimensional multi-sensory data, shape detection, texture detection, color detection, face detection, skeleton detection, digital human marker detection, and robot marker detection are performed, and the visual features of each active object in the target three-dimensional space are obtained according to the detection results. Based on multiple sets of four-dimensional multi-sensory data, velocity detection, acceleration detection, trajectory smoothness detection, motion pattern detection, and stationary time detection are performed, and the motion characteristics of each active object in the target three-dimensional space are obtained according to the detection results. Based on multiple sets of four-dimensional multi-sensory data, voice detection, gesture recognition detection, touch interaction detection, eye contact detection, intention and motivation detection, and emotion expression detection are performed, and the interaction characteristics of each active object in the target three-dimensional space are obtained according to the detection results. Body temperature and heat distribution are detected based on multiple sets of four-dimensional multi-sensory data, and the thermal characteristics of each active object in the target three-dimensional space are obtained based on the detection results.
[0009] In some embodiments, the multimodal fusion based on the classification result of the active object and the four-dimensional multisensory data to obtain spatiotemporal context data in the target three-dimensional space includes: The four-dimensional multi-sensory data is time-synchronized to a unified timestamp. Based on the classification results of the activity objects, the time-synchronized four-dimensional multi-sensory data is spatially aligned to a unified spatial world coordinate system. The spatially aligned four-dimensional multi-sensory data is subjected to data verification processing to filter out data that does not meet the verification criteria. The validated four-dimensional multi-sensory data is then standardized to conform to a unified format. The classification results of the activity object and the standardized four-dimensional multi-sensory data are fused in a multimodal manner to obtain the spatiotemporal context data of the target in three-dimensional space.
[0010] In some embodiments, the data verification process of the spatially aligned classification results of the active object and the four-dimensional multi-sensory data includes: The classification results and four-dimensional multi-sensory data of the spatially aligned active objects are subjected to integrity verification, type verification, range verification, consistency verification, and format verification. Among them, the integrity verification is used to check whether the required fields exist, the type verification is used to check whether the data type is correct, the range verification is used to check whether the data value is within a reasonable range, the consistency verification is used to check the consistency between data, and the format verification is used to check whether the data format conforms to the specification. The process of standardizing the classification results of the activity objects after data verification and the four-dimensional multi-sensory data includes: The classification results and four-dimensional multi-sensory data of the activity objects after data verification are subjected to unit standardization, coordinate system standardization, data format standardization, coding standardization and precision standardization. Among them, unit standardization is used to unify units, coordinate system standardization is used to unify coordinate systems, data format standardization is used to unify data formats, coding standardization is used to unify coding, and precision standardization is used to unify numerical precision.
[0011] In some embodiments, storing the spatiotemporal context data within the target three-dimensional space to obtain spatial representation data of the target three-dimensional space includes: The spatiotemporal context data within the target three-dimensional space is stored according to the preset SpatiotemporalContextV2 unified data structure. The SpatiotemporalContextV2 data structure includes a core identifier, four-dimensional information, multi-sensory data, a list of active objects, and metadata. The core identifier includes a context ID and a version number. The four-dimensional information integrates the time dimension, the three-dimensional spatial dimension, and the location information of the active objects. The profile of each active object in the list of active objects integrates its basic information, physical attributes, capability attributes, status information, historical information, and relationship information. The data of the SpatiotemporalContextV2 structure is associated with the spatial identifier of the target three-dimensional space to form a structured storage with space as the dimension, thus obtaining the spatial representation data.
[0012] In some embodiments, forming a spatially dimensional structured storage includes: A sharded storage mechanism is established for different data modules in the SpatiotemporalContextV2 structure. The sharded storage mechanism performs time sharding of four-dimensional information according to time windows, with each shard corresponding to spatiotemporal data of a fixed duration; it performs type sharding of multi-sensory data according to the perception types of vision, hearing, touch, smell, taste, EEG, ECG, EMG, and EOG; and it performs object sharding of the list of active objects according to the object types of human, digital human, and robot roles. Each shard establishes a global association index through context ID and spatial identifier to achieve fast aggregation of sharded data and cross-shard query.
[0013] In some embodiments, associating the data of the SpatiotemporalContextV2 structure with the spatial identifier of the target three-dimensional space includes: When storing data in the SpatiotemporalContextV2 structure, a version management field and an extended field are added to the data. The version management field is used to record the data structure version and update log. The extended field uses a key-value pair format to reserve an interface for compatibility with future additions of perception dimension data, activity object types, or metadata fields. The storage process establishes multi-dimensional indexes, including a time index built by timestamp buckets, a spatial index built by quadtree partitioning based on spatial bounding boxes, an object index associated with object ID and object type, and a multi-sensory data feature index that builds inverted indexes for key features such as brightness, volume, and temperature. This enables the spatial representation data to support fast fusion queries of multi-dimensional combinations, and the version management field enables forward and backward compatibility and expansion of different versions of data.
[0014] In some embodiments, after storing the spatiotemporal context data within the target three-dimensional space to obtain spatial representation data of the target three-dimensional space, the spatial representation method based on a four-dimensional multi-sensory spatiotemporal data framework further includes: Receive multi-dimensional query parameters, wherein the query parameters include at least any two of the following combinations: time dimension parameters, spatial dimension parameters, activity object dimension parameters, and multi-sensory data dimension parameters; The query parameters are parsed, and the spatial representation data is filtered according to the corresponding dimensions. This includes filtering data based on time dimension parameters to match timestamps or time windows, filtering data based on spatial dimension parameters to match spatial ranges or spatial types, filtering data based on activity object dimension parameters to match object types or attributes, and filtering data based on multi-sensory data dimension parameters to match perceptual feature thresholds. At least two of these parameters are then used to obtain the corresponding filtering results. After performing an intersection operation on the filtering results of each dimension to achieve multimodal fusion, the fusion results are sorted according to preset rules, and multimodal query results are returned based on the sorting results.
[0015] To achieve the above objectives, a second aspect of the present disclosure proposes a spatial representation system based on a four-dimensional multi-sensory spatiotemporal data framework, including multiple data acquisition devices and processing devices; Among them, multiple data acquisition devices are set in the target three-dimensional space, and each data acquisition device is used to acquire data in the corresponding modal perception dimension; The processing device is used to execute the spatial representation method based on a four-dimensional multi-sensory spatiotemporal data framework as described in the first aspect embodiment above.
[0016] To achieve the above objectives, a third aspect of this disclosure proposes a spatial representation system based on a four-dimensional multi-sensory spatiotemporal data framework, comprising: The four-dimensional multi-sensory data acquisition module is used to acquire data from acquisition devices with multiple different modal sensing dimensions along multiple different times in the target three-dimensional space, and obtain four-dimensional multi-sensory data along the time dimension in the target three-dimensional space. The active object detection and classification module is used to detect and classify active objects in the target three-dimensional space based on the four-dimensional multi-sensory data, and obtain the classification result of the active objects; The spatiotemporal context construction module is used to perform multimodal fusion based on the classification result of the active object and the four-dimensional multisensory data to obtain the spatiotemporal context data in the target three-dimensional space; The spatial representation module is used to store the spatiotemporal context data within the target three-dimensional space to obtain the spatial representation data of the target three-dimensional space.
[0017] To achieve the above objectives, a fourth aspect of the present disclosure provides an electronic device, the electronic device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the spatial representation method based on a four-dimensional multi-sensory spatiotemporal data framework described in the first aspect embodiment.
[0018] To achieve the above objectives, a fifth aspect of the present disclosure provides a storage medium, which is a computer-readable storage medium storing a computer program that, when executed by a processor, implements the spatial representation method based on a four-dimensional multi-sensory spatiotemporal data framework described in the first aspect embodiment.
[0019] The beneficial effects of the embodiments disclosed herein include: By first acquiring data from multimodal sensing devices at different times within the target's three-dimensional space, four-dimensional multi-sensory data is formed, integrating the time dimension and multiple sensing dimensions. This breaks through the limitations of traditional methods, which suffer from single data dimensions and incomplete information, from the source of data acquisition. Then, based on the four-dimensional multi-sensory data, objects moving within the space are detected and classified, accurately identifying and defining relevant information about these objects. This avoids incomplete representations caused by missing object information. Subsequently, multimodal fusion is performed by combining the classification results of the objects with the four-dimensional multi-sensory data to generate spatiotemporal context data that reflects the spatiotemporal correlation characteristics of the target's three-dimensional space. This achieves the organic integration of multi-dimensional and multi-type data, allowing the data to comprehensively and realistically reflect the actual situation of the three-dimensional space in the time dimension. Finally, this spatiotemporal context data is stored as spatial representation data, ensuring that the final output spatial representation data is integrated and complete, reflecting the true spatiotemporal state of the space. This fully meets the actual data needs of intelligent space applications such as enterprise offices and commercial exhibitions. The entire process from data acquisition, processing, fusion to final output solves the problem of incomplete spatial representation data, thereby improving the accuracy and realism of spatial representation. Attached Figure Description
[0020] Figure 1 This is a flowchart illustrating the spatial representation method based on a four-dimensional multi-sensory spatiotemporal data framework provided in this embodiment of the disclosure; Figure 2 This is a flowchart illustrating the process of collecting four-dimensional and five-sense data according to an embodiment of this disclosure; Figure 3 yes Figure 1 A flowchart further includes step S102; Figure 4 yes Figure 3 A flowchart further included in step S301; Figure 5 yes Figure 1 A flowchart further includes step S103; Figure 6 yes Figure 1 A flowchart further includes step S104; Figure 7 yes Figure 1 A flowchart illustrating the further steps following step S104; Figure 8 This is a schematic diagram of the framework of a spatial representation system based on a four-dimensional multi-sensory spatiotemporal data framework provided in an embodiment of this disclosure; Figure 9 This is a schematic diagram of the functional modules of a spatial representation system based on a four-dimensional multi-sensory spatiotemporal data framework provided in an embodiment of this disclosure; Figure 10This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this disclosure. Detailed Implementation
[0021] The accompanying drawings in the embodiments clearly and completely describe the technical solutions in the embodiments of this disclosure. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0022] It is understood that in the specific embodiments of this disclosure, data and related data obtained by multiple acquisition devices with different modal perception dimensions are involved. When the above embodiments of this disclosure are applied to specific products or technologies, permission or consent from the subject is required, and the collection, use and processing of related data must comply with relevant laws, regulations and standards.
[0023] Furthermore, when this embodiment of the disclosure needs to retrieve data and related data obtained by multiple acquisition devices with different modal perception dimensions, it will obtain separate permission or separate consent for the data and related data obtained by multiple acquisition devices with different modal perception dimensions through pop-up windows or jumps to a confirmation page. After clearly obtaining separate permission or separate consent for the data and related data obtained by multiple acquisition devices with different modal perception dimensions, it will then obtain the necessary data and related data obtained by multiple acquisition devices with different modal perception dimensions to enable the embodiments of this disclosure to operate normally.
[0024] In this disclosure, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0025] Please see Figure 1 , Figure 1 This is a flowchart illustrating the spatial representation method based on a four-dimensional multi-sensory spatiotemporal data framework provided in this embodiment. This spatial representation method based on a four-dimensional multi-sensory spatiotemporal data framework can be applied to a spatial representation system (hereinafter referred to as the system) based on a four-dimensional multi-sensory spatiotemporal data framework. The spatial representation method based on the four-dimensional multi-sensory spatiotemporal data framework includes steps S101 to S104: Step S101: Acquire data from multiple acquisition devices with different modal sensing dimensions along multiple different times in the target three-dimensional space to obtain four-dimensional multi-sensory data along the time dimension in the target three-dimensional space. Step S102: Detect and classify active objects in the target's three-dimensional space based on four-dimensional multi-sensor data to obtain the classification results of the active objects; Step S103: Based on the classification results of the active object and the four-dimensional multi-sensory data, perform multimodal fusion to obtain the spatiotemporal context data of the target in the three-dimensional space; Step S104: Store the spatiotemporal context data within the target's three-dimensional space to obtain the spatial representation data of the target's three-dimensional space.
[0026] Regarding step S101 above, the target three-dimensional space refers to the physical or virtual three-dimensional area that needs to be spatially represented, such as corporate office areas, commercial exhibition halls, cultural and tourism venues, family residences, shooting studios, teaching venues, etc., and its range can be defined by parameters such as spatial coordinates and bounding boxes.
[0027] Different modal perception dimensions include at least perception dimensions based on the five human senses, such as at least visual, auditory, tactile, olfactory, and gustatory perception dimensions. Each dimension corresponds to a specific type of physical perception, which can comprehensively capture environmental and interactive information within the space. This embodiment only uses five perception dimensions as an example for illustration. Under the premise of meeting the requirements of this embodiment, the number of perception dimensions can be less than five, such as at least two of visual, auditory, tactile, olfactory, and gustatory perception. The number of perception dimensions can also be more, and data from other dimensions can be acquired in addition to visual, auditory, tactile, olfactory, and gustatory perception dimensions, such as data from electroencephalogram (EEG), electrocardiogram (ECG), electromyogram (EMG), and electrooculogram (EOG). This embodiment does not impose specific limitations on this.
[0028] Acquisition devices, also known as data acquisition devices, are specialized data acquisition hardware corresponding to various sensory dimensions. These include visual acquisition devices, such as cameras, depth cameras, and LiDAR; auditory acquisition devices, such as microphones and audio sensors; tactile acquisition devices, such as temperature sensors, pressure sensors, and air quality sensors; olfactory acquisition devices, such as chemical sensors and odor sensors; gustatory acquisition devices, such as taste sensors and component analysis sensors; and electroencephalogram (EEG), electrocardiogram (ECG), electromyogram (EMG), and electrooculogram (EOG) acquisition devices to acquire data from other dimensions. These acquisition devices are used to accurately collect raw data for the corresponding dimensions.
[0029] Data acquired by acquisition devices along multiple time periods and using multiple modal sensing dimensions within the target's three-dimensional space can be termed four-dimensional sensing data under each modality. Combining these four-dimensional sensing data under each modality yields four-dimensional multi-sensory data along the time dimension within the target's three-dimensional space. Therefore, four-dimensional multi-sensory data is a composite data type that integrates the time dimension, the three-dimensional spatial dimension, and the multi-sensory dimension. It contains both the original information of each sensing dimension within the space and is associated with the timestamps of data acquisition, enabling it to dynamically reflect the multi-dimensional state of the target's three-dimensional space at different points in time, thus overcoming the limitation of traditional spatial data that only focuses on a single dimension.
[0030] It should be noted that the embodiments of this disclosure collect multimodal sensory data along multiple different times in order to capture the temporal dynamic changes of spatial state and avoid the defect that static data cannot reflect the spatial temporal characteristics; integrating multi-sensory dimension data realizes all-round perception of the spatial environment, solves the problem of single dimension and one-sided information in traditional spatial data, and lays the data foundation for the subsequent construction of a complete and realistic spatial representation.
[0031] Regarding step S102 above, the active object refers to a subject with dynamic characteristics within the target three-dimensional space. For example, the active object can primarily include three types of roles: Human, Avatar, and Robot. Its core characteristic is the ability to move, interact, or change state, making it a key element to capture in spatial representation. In addition, the active object can also include other living organisms, such as pets, and dynamic environmental entities from other environments; this disclosure does not impose specific limitations on this aspect.
[0032] Among them, Human, Avatar, and Robot are core active object types within the target three-dimensional space that possess dynamic characteristics such as positional movement, interactive behavior, or state changes. These three types of subjects correspond to the categories of physical real life forms, virtual digital characters, and intelligent automated devices, respectively, covering all key dynamic subjects in physical three-dimensional space, virtual three-dimensional space, and virtual-real fusion space. Human refers to natural persons existing in the target three-dimensional physical space; Digital Avatar refers to virtual digital characters generated based on computer graphics, virtual engines, and artificial intelligence algorithms, which can exist in the target virtual three-dimensional space (such as digital twin space, metaverse scene, virtual exhibition space), or be mapped to the physical three-dimensional space through projection, display terminals, AR overlay, etc.; Robot refers to intelligent / automated physical devices equipped with mechanical structures, drive modules, sensing units, and control programs, which are non-human entities active subjects within the physical three-dimensional space, encompassing various types such as service robots, mobile robots, interactive robots, operational robots, and flying robots.
[0033] Detection refers to the process of identifying the existence, location, and basic shape of moving objects in a target three-dimensional space through multimodal data analysis. It relies on cross-validation of data from multiple dimensions, including vision, motion, and interaction, to ensure the accuracy of object detection. Classification refers to the process of categorizing moving objects into predefined classes based on their characteristic differences. The classification results directly determine the dimensions and methods of subsequent object profiling, ensuring that information about different types of objects can be accurately represented.
[0034] It should be noted that the embodiments of this disclosure are based on four-dimensional multi-sensory data for the detection and classification of active objects. This can make full use of the complementarity of multi-dimensional data, avoid the object recognition bias caused by single data types, accurately define the type and basic attributes of active objects in space, solve the problem of missing or misjudged information of active objects in traditional spatial representation, and provide key support for constructing spatial representations that include the dynamic features of objects.
[0035] Regarding step S103 above, multimodal fusion refers to the process of integrating the classification results of the activity object with four-dimensional multisensory data. For example, through a series of operations such as time synchronization, spatial alignment, data verification, and standardization, the format differences and logical conflicts of data from different sources are eliminated, and the organic integration of data is achieved.
[0036] Spatiotemporal context data is the output data after multimodal fusion. It is structured data containing four-dimensional correlation information of time, space, and objects, which can comprehensively reflect the environmental state, object dynamics, and the relationship between the two in the three-dimensional space of the target.
[0037] It should be noted that the multimodal fusion process in this embodiment solves the compatibility problem between different types of data. By establishing a unified data logic and format standard, it integrates scattered multisensory data and object information into a coherent spatiotemporal context, enabling the data to truly reflect the dynamic interaction scene in space, avoiding spatial representation distortion caused by data fragmentation, and greatly improving the accuracy of spatial representation.
[0038] Regarding step S104 above, storage refers to the process of storing spatiotemporal context data in a structured manner according to a preset data structure, such as by establishing a reasonable storage mechanism and index system to ensure the security, accessibility and scalability of the data.
[0039] Spatial representation data is the final result formed after storage. It is a structured data set that can comprehensively and realistically reflect the environmental state, object dynamics, and interaction relationships of the target three-dimensional space in the time dimension. It can directly provide data support for intelligent space applications in multiple scenarios such as enterprise office, commercial exhibition, cultural tourism display, product sales, teaching and training, family and family inheritance.
[0040] It should be noted that in this embodiment, the spatiotemporal context data is stored as spatial representation data, which enables long-term data retention and efficient retrieval. The structured storage mechanism ensures the orderly organization of the data, solving the problems of chaotic and difficult-to-reuse traditional spatial data storage. This allows the spatial representation data to continuously provide accurate and complete data support for various intelligent spatial applications.
[0041] In summary, this embodiment of the present disclosure, through the execution of steps S101 to S104, employs a spatial representation method based on a four-dimensional multi-sensory spatiotemporal data framework. This method first acquires data from multi-modal sensing dimension acquisition devices at different times within the target three-dimensional space, forming four-dimensional multi-sensory data that integrates the temporal and multi-sensory dimensions. This overcomes the limitations of traditional methods, which suffer from single data dimensions and incomplete information, by starting with the data acquisition source. Then, based on the four-dimensional multi-sensory data, it detects and classifies active objects within the space, accurately identifying and defining relevant information about these objects. This avoids incomplete representations caused by missing object information. Finally, it combines the classification results of the active objects with the four-dimensional multi-sensory data to perform multi-modal fusion, generating a representation that accurately reflects the target space. Spatiotemporal context data with three-dimensional spatial spatiotemporal correlation characteristics has achieved the organic integration of multi-dimensional and multi-type data, enabling the data to comprehensively and realistically reflect the actual situation of three-dimensional space in the time dimension. Finally, the spatiotemporal context data is stored as spatial representation data, so that the final output spatial representation data is integrated and can reflect the real spatiotemporal state of space. It can fully meet the actual data needs of intelligent space applications such as enterprise offices, commercial exhibitions, cultural tourism displays, product sales, teaching and training, and family and family inheritance. The entire process from data collection, processing, fusion to final output solves the problem of incomplete spatial representation data, thereby improving the accuracy and authenticity of spatial representation.
[0042] The following is a detailed description of the further contents included in steps S101 to S104 in the embodiments of this disclosure.
[0043] In some embodiments, the process of acquiring data from multiple acquisition devices with different modal sensing dimensions to obtain four-dimensional multi-sensory data along the time dimension in the target three-dimensional space in step S101 above may include: Data is acquired from at least two acquisition devices across multiple dimensions, including visual, auditory, tactile, olfactory, gustatory, EEG, ECG, EMG, EEG, and higher-level sensory dimensions, to obtain four-dimensional multi-sensory data along the time dimension within the target three-dimensional space.
[0044] In one embodiment, this embodiment adds two new human physiological electrical signal perception dimensions, namely electroencephalogram (EEG) and electrocardiogram (ECG), to the traditional five human sense perception dimensions. By selecting at least two of the above seven dimensions and corresponding acquisition devices to carry out data acquisition, and combining the time dimension and three-dimensional spatial coordinate information, four-dimensional multi-sensory data of the target three-dimensional space is finally generated, further enriching the perception dimensions and information coverage depth of the four-dimensional multi-sensory data.
[0045] Among them, the visual dimension corresponds to visual acquisition devices, including cameras, depth cameras, LiDAR, infrared imagers, etc., which are used to acquire visual data such as color images, depth information, spatial point clouds, and infrared thermal imaging in the three-dimensional space of the target, and capture core visual features such as spatial shape, object outline, spatial layout, and appearance of moving objects. The auditory dimension corresponds to auditory acquisition devices, including microphones, array microphones, audio sensors, sound level meters, etc., which are used to collect acoustic data such as ambient noise, voice interaction, and sound generated by object collisions in the space, and to characterize the state of the spatial acoustic environment and the interactive behavior of objects emitting sound. The tactile dimension corresponds to tactile-related data acquisition devices, including temperature sensors, pressure sensors, air quality sensors, humidity sensors, vibration sensors, etc., which are used to collect physical quantity data such as spatial temperature and humidity, air pressure, air quality, contact pressure, and vibration intensity to reflect the tactile perception attributes of the spatial physical environment. The olfactory dimension corresponds to olfactory collection devices, including chemical gas sensors, odor recognition sensors, and volatile organic compound detection sensors, which are used to collect data such as gas composition, odor concentration, and volatile substance content in the space to characterize the odor environment of the space. The taste dimension corresponds to taste-related data acquisition devices, including taste sensors, material composition analysis sensors, and chromatographic detection devices, which are used to collect data such as the taste attributes and chemical composition ratios of substances that can be touched in the space, supplementing the taste perception information of the material characteristics in the space. The EEG dimension corresponds to EEG signal acquisition equipment, including EEG acquisition electrodes, EEG signal acquisition instruments, head-mounted EEG acquisition terminals, etc., which are used to collect EEG waveforms, EEG frequency band energy and other neural electrical signal data of the human body in the target three-dimensional space, reflecting the human body's cognitive state, spatial perception response and neural interaction characteristics in space. The ECG dimension corresponds to ECG signal acquisition devices, including ECG electrodes, ECG acquisition terminals, wearable ECG monitoring devices, etc., which are used to collect cardiac physiological data such as ECG waveforms, heart rate, and heart rate variability of the human body in space, and characterize the physiological state and environmental adaptation characteristics of the human body in the target three-dimensional space.
[0046] The electromyography (EMG) dimension corresponds to EMG signal acquisition devices, including surface electrodes, needle electrodes, EMG signal acquisition instruments, and wireless surface EMG acquisition terminals. These devices are used to acquire neuromuscular data such as EMG waveforms, EMG amplitudes, and muscle activation sequences of specific muscles or muscle groups in a target three-dimensional space. This data reflects the muscle activation state, movement intention, fatigue level, and mechanical output characteristics in human-computer interaction within the space.
[0047] The electrooculogram (EOG) dimension corresponds to EOG signals and eye-tracking acquisition devices, including EOG electrodes, video eye trackers, and head-mounted eye-tracking terminals. These devices are used to synchronously acquire visual behavioral data such as EOG waveforms, eye movement trajectories, fixation points, and pupil diameter changes of the human body in space. This data characterizes the human body's visual attention allocation, cognitive processing load, regions of interest, and environmental information acquisition characteristics in the target three-dimensional space.
[0048] Electromyography (EMG) dimension, also known as skin conductance dimension, corresponds to skin conductance signal acquisition devices, including skin conductance sensors, skin conductance signal acquisition modules, wearable skin conductance monitoring terminals, etc. It is used to collect autonomic neural activity data such as the conductance level (SCL) and conductance response (SCR) of the human skin surface in the target three-dimensional space. It reflects the human body's emotional arousal, cognitive stress level and the excitation state of the sympathetic nervous system in space, and characterizes its immediate physiological and psychological response characteristics to spatial environmental stimuli.
[0049] A higher level of perception corresponds to a multi-channel physiological signal synchronous acquisition system, which integrates multi-dimensional sensors such as EEG, ECG, EMG, EOG, and TEG. It is used to synchronously collect and fuse multimodal physiological data of the human body in space. Through information fusion and machine learning algorithms, it comprehensively analyzes the human body's cognitive state, emotional valence, psychological load, and overall human-machine-environment adaptability in complex environments, achieving a leap from single physiological indicators to comprehensive perception.
[0050] This disclosure specifies that at least two of the above-mentioned dimensions should be selected for data collection. No specific restrictions are placed on the combination of dimensions. It can be a combination of traditional sensing dimensions, or a combination of traditional sensing dimensions and physiological electrical signal dimensions. Three or more dimensions can be selected to form a multi-dimensional sensing array according to the application scenario requirements. This ensures the flexibility and adaptability of the data collection scheme, meets the differentiated sensing needs of different intelligent space scenarios, and avoids information collection blind spots caused by single dimensions or fixed dimension combinations.
[0051] It should be noted that, by incorporating human physiological data such as the five senses, electroencephalogram (EEG), electrocardiogram (ECG), electromyogram (EMG), and electrooculogram (EOG), this embodiment of the present disclosure breaks through the limitation of traditional spatial representation which only collects spatial physical environment data. It realizes the synchronous collection and correlation of spatial environment information and human physiological response information, and captures the dynamic state of the target's three-dimensional space from multiple levels such as physical environment, human interaction, and physiological perception. This further makes up for the problem of single spatial representation data dimension and one-sided information coverage in related technologies, and provides a richer and more complete original data foundation for subsequent object detection and classification and multimodal data fusion.
[0052] Furthermore, when four-dimensional multi-sensory data includes visual, auditory, tactile, olfactory, and gustatory dimensions—a total of five senses—please refer to [link to relevant documentation]. Figure 2 , Figure 2 This is a flowchart illustrating the process of acquiring four-dimensional five-sense data according to an embodiment of this disclosure. In some embodiments, the process of acquiring data from at least two acquisition devices among multiple dimensions including visual, auditory, tactile, olfactory, gustatory, EEG, ECG, EMG, and EEG to obtain four-dimensional multi-sensory data along the time dimension in the target three-dimensional space may further include steps S201 to S206: Step S201: Collect at least one of the following in the target three-dimensional space: screen content, spatial structure, spatial objects, lighting, character objects, eye movement information, and pose estimation information, through a visual acquisition device, to obtain visual data in the target three-dimensional space along the time dimension. Step S202: Acquire at least one of the following in the target three-dimensional space: sound source location and orientation, audio content, ambient sound, audio tone, and audio quality, through an auditory acquisition device, to obtain auditory data in the target three-dimensional space along the time dimension. Step S203: Collect at least one of temperature, physical contact, air quality, vibration information, and object texture in the target three-dimensional space using a tactile acquisition device to obtain tactile data in the target three-dimensional space along the time dimension; Step S204: Collect at least one of the following information in the target three-dimensional space: odor, air composition, and pleasantness information, through the olfactory acquisition device, to obtain olfactory data in the target three-dimensional space along the time dimension; Step S205: Collect at least one of taste, taste description, and food material in the target three-dimensional space using a taste sensor to obtain taste data in the target three-dimensional space along the time dimension; Step S206: Based on visual data, auditory data, tactile data, olfactory data, and gustatory data, obtain four-dimensional five-sense data along the time dimension in the target three-dimensional space.
[0053] In the above steps, the visual acquisition device can continuously acquire data within the target three-dimensional space according to a preset sampling frequency, obtaining temporal visual data bound to timestamps and spatial coordinates. Specifically, screen content refers to visual information such as text, images, videos, and interface interactions output by the display terminal within the space; spatial structure refers to topological information such as walls, beams, columns, partitions, and area divisions; spatial objects refer to the position and shape information of static entities such as furniture, equipment, and furnishings within the space; lighting refers to optical environment information such as light source brightness, color temperature, illumination range, and temporal changes in brightness within the space; character objects refer to the shape, outline, and identity differentiation information of active subjects such as humans, digital humans, and robots within the space; eye movement information refers to the eye movement information of character objects, including visual attention characteristics such as gaze direction, gaze point, and gaze duration; and posture estimation information refers to the body characteristics of character objects such as limb joint angles, movement forms, and behavioral postures. The visual data formed by acquiring at least one of the above information in this step can comprehensively depict the spatial visual environment and the visual characteristics of active objects.
[0054] Auditory acquisition devices can collect and analyze acoustic information from the entire three-dimensional space of a target environment, obtaining temporally and spatially labeled temporal auditory data. This data includes: sound source localization and orientation (the three-dimensional coordinates, orientation, and propagation direction of the sound source); audio content (sound information with analyzable semantics, such as voice dialogues, device announcements, and warning sounds); ambient sound (background noise, natural sounds, and collision sounds); audio tonality (pitch, timbre, loudness, and other acoustic attributes); and audio quality (signal-to-noise ratio, distortion, and clarity). By collecting auditory data from at least one of these acoustic information sources, the acoustic environment and sound-based interactive behaviors in a space can be accurately characterized, supplementing the perception of non-visual dimensions of the space.
[0055] Tactile data acquisition devices can monitor tactile perception-related physical indicators within a target's three-dimensional space in real time, obtaining temporal tactile data. These include: temperature (thermal parameters such as ambient temperature and object surface temperature); physical contact (contact pressure, contact area, and contact state information between the moving object and other objects or spatial interfaces); air quality (air quality parameters such as PM2.5, harmful gas concentrations, and oxygen content); vibration information (dynamic information such as vibration frequency and amplitude of the ground, equipment, and objects); and object texture (physical characteristics such as surface roughness and material feel of physical objects). This step, by acquiring tactile data from at least one of these parameters, can quantitatively reflect the tactile attributes of the spatial physical environment, thus improving the physical perception representation of the spatial environment.
[0056] Olfactory data acquisition devices can analyze odors and related components within a target three-dimensional space to obtain time-series olfactory data. Odor refers to the type, concentration level, and distribution of odors within the space; air composition refers to volatile organic compounds, characteristic gas molecules, and other odor-related chemical components within the space; and pleasantness information refers to related data such as human olfactory comfort levels and sensory pleasure scores obtained based on odor component analysis. Olfactory data obtained from collecting at least one of the above information can characterize the spatial odor environment, filling the data gap in the olfactory dimension of traditional spatial representation.
[0057] Taste sensors can sample and analyze detectable taste-related substances within a target's three-dimensional space to obtain time-series taste data. Taste refers to the basic taste characteristics of a substance, such as sour, sweet, bitter, salty, and umami; taste description refers to quantitative descriptive information such as taste intensity, taste levels, and duration; food texture refers to taste-related attributes such as the food's ingredients and texture. This step, by collecting taste data from at least one of the above information sources, enables the digital representation of the taste characteristics of substances within a space, thus improving the data coverage across the five senses.
[0058] This embodiment of the disclosure can integrate acquired visual, auditory, tactile, olfactory, and gustatory temporal data to obtain four-dimensional five-sensory data that fuses the temporal dimension, three-dimensional spatial dimension, and five-sensory modality dimension. This data, as a typical subset of four-dimensional multi-sensory data, possesses the characteristics of spatiotemporal uniformity, complete modality, and standardized format, providing standardized raw data support for subsequent object detection and classification, and multi-modal fusion. It should be noted that this embodiment, by breaking down the five-sensory data acquisition into a refined, dimensionally segmented execution process, clarifies the acquisition boundaries and data connotations of each sensory dimension, solving the problems of fuzzy content and missing dimensions in multi-sensory acquisition. This enables the generated four-dimensional five-sensory data to comprehensively recreate the human five-sensory perception scene in the target three-dimensional space, further improving the completeness of the raw spatial representation data from the acquisition stage.
[0059] Please see Figure 3 , Figure 3 yes Figure 1 The flowchart further includes step S102. In some embodiments, the process of detecting and classifying moving objects in the target three-dimensional space based on four-dimensional multi-sensor data to obtain the classification result of the moving objects may further include steps S301 to S303: Step S301: Extract the visual features, motion features, interaction features, and thermal features of each active object in the target's three-dimensional space based on multiple four-dimensional multi-sensory data. Step S302: Scoring is performed based on the visual features, motion features, interaction features and thermal features corresponding to each activity object, and the scores are weighted and summed to obtain the multimodal scores of each activity object. Step S303: Determine the classification result of each activity object based on the multimodal score of each activity object.
[0060] In the above steps, the embodiments of this disclosure can rely on multiple sets of four-dimensional multi-sensory data that integrate time dimension, three-dimensional spatial dimension and five-sense perception dimension. Through targeted detection and analysis, four types of core features of each active object can be comprehensively extracted. These include capturing visual features such as shape, texture and face from visual data, mining motion features such as speed, trajectory and motion pattern from spatiotemporal correlation data, recognizing interaction features such as voice, gesture and emotional expression from multi-sensory interaction data, and extracting thermal features such as body temperature and thermal distribution pattern from temperature and thermal distribution data. Finally, through the cross-support of multi-dimensional data, a complete characterization of the active object's attributes is formed, providing a comprehensive and reliable feature basis for subsequent accurate classification.
[0061] Next, in this embodiment of the disclosure, the visual features, motion features, interaction features, and thermal features of each active object are matched and quantified with preset category standards. For example, when facial and skeletal features are detected, the visual feature score of the human category increases; when the movement speed is between 0.5 and 2 meters per second and the trajectory is smooth, the motion feature score of the human category increases; when voice and gesture interaction capabilities are available, the interaction feature scores of both human and digital human categories increase; when the body temperature is between 36 and 37 degrees Celsius, the thermal feature score of the human category increases.
[0062] Weighted summation assigns different weights to visual, motion, interaction, and thermal features based on their importance to the classification result. These weights can be dynamically adjusted according to the actual application scenario. The score of each feature is multiplied by its corresponding weight and then summed to obtain the multimodal total score for each active object.
[0063] The classification result is determined by comparing the multimodal total score of the activity object in the three categories of human, digital human, and robot. The category with the highest total score is taken as the final classification result of the activity object, ensuring that the classification result can accurately reflect the true type of the activity object.
[0064] It should be noted that the multi-feature extraction in this embodiment achieves a comprehensive characterization of the active object, avoiding misclassification caused by a single feature. The weighted scoring classification logic fully considers the classification importance of different features, making the classification results more scientific and accurate, and providing accurate basic information about the object for subsequent object profiling and spatial representation.
[0065] Please see Figure 4 , Figure 4 yes Figure 3The flowchart further includes step S301. In some embodiments, the process of extracting visual features, motion features, interaction features, and thermal features of various active objects in the target three-dimensional space based on multiple four-dimensional multi-sensory data may further include steps S401 to S404: Step S401: Based on multiple four-dimensional multi-sensory data, shape detection, texture detection, color detection, face detection, skeleton detection, digital human marker detection, and robot marker detection are performed, and the visual features of each active object in the target three-dimensional space are obtained according to the detection results. Step S402: Based on multiple four-dimensional multi-sensory data, velocity detection, acceleration detection, trajectory smoothness detection, motion pattern detection, and stationary time detection are performed, and the motion characteristics of each active object in the target three-dimensional space are obtained according to the detection results. Step S403: Based on multiple four-dimensional multi-sensory data, perform voice detection, gesture recognition detection, touch interaction detection, eye contact detection, intention and motivation detection, and emotion expression detection, and obtain the interaction features of each active object in the target three-dimensional space according to the detection results; Step S404: Body temperature detection and heat distribution detection are performed based on multiple four-dimensional multi-sensory data, and the thermal characteristics of each active object in the target three-dimensional space are obtained according to the detection results.
[0066] In the above steps, visual feature extraction includes obtaining the visual features of the moving object by performing shape detection, texture detection, color detection, face detection, skeleton detection, digital human tag detection, and robot tag detection on four-dimensional multi-sensory data. Shape detection distinguishes different shapes such as human skeletons and robot outlines; texture detection identifies surface features such as skin texture and metallic texture; color detection determines color attributes such as skin color and robot-specific colors; face detection and skeleton detection are used to identify human features; digital human tag detection captures the unique identifiers of digital humans, comprehensively characterizing the visual attributes of the moving object; and robot tag detection refers to extracting unique visual features of the physical robot, such as its model identification, mechanical structure features, and hardware component markings.
[0067] Motion feature extraction includes extracting the motion features of moving objects through velocity detection, acceleration detection, trajectory smoothness detection, motion pattern detection, and stationary time detection. Velocity detection records the object's moving speed, acceleration detection captures the rate of change of velocity, trajectory smoothness detection analyzes the continuity of the moving trajectory, motion pattern detection distinguishes different motion modes such as random motion and regular motion, and stationary time detection counts the duration of the object's stationary state, accurately reflecting the motion patterns of the moving object.
[0068] Interaction feature extraction involves using speech detection, gesture recognition detection, touch interaction detection, eye contact detection, intention / motivation detection, and emotion expression detection to obtain the interaction features of the active object. Speech detection determines whether the object has the ability to output speech; gesture recognition detects the object's hand gestures; touch interaction detection records the object's touch behavior; eye contact detection captures the object's eye contact state; intention / motivation detection identifies the object's intentional actions, or combines EEG data to analyze the object's intentional interaction commands; and emotion expression detection identifies the object's facial emotional features, comprehensively reflecting the active object's interaction capabilities.
[0069] Thermal feature extraction involves extracting the thermal characteristics of an active object through body temperature detection and heat distribution detection. Body temperature detection records the object's surface temperature, while heat distribution detection analyzes the object's heat distribution pattern. The human body possesses a stable body temperature of 36 to 37 degrees Celsius and a specific heat distribution, which are key features that distinguish the human body from other objects.
[0070] Please see Figure 5 , Figure 5 yes Figure 1 The flowchart further includes step S103. In some embodiments, the process of obtaining spatiotemporal context data in the target three-dimensional space by multimodal fusion based on the classification results of the active object and four-dimensional multisensory data may further include steps S501 to S505: Step S501: Perform time synchronization processing on the four-dimensional multi-sensory data to synchronize it to a unified timestamp; Step S502: Based on the classification results of the activity objects, the time-synchronized four-dimensional multi-sensory data is spatially aligned to a unified spatial world coordinate system. Step S503: Perform data verification processing on the spatially aligned four-dimensional multi-sensory data to filter out data that does not meet the verification criteria; Step S504: Perform data standardization processing on the verified four-dimensional multi-sensory data to standardize the data into a unified format; Step S505: The classification results of the activity object and the standardized four-dimensional multi-sensory data are fused in a multimodal manner to obtain the spatiotemporal context data in the target three-dimensional space.
[0071] In the above steps, time synchronization processing unifies the timestamps of different acquisition devices and different data types to the same benchmark, such as the system timestamp. By calculating the time difference or using interpolation algorithms, it ensures that the spatial environment data and object status data at the same point in time are accurately correlated, avoiding spatiotemporal misalignment caused by acquisition delay.
[0072] In one embodiment, the four-dimensional multi-sensory data includes four-dimensional sensing data under various different modalities, and the time synchronization processing in step S501 may include: Based on the corresponding timestamps, the time consistency between four-dimensional sensing data in different modalities is judged to obtain the corresponding judgment results. For data whose judgment results indicate time consistency, the four-dimensional sensing data in different modalities are matched according to time windows based on the corresponding timestamps to perform time synchronization processing, resulting in multiple time-synchronized four-dimensional sensing data. For data whose judgment results indicate time inconsistency, the time difference of the four-dimensional sensing data in the corresponding modalities is performed to obtain time-synchronized four-dimensional sensing data.
[0073] The time consistency judgment is based on the timestamp difference of the four-dimensional sensing data of each modality. A preset time difference threshold is used. If the timestamp difference between different modalities of four-dimensional sensing data is less than the threshold, it is judged as time consistent; if the timestamp difference is greater than or equal to the threshold, it is judged as time inconsistent. The time window is a preset fixed-length time interval. The four-dimensional sensing data with bound timestamps are divided into time windows according to the chronological order. Different modalities of four-dimensional sensing data within the same time window are matched to ensure that the sensing data of different modalities are consistent in the time dimension. The matched sensing data is the time-synchronized four-dimensional sensing data. Time interpolation is a time correction method for four-dimensional sensing data with inconsistent times. According to a preset interpolation algorithm, the time dimension of the four-dimensional sensing data with timestamp deviation is interpolated and completed to ensure that the timestamp of the interpolated sensing data is consistent with the sensing data of other modalities. For example, the interpolation algorithm in this embodiment is a linear interpolation algorithm, which can calculate the sensing data corresponding to the target timestamp based on the sensing data and timestamps at adjacent time points. The interpolated sensing data is the four-dimensional sensing data after time synchronization.
[0074] It should be noted that the embodiments of this disclosure achieve time unification of four-dimensional sensing data of different modalities through time consistency judgment and differentiated time synchronization processing, eliminating the time difference of the data. The time-synchronized four-dimensional sensing data provides a data foundation with time dimension consistency for subsequent spatial alignment processing.
[0075] Spatial alignment processing unifies the spatial coordinates of all data into a preset world coordinate system, such as the world right-handed coordinate system, where the X-axis is to the right, the Y-axis is forward, and the Z-axis is upward. Through coordinate transformation algorithms, the local coordinates of different acquisition devices are converted into global coordinates, ensuring that the relative relationship between the object's position and the spatial environment is accurate.
[0076] In one embodiment, the spatial alignment process in step S402 may include: Based on the virtual and real spaces corresponding to each 4D sensing data, a matching coordinate transformation method is determined; the initial spatial coordinates corresponding to the 4D sensing data under different modalities are determined, and under the corresponding coordinate transformation method, the initial spatial coordinates corresponding to the 4D sensing data under different modalities are transformed to the same world coordinate system, resulting in the transformed spatial coordinates of each 4D sensing data; spatial registration processing is performed on the transformed spatial coordinates corresponding to the 4D sensing data under each modality, resulting in the registered spatial coordinates of each 4D sensing data; based on the registered spatial coordinates corresponding to each 4D sensing data, a spatial mapping relationship is established between the 4D sensing data under each modality for spatial alignment processing, resulting in multiple spatially aligned 4D sensing data.
[0077] The coordinate transformation method is adapted to both physical and virtual spaces. Physical space sensing data uses a physical coordinate transformation method, calculating the transformation matrix based on the position and orientation of the acquisition device in the physical space. Virtual space sensing data uses a virtual coordinate transformation method, calculating the transformation matrix based on the coordinate system parameters of the virtual scene. This ensures that sensing data from different physical and virtual spaces can be accurately transformed to a unified world coordinate system. The initial spatial coordinates are the spatial coordinates corresponding to the local coordinate system of each modal acquisition device. Different acquisition devices have different origins and coordinate axis directions in their local coordinate systems, leading to differences in the initial spatial coordinates. The world coordinate system is a pre-defined, unified global spatial coordinate system, using the same coordinate measurement standard and coordinate axis directions, serving as the unified spatial coordinate reference for all modal sensing data. To transform the initial spatial coordinates to the world coordinate system, a transformation matrix needs to be calculated according to the corresponding coordinate transformation method. For example, the transformation matrix is a 4x4 homogeneous transformation matrix from the source coordinate system to the world coordinate system. The matrix contains a rotation matrix and a translation vector. The rotation matrix is used to adjust the direction of the coordinate axes, and the translation vector is used to adjust the origin of the coordinate system. By performing operations on the initial spatial coordinates and the transformation matrix, the transformed spatial coordinates in the world coordinate system can be obtained.
[0078] Spatial registration processing is a further correction of the transformed spatial coordinates, enabling more accurate matching of sensing data from different modalities in spatial dimensions. In some embodiments, the process of performing spatial registration processing on the transformed spatial coordinates corresponding to the four-dimensional sensing data in each modality to obtain the registered spatial coordinates of each four-dimensional sensing data may further include: Initial cross-modal common feature points among four-dimensional sensing data in each modality are determined. Based on the transformed spatial coordinates of the four-dimensional sensing data in each modality, the initial cross-modal common feature points are filtered to obtain the filtered target cross-modal common feature points. Based on the spatial coordinates of the target cross-modal common feature points and a preset transformation algorithm, the corresponding target transformation matrix is calculated. Based on the target transformation matrix, the transformed spatial coordinates of the four-dimensional sensing data in each modality are spatially registered to obtain the registered spatial coordinates of each four-dimensional sensing data.
[0079] In the above steps, initial cross-modal common feature points refer to feature points in four-dimensional perception data from different modalities that point to the same spatial entity. The speaker position detected in the visual modality and the sound source position in the auditory modality are the initial cross-modal common feature points. Based on the transformed spatial coordinates, the spatial distance between each initial cross-modal common feature point is calculated. A preset spatial distance threshold is set, and initial cross-modal common feature points with spatial distances less than the threshold are selected as target cross-modal common feature points to ensure the spatial consistency of feature points.
[0080] The preset transformation algorithm is the iterative nearest-point algorithm. Based on the spatial coordinates of common feature points across modalities of the target, the optimal target transformation matrix is calculated using the iterative nearest-point algorithm. The target transformation matrix is used to perform spatial correction on the transformed spatial coordinates. The transformed spatial coordinates are then processed with the target transformation matrix to obtain the registered spatial coordinates, achieving precise spatial matching of sensing data from different modalities. Spatial mapping refers to the spatial association established between four-dimensional sensing data from different modalities based on the registered spatial coordinates. Based on the registered spatial coordinates, the spatial distance between sensing data from different modalities is calculated. Sensing data with a spatial distance less than a preset threshold are spatially mapped, thus establishing a spatial association between sensing data from different modalities and completing spatial alignment. In this way, through coordinate transformation, spatial registration, and spatial mapping, spatial unification of four-dimensional sensing data from different modalities is achieved, eliminating spatial differences in the data.
[0081] Data validation processing includes integrity validation to check for missing required fields, type validation to confirm that data types conform to specifications, range validation to ensure that data values are within a reasonable range, consistency validation to verify the logical consistency between data, such as whether the location of an object is within spatial boundaries, and format validation to ensure that data formats such as coordinates and time are consistent and to filter out invalid or abnormal data.
[0082] Data standardization includes unit standardization, coordinate system standardization, data format standardization, encoding standardization, and precision standardization, eliminating format differences between different data. In one embodiment, data standardization can also be completed before time synchronization, and this disclosure does not impose specific limitations on this aspect.
[0083] Multimodal fusion, based on the above processing, integrates object classification results with five senses data and spatiotemporal information into a unified data structure, establishes a time, space and object association mapping, and forms complete spatiotemporal context data.
[0084] It should be noted that the multi-step fusion process in the disclosed embodiments ensures the accuracy, consistency and compatibility of the data, eliminates the inherent differences of multimodal data from four dimensions: time, space, data quality and format, and solves the problem of difficulty in traditional multimodal data fusion; the resulting spatiotemporal context data can fully reflect the dynamic scene in space, providing core data support for building high-quality spatial representation.
[0085] Furthermore, in step S503 above, during the data verification process of the spatially aligned classification results of the active objects and the four-dimensional multi-sensory data, the following may be further included: The classification results of spatially aligned active objects and four-dimensional multi-sensory data are subjected to integrity verification, type verification, range verification, consistency verification, and format verification. Among them, integrity verification is used to check whether the required fields exist, type verification is used to check whether the data type is correct, range verification is used to check whether the data values are within a reasonable range, consistency verification is used to check the consistency between data, and format verification is used to check whether the data format conforms to the specification.
[0086] For example, integrity, type, range, consistency, and format validations are performed on the spatially aligned data to filter out invalid or abnormal data: Integrity verification includes checking the existence of required fields such as timestamp, space_id, space_coordinate, vision, hearing, touch, smell, and taste. Data missing any required field is considered invalid. Type validation includes confirming that the data types of each field conform to preset specifications, such as timestamp being an integer or floating-point number, space_id being a string, and space_coordinate being a dictionary type, to avoid affecting subsequent processing due to incorrect data types; Range verification includes ensuring that each data value is within a reasonable range. For example, in one case, the timestamp is between 0 and the next 24 hours, the x and y values of the spatial coordinates are between 0 and 5 meters, the z value is between 0 and 5 meters, the brightness is between 0 and 5000, the volume is between 0 and 140 decibels, and the temperature is between 0 and 30 degrees Celsius. Data that exceeds the reasonable range is marked as abnormal and filtered. Consistency verification includes verifying the logical consistency between data. For example, the value of the time window must be between 0 and 3600 seconds, and the location of the active object must be within the spatial bounding box to avoid logically contradictory data. Format verification includes ensuring that the coordinate data packet has three dimensions (x, y, z) that are all numbers, and that the time format conforms to the preset specifications, to avoid data parsing failure due to format errors.
[0087] Furthermore, in step S504 above, during the process of standardizing the classification results of the verified activity objects and the four-dimensional multi-sensory data to standardize the data into a unified format, the following may be further included: The classification results of the activity objects after data verification and the four-dimensional multi-sensory data are processed for unit standardization, coordinate system standardization, data format standardization, coding standardization and precision standardization. Among them, unit standardization is used to unify units, coordinate system standardization is used to unify coordinate systems, data format standardization is used to unify data formats, coding standardization is used to unify coding, and precision standardization is used to unify numerical precision.
[0088] For example, after data verification, the units, coordinate system, data format, encoding, and precision are standardized to eliminate format differences between different data. Unit standardization includes unifying the unit of length to meters, the unit of time to seconds, and the unit of temperature to degrees Celsius. Through preset unit conversion rules, raw data with different units are converted into a unified unit to ensure data comparability. Coordinate system standardization includes mandating that all spatial data adopt a unified right-handed coordinate system. If the original data is in a left-handed coordinate system, coordinate system transformation is achieved by flipping the Y-axis to ensure the consistency of spatial data. Data format standardization includes unifying all data into JSON format, retaining 3 decimal places for floating-point numbers, and converting time formats to ISO 8601 format, so that the data has a unified parsing rule; Encoding standardization includes unifying text encoding to UTF-8, color encoding to RGB or RGBA format, and audio encoding to PCM format to avoid data parsing anomalies caused by encoding differences; Precision standardization includes retaining coordinate data to 3 decimal places, time data to the millisecond level, angle data to 2 decimal places, and velocity and acceleration data to 2 decimal places, ensuring the consistency of data precision.
[0089] Please see Figure 6 , Figure 6 yes Figure 1The flowchart further includes step S104. In some embodiments, the process of storing the spatiotemporal context data within the target three-dimensional space to obtain the spatial representation data of the target three-dimensional space may further include steps S601 to S602: Step S601: Store the spatiotemporal context data in the target three-dimensional space according to the preset SpatiotemporalContextV2 unified data structure; The SpatiotemporalContextV2 data structure includes a core identifier, four-dimensional information, multi-sensory data, a list of active objects, and metadata. The core identifier includes a context ID and a version number. The four-dimensional information integrates the time dimension, the three-dimensional spatial dimension, and the location information of the active objects. The profile of each active object in the list of active objects integrates its basic information, physical attributes, capability attributes, status information, historical information, and relationship information. Step S602: Associate the data of the SpatiotemporalContextV2 structure with the spatial identifier of the target three-dimensional space to form a structured storage with space as the dimension, and obtain spatial representation data.
[0090] In the above steps, the SpatiotemporalContextV2 unified data structure is a structured data framework specifically designed in this embodiment. The context ID in the core identifier is used to uniquely identify each piece of spatiotemporal context data; the version number records the data structure version, supporting subsequent compatibility and expansion; four-dimensional information integrates the time dimension, three-dimensional spatial dimension, and the location of the active object, comprehensively covering spatiotemporal association information; multi-sensory data is categorized and stored as standardized sensory data according to visual, auditory, tactile, olfactory, and gustatory senses; each object profile in the active object list integrates basic information, physical attributes, capability attributes, status information, historical information, and relationship information; metadata stores auxiliary information such as data source, acquisition device model, and data quality score.
[0091] Spatial identifier association refers to binding SpatiotemporalContextV2 data with a unique identifier in the target three-dimensional space to form a storage system based on space, enabling centralized management and rapid retrieval of all spatiotemporal context data in the same space.
[0092] It should be noted that this embodiment solves the problems of unstandardized and disorganized traditional spatial data storage by using a unified SpatiotemporalContextV2 data structure; the spatial identifier-associated storage method ensures the orderliness and accessibility of the data, enabling spatial representation data to be centrally managed according to spatial dimensions, providing efficient data support for subsequent multi-scenario intelligent applications.
[0093] Furthermore, the process of forming a spatially oriented structured storage in step S602 above may further include: A sharded storage mechanism is established for different data modules in the SpatiotemporalContextV2 structure. The sharded storage mechanism performs time sharding of four-dimensional information according to time windows, with each shard corresponding to spatiotemporal data of a fixed duration; it performs type sharding of multi-sensory data according to the perception types such as vision, hearing, touch, smell, taste, EEG, ECG, EMG, and EEG; and it performs object sharding of the list of active objects according to the object types mainly of human, digital human, and robot roles. Each shard establishes a global association index through context ID and spatial identifier to realize fast aggregation of sharded data and cross-shard query.
[0094] In the above embodiments, time sharding is a storage partitioning method based on the temporal attributes of four-dimensional information. It divides continuous time windows into units of a preset fixed duration, assigning the four-dimensional information within the SpatiotemporalContextV2 structure, such as spatiotemporal coordinates, temporal states, and dynamic changes, to the corresponding time window shards according to the data acquisition timestamp. Each time shard independently carries spatiotemporal associated data for a fixed time period. This sharding rule divides long-series, large-capacity spatiotemporal data into discrete temporal storage units, avoiding the problems of high read / write pressure and redundant retrieval range caused by centralized storage of the entire spatiotemporal data. It facilitates the rapid location of spatial representation information for the target time period according to the time dimension.
[0095] Perceptual type sharding is a storage partitioning method based on the modal attributes of multi-sensory data. It divides data into shards according to multiple sensory dimensions, including vision, hearing, touch, smell, taste, EEG, ECG, EMG, EEG, and EMG. The raw acquisition data, parsed and processed data, and feature extraction data corresponding to each sensory type in the SpatiotemporalContextV2 structure are categorized into matching modal shards. This sharding rule aligns with the dimensional differences and processing logic of multi-modal data, enabling the classification and aggregation of data across different sensory dimensions. This facilitates targeted retrieval, verification, updating, or secondary analysis of data for a single sensory type, improving the storage specificity and maintenance convenience of multi-sensory data.
[0096] Object sharding is a storage partitioning method based on the category attributes of activity objects. Using three core activity object categories—humans, digital humans, and robots—as classification benchmarks, it assigns feature data, behavioral data, interaction data, and spatiotemporal correlation data of each type of object in the SpatiotemporalContextV2 structure to the corresponding object shards, achieving independent aggregation of data for different types of activity objects. This sharding rule is compatible with the activity object classification results and can accurately extract the spatial correlation information of the target subject according to object type.
[0097] Each time segment, perception type segment, and object segment adopts a storage logic of physical splitting and logical association. A global association index is established through the context ID built into the SpatiotemporalContextV2 structure and the spatial identifier exclusive to the target 3D space. The context ID is used to bind segmented data of different dimensions under the same spatiotemporal context to achieve unique association of data of multiple modules in a single spatiotemporal scene. The spatial identifier is used to distinguish segmented data of different target 3D spaces to achieve isolation and differentiation of cross-space data.
[0098] It should be noted that the sharded storage mechanism in this embodiment solves the core contradiction between storing and querying large-scale spatiotemporal context data. That is, it solves the problem of dealing with the storage pressure brought by the continuous accumulation of four-dimensional multi-sensory data and information on active objects, while ensuring the efficient response of multi-dimensional combined queries. At the same time, it avoids the break in association caused by data dispersion after sharding, and ultimately provides the underlying support for large-capacity storage and fast querying for intelligent space applications.
[0099] Specifically, the sharded storage mechanism significantly alleviates storage pressure through targeted sharding design. Sharding four-dimensional information by time window breaks down continuous spatiotemporal data into independent units of fixed duration, avoiding storage redundancy caused by excessively large single data blocks and adapting to long-term, high-frequency data acquisition scenarios. Sharding multi-sensory data by five sense types allows for independent storage of data from different sensory dimensions, enabling the retrieval of specific dimension information without loading all sensory data, such as querying only visual or tactile data, reducing storage resources occupied by invalid data. Sharding the list of active objects by object type allows for focused data management of specific objects, primarily humans, digital humans, and robots, avoiding storage chaos caused by mixed information from different object types, and is particularly suitable for complex spatial scenarios where multiple objects coexist.
[0100] Secondly, the sharding design significantly improves query response efficiency. Time sharding allows queries by time period to quickly filter data without traversing the entire dataset, simply by locating the shard corresponding to the time window, greatly reducing the time spent on time-based filtering. Type sharding supports on-demand invocation; for example, in intelligent navigation scenarios, only visual data and location information need to be queried, without loading irrelevant data such as smell and taste, reducing data transmission and parsing costs. Object sharding enables queries targeting specific types of objects, such as those only counting human activity data within a space, to directly locate the corresponding shard, avoiding full data scanning and adapting to the rapid query needs of scenarios such as behavioral analysis and security monitoring.
[0101] Furthermore, the global relational index ensures the integrity and relevance of the data. By establishing cross-shard relationships through context IDs and spatial identifiers, it solves the pain points of data dispersion and difficulty in aggregation in traditional sharded storage. Even if four-dimensional information, multi-sensory data, and active objects belong to different shards, they can be quickly associated with all data in the same spatiotemporal context through the index, ensuring the integrity of the results of multimodal fusion queries and not missing any related dimensions.
[0102] Finally, this mechanism boasts exceptional adaptability and scalability. The sharding rules are deeply integrated with the SpatiotemporalContextV2 data structure, adapting to the data storage needs of various scenarios such as enterprise offices, commercial exhibitions, cultural tourism displays, product sales, education and training, and family and clan inheritance. It can also flexibly expand the number of shards as data volume grows without requiring a complete reconstruction of the overall storage architecture. Furthermore, the independent management feature after sharding ensures that when adding new perception dimensions or object types, only the corresponding shard needs to be added, without affecting the normal use of existing shards.
[0103] Furthermore, in step S602 above, the process of associating the data of the SpatiotemporalContextV2 structure with the spatial identifier of the target three-dimensional space may further include: When storing data in the SpatiotemporalContextV2 structure, add version management fields and extended fields to the data. The version management fields are used to record the data structure version and update log, while the extended fields reserve interfaces in the form of key-value pairs to be compatible with future additions of perception dimension data, activity object types or metadata fields. The storage process establishes multi-dimensional indexes, including time indexes built by timestamp buckets, spatial indexes built by quadtree partitioning based on spatial bounding boxes, object indexes associated with object IDs and object types, and multi-sensory data feature indexes that build inverted indexes for key features such as brightness, volume, and temperature. This enables spatial representation data to support fast fusion queries of multi-dimensional combinations, and version management fields enable forward and backward compatibility and expansion of different versions of data.
[0104] In the above embodiments, the version management field records the data structure version number, update time, and update content. When the data structure is upgraded, the version number can be used to identify the version type of historical data, ensuring that historical data can be parsed normally and achieving forward and backward compatibility.
[0105] The extended fields are in key-value pair format, allowing new fields to be added without modifying the original data structure, adapting to new scenarios and requirements in the future, and improving system flexibility.
[0106] The time index is built by timestamp buckets, dividing data of different time ranges into different buckets, and supports quick data filtering by time period; the spatial index uses quadtree partitioning based on spatial bounding boxes, and supports quick retrieval by spatial range; the object index associates object ID with type, and supports quick query by object ID and type; the multi-sensory data feature index builds inverted indexes for key features such as brightness, volume, and temperature, and supports quick filtering by feature threshold.
[0107] It should be noted that this embodiment addresses the core pain points of existing spatial representation systems, such as rigid data structures, weak scalability, and low efficiency of multi-dimensional queries. Through a compatible expansion mechanism and efficient index design, it ensures both the adaptability of existing data to future new requirements and the rapid response of multi-dimensional combined queries.
[0108] Specifically, the design of version management fields and extended fields fundamentally breaks through the limitations of traditional data structures that are fixed at once and difficult to iterate. Version management fields, by recording data structure versions and update logs, achieve forward and backward compatibility between different versions of data. Older version data can be parsed by newer versions of the system without reconstruction, and newly added fields in new versions will not affect the normal operation of older systems, effectively protecting historical data assets and reducing the maintenance costs of system iteration. Extended fields reserve interfaces in key-value pair format, allowing for the flexible addition of perception dimension data, activity object types, or metadata fields without modifying the original data structure. For example, when adding pain perception dimensions or virtual pet activity object types in the future, the entire storage system does not need to be reconstructed; it can be adapted simply through extended fields. This allows the system to quickly respond to new scenarios and new requirements, extending the lifecycle of the technical solution.
[0109] The construction of multi-dimensional indexes specifically addresses the problem of low efficiency in multimodal data combination queries. The time index is built by timestamp binning, enabling rapid filtering of data within specific time periods and avoiding full data scanning. The spatial index, based on quadtree partitioning, accurately locates information within specific spatial ranges, adapting to spatial range query scenarios. The object index associates object IDs with types, quickly identifying data for specific object types such as people, digital humans, and robots, improving object-dimensional query efficiency. The multi-sensory data feature index builds inverted indexes for key features such as brightness, volume, and temperature, supporting rapid filtering by perception thresholds to meet the precise query needs of multi-sensory dimensions. These indexes work together, allowing multi-dimensional combination queries to quickly aggregate results without traversing the entire dataset, significantly improving query response speed and fully meeting the complex query needs of various scenarios such as intelligent navigation, behavior analysis, and environmental optimization.
[0110] Overall, the embodiments disclosed herein not only ensure the system's compatibility and scalability, allowing the data structure to adapt to future iterations and changes in multiple scenarios, but also improve the availability of data through efficient indexing, enabling spatial representation data to quickly support the decision-making needs of various intelligent spatial applications.
[0111] Please see Figure 7 , Figure 7 yes Figure 1 A flowchart illustrating further steps following step S104. In some embodiments, steps S701 to S703 may also be included after step S104: Step S701: Receive multi-dimensional query parameters. The query parameters shall include at least any two of the following combinations: time dimension parameters, spatial dimension parameters, activity object dimension parameters, and multi-sensory data dimension parameters. Step S702: Parse the query parameters and perform corresponding dimension filtering on the spatial representation data, including filtering data that matches timestamps or time windows based on time dimension parameters, filtering data that matches spatial ranges or spatial types based on spatial dimension parameters, filtering data that matches object types or attributes based on activity object dimension parameters, and filtering data that matches perceptual feature thresholds based on multi-sensory data dimension parameters, and obtain the corresponding filtering results. Step S703: After performing an intersection operation on the filtering results of each dimension to achieve multimodal fusion, sort the fusion results according to preset rules, and return the multimodal query results based on the sorting results.
[0112] In the above steps, the time dimension parameters include time range, time point, and time window duration; the spatial dimension parameters include spatial ID, spatial range, and spatial type; the activity object dimension parameters include object type, object ID, and object attributes; and the multi-sensory data dimension parameters include the feature thresholds for each sensory dimension.
[0113] Query parameter parsing refers to extracting the filtering conditions for each dimension and clarifying the matching rules for each dimension.
[0114] Multi-dimensional filtering refers to filtering data that meets certain conditions based on various dimension indexes: time filtering filters data within a time period based on the time index, spatial filtering filters data within a spatial range based on the spatial index, object filtering filters data based on object type or attribute based on the object index, and multi-sensory filtering filters data based on feature index and feature threshold.
[0115] Result fusion involves performing an intersection operation on the filtering results from each dimension to obtain data that simultaneously meets all filtering conditions. The results are then sorted according to preset rules and returned based on pagination parameters.
[0116] It should be noted that this embodiment overcomes the limitation of traditional spatial data only supporting single-dimensional queries by using multi-dimensional combined queries, and can accurately meet the complex query needs of intelligent spatial applications; intersection operation ensures the accuracy of query results, and sorting and pagination improve the usability of results, enabling spatial representation data to quickly and accurately support intelligent applications.
[0117] In summary, the spatial representation method based on a four-dimensional multi-sensory spatiotemporal data framework in this disclosure is the first to propose integrating the time dimension, three-dimensional spatial dimension, and the orientation information of the active object into a unified four-dimensional information structure, achieving spatiotemporal integrated representation, solving the problem of incomplete data, and realizing complete recording of all representations within the spatial scope through the four-dimensional multi-sensory data framework, providing sufficient data support for intelligent space applications. It is also the first to design a complete multi-sensory data container including vision, hearing, touch, smell, taste, EEG, ECG, EMG, and EEG, realizing unified management of multimodal sensory data and achieving multimodal data fusion. Through the unified multi-sensory data container, effective fusion of visual, auditory, tactile, olfactory, gustatory, EEG, ECG, EMG, and EEG data is achieved, supporting multimodal intelligent decision-making. For the first time, the system categorizes active subjects into three main roles: humans, digital humans, and robots. Each category includes complete kinematic and profiling information. It also introduces the SpatiotemporalContextV2 unified data structure, integrating four-dimensional information, multi-sensory data, and active subjects to form a complete spatiotemporal context, supporting integrated spatiotemporal analysis. Furthermore, it supports multi-scenario applications, including enterprise offices, commercial exhibitions, cultural tourism displays, product sales, education and training, and family and clan heritage applications, all through a unified data structure framework. Version management and extended fields ensure backward compatibility, supporting future expansion for new requirements.
[0118] Please see Figure 8 The present disclosure also provides a spatial representation system based on a four-dimensional multi-sensory spatiotemporal data framework, including multiple data acquisition devices 801 and a processing device 802; Among them, multiple data acquisition devices 801 are set in the target three-dimensional space, and each data acquisition device 801 is used to acquire data in the corresponding modal perception dimension; The processing device 802 is used to execute the spatial representation method based on the four-dimensional multi-sensory spatiotemporal data framework as described in any of the above embodiments.
[0119] In summary, by first acquiring data from multimodal sensing devices at different times within the target's three-dimensional space, four-dimensional multisensory data integrating the temporal and multisensory dimensions is formed. This breakthrough overcomes the limitations of traditional methods, which suffer from single data dimensions and incomplete information, at the data acquisition source. Then, based on this four-dimensional multisensory data, active objects within the space are detected and classified, accurately identifying and defining relevant information about these objects. This avoids incomplete representations caused by missing object information. Subsequently, multimodal fusion is performed by combining the classification results of the active objects with the four-dimensional multisensory data to generate spatiotemporal context data that reflects the spatiotemporal correlation characteristics of the target's three-dimensional space, thus achieving... The system organically integrates multi-dimensional and multi-type data, enabling the data to comprehensively and realistically reflect the actual situation of three-dimensional space in the time dimension. Finally, the spatiotemporal context data is stored as spatial representation data, so that the final output spatial representation data is integrated and can reflect the real spatiotemporal state of space. It can fully meet the actual data needs of intelligent space applications such as enterprise offices, commercial exhibitions, cultural tourism displays, product sales, teaching and training, and family and family inheritance. The entire process from data collection, processing, fusion to final output solves the problem of incomplete spatial representation data, thereby improving the accuracy and authenticity of spatial representation.
[0120] Please see Figure 9 This disclosure also provides a spatial representation system based on a four-dimensional multi-sensory spatiotemporal data framework, which can implement the above-mentioned spatial representation method based on the four-dimensional multi-sensory spatiotemporal data framework. The spatial representation system based on the four-dimensional multi-sensory spatiotemporal data framework includes: The four-dimensional multi-sensory data acquisition module 901 is used to acquire data from acquisition devices with multiple different modal sensing dimensions along multiple different times in the target three-dimensional space, and obtain four-dimensional multi-sensory data along the time dimension in the target three-dimensional space. The moving object detection and classification module 902 is used to detect and classify moving objects in the target three-dimensional space based on four-dimensional multi-sensory data, and obtain the classification results of the moving objects; The spatiotemporal context construction module 903 is used to perform multimodal fusion based on the classification results of the active object and four-dimensional multisensory data to obtain the spatiotemporal context data in the target's three-dimensional space; The spatial representation module 904 is used to store the spatiotemporal context data within the target's three-dimensional space to obtain the spatial representation data of the target's three-dimensional space.
[0121] In summary, the spatial representation system based on a four-dimensional multi-sensory spatiotemporal data framework executes the spatial representation method based on this framework described in the above embodiments. It first acquires data from multimodal sensing devices at different times within the target three-dimensional space, forming four-dimensional multi-sensory data that integrates the temporal and multi-sensory dimensions. This overcomes the limitations of traditional methods, which suffer from single data dimensions and incomplete information, at the data acquisition source. Then, based on this four-dimensional multi-sensory data, it detects and classifies objects moving within the space, accurately identifying and defining relevant information about these objects. This avoids incomplete representations caused by missing object information. Finally, it combines the classification results of the moving objects with the four-dimensional multi-sensory data to perform multimodal fusion, generating a representation system capable of... The spatiotemporal context data, which reflects the spatiotemporal correlation characteristics of the target's three-dimensional space, achieves the organic integration of multi-dimensional and multi-type data. This allows the data to comprehensively and realistically reflect the actual situation of three-dimensional space in the time dimension. Finally, the spatiotemporal context data is stored as spatial representation data, making the final output spatial representation data a complete data that is integrated and can reflect the true spatiotemporal state of space. This can fully meet the actual data needs of intelligent space applications such as enterprise offices, commercial exhibitions, cultural tourism displays, product sales, teaching and training, and family and family inheritance. The entire process from data collection, processing, fusion to final output solves the problem of incomplete spatial representation data, thereby improving the accuracy and authenticity of spatial representation.
[0122] The specific implementation of the spatial representation system based on the four-dimensional multi-sensory spatiotemporal data framework is basically the same as the specific embodiment of the spatial representation method based on the four-dimensional multi-sensory spatiotemporal data framework described above, and will not be repeated here. Subject to meeting the requirements of the embodiments of this disclosure, the spatial representation system based on the four-dimensional multi-sensory spatiotemporal data framework can also be equipped with other functional modules to implement the spatial representation method based on the four-dimensional multi-sensory spatiotemporal data framework described above.
[0123] This disclosure also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned spatial representation method based on a four-dimensional multi-sensory spatiotemporal data framework. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0124] Please see Figure 10 , Figure 10 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 1001 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this disclosure. The memory 1002 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1002 can store operating devices and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1002 and is called and executed by the processor 1001 to execute the spatial representation method based on a four-dimensional multi-sensory spatiotemporal data framework according to the embodiments of this disclosure. Input / output interface 1003 is used to implement information input and output; The communication interface 1004 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 1005 transmits information between various components of the device (e.g., processor 1001, memory 1002, input / output interface 1003, and communication interface 1004); The processor 1001, memory 1002, input / output interface 1003 and communication interface 1004 are connected to each other within the device via bus 1005.
[0125] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described spatial representation method based on a four-dimensional multi-sensory spatiotemporal data framework.
[0126] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0127] The embodiments described in this disclosure are for the purpose of more clearly illustrating the technical solutions of this disclosure and do not constitute a limitation on the technical solutions provided by this disclosure. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by this disclosure are also applicable to similar technical problems.
[0128] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this disclosure, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0129] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0130] Those skilled in the art will understand that all or some of the steps, apparatuses, or functional modules / units in the methods disclosed above can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0131] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in this disclosure and the foregoing drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, apparatus, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0132] It should be understood that in this disclosure, "at least one item" means one or more, and "more than one" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0133] In the several embodiments provided in this disclosure, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0134] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0135] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0136] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0137] The preferred embodiments of the present disclosure have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present disclosure. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and spirit of the present disclosure shall be within the scope of the claims of the present disclosure.
Claims
1. A spatial representation method based on a four-dimensional multi-sensory spatiotemporal data framework, characterized in that, include: Data acquired by multiple acquisition devices with different modal sensing dimensions along multiple different time periods within the target three-dimensional space are obtained to obtain four-dimensional multi-sensory data along the time dimension within the target three-dimensional space. Based on the four-dimensional multi-sensory data, active objects in the target three-dimensional space are detected and classified to obtain the classification results of the active objects; Based on the classification results of the activity object and the four-dimensional multi-sensory data, multimodal fusion is performed to obtain the spatiotemporal context data of the target in three-dimensional space; The spatiotemporal context data within the target three-dimensional space is stored to obtain the spatial representation data of the target three-dimensional space.
2. The spatial representation method based on a four-dimensional multi-sensory spatiotemporal data framework according to claim 1, characterized in that, The acquisition of data from multiple different modal sensing dimensions by acquisition devices yields four-dimensional multi-sensory data along the time dimension within the target's three-dimensional space, including: Data is acquired from at least two of the following dimensions: visual, auditory, tactile, olfactory, gustatory, EEG, ECG, EMG, and EEG, to obtain four-dimensional multi-sensory data along the time dimension in the target's three-dimensional space.
3. The spatial representation method based on a four-dimensional multi-sensory spatiotemporal data framework according to claim 2, characterized in that, When the four-dimensional multi-sensory data includes visual, auditory, tactile, olfactory, and gustatory dimensions, the acquisition of data from at least two of the following dimensions—visual, auditory, tactile, olfactory, gustatory, EEG, ECG, EMG, and EEG—to obtain four-dimensional multi-sensory data along the time dimension within the target three-dimensional space includes: The visual data in the target three-dimensional space is obtained by acquiring at least one of the following: screen content, spatial structure, spatial objects, lighting, character objects, eye movement information, and posture estimation information, along the time dimension: The auditory acquisition device collects at least one of the following in the target three-dimensional space: sound source location and orientation, audio content, ambient sound, audio tone, and audio quality, to obtain auditory data in the target three-dimensional space along the time dimension. The tactile data is obtained by collecting at least one of the following in the target three-dimensional space: temperature, physical contact, air quality, vibration information, and object texture, along the time dimension: tactile data is collected by the tactile acquisition device in the target three-dimensional space. The olfactory acquisition device collects at least one of the odor, air composition, and pleasure information in the target three-dimensional space to obtain olfactory data in the target three-dimensional space along the time dimension. The taste sensor collects at least one of the taste, taste description, and food texture in the target three-dimensional space to obtain taste data in the target three-dimensional space along the time dimension. Based on the visual data, auditory data, tactile data, olfactory data, and gustatory data, four-dimensional five-sense data are obtained along the time dimension in the three-dimensional space of the target.
4. The spatial representation method based on a four-dimensional multi-sensory spatiotemporal data framework according to claim 1, characterized in that, The detection and classification of active objects within the target's three-dimensional space based on the four-dimensional multi-sensory data, to obtain the classification results of the active objects, includes: Based on multiple sets of four-dimensional multi-sensory data, the visual features, motion features, interaction features, and thermal features of each active object in the target three-dimensional space are extracted; Scoring is performed based on the visual features, motion features, interaction features, and thermal features corresponding to each of the aforementioned activity objects, and a weighted sum is calculated based on the scoring results to obtain the multimodal score of each of the aforementioned activity objects; The classification result of each activity object is determined based on the multimodal score of each activity object.
5. The spatial representation method based on a four-dimensional multi-sensory spatiotemporal data framework according to claim 4, characterized in that, The extraction of visual features, motion features, interaction features, and thermal features of each active object in the target three-dimensional space based on multiple sets of four-dimensional multi-sensory data includes: Based on multiple sets of four-dimensional multi-sensory data, shape detection, texture detection, color detection, face detection, skeleton detection, digital human marker detection, and robot marker detection are performed, and the visual features of each active object in the target three-dimensional space are obtained according to the detection results. Based on multiple sets of four-dimensional multi-sensory data, velocity detection, acceleration detection, trajectory smoothness detection, motion pattern detection, and stationary time detection are performed, and the motion characteristics of each active object in the target three-dimensional space are obtained according to the detection results. Based on multiple sets of four-dimensional multi-sensory data, voice detection, gesture recognition detection, touch interaction detection, eye contact detection, intention and motivation detection, and emotion expression detection are performed, and the interaction characteristics of each active object in the target three-dimensional space are obtained according to the detection results. Body temperature and heat distribution are detected based on multiple sets of four-dimensional multi-sensory data, and the thermal characteristics of each active object in the target three-dimensional space are obtained based on the detection results.
6. The spatial representation method based on a four-dimensional multi-sensory spatiotemporal data framework according to claim 1, characterized in that, The multimodal fusion based on the classification result of the active object and the four-dimensional multisensory data yields the spatiotemporal context data within the target's three-dimensional space, including: The four-dimensional multi-sensory data is time-synchronized to a unified timestamp. Based on the classification results of the activity objects, the time-synchronized four-dimensional multi-sensory data is spatially aligned to a unified spatial world coordinate system. The spatially aligned four-dimensional multi-sensory data is subjected to data verification processing to filter out data that does not meet the verification criteria. The validated four-dimensional multi-sensory data is then standardized to conform to a unified format. The classification results of the activity object and the standardized four-dimensional multi-sensory data are fused in a multimodal manner to obtain the spatiotemporal context data of the target in three-dimensional space.
7. The spatial representation method based on a four-dimensional multi-sensory spatiotemporal data framework according to claim 6, characterized in that, The data validation process for the spatially aligned classification results of the active objects and the four-dimensional multi-sensory data includes: The classification results and four-dimensional multi-sensory data of the spatially aligned active objects are subjected to integrity verification, type verification, range verification, consistency verification, and format verification. Among them, the integrity verification is used to check whether the required fields exist, the type verification is used to check whether the data type is correct, the range verification is used to check whether the data value is within a reasonable range, the consistency verification is used to check the consistency between data, and the format verification is used to check whether the data format conforms to the specification. The process of standardizing the classification results of the activity objects after data verification and the four-dimensional multi-sensory data includes: The classification results and four-dimensional multi-sensory data of the activity objects after data verification are subjected to unit standardization, coordinate system standardization, data format standardization, coding standardization and precision standardization. Among them, unit standardization is used to unify units, coordinate system standardization is used to unify coordinate systems, data format standardization is used to unify data formats, coding standardization is used to unify coding, and precision standardization is used to unify numerical precision.
8. The spatial representation method based on a four-dimensional multi-sensory spatiotemporal data framework according to claim 1, characterized in that, The step of storing the spatiotemporal context data within the target three-dimensional space to obtain spatial representation data of the target three-dimensional space includes: The spatiotemporal context data within the target three-dimensional space is stored according to the preset SpatiotemporalContextV2 unified data structure. The SpatiotemporalContextV2 data structure includes a core identifier, four-dimensional information, multi-sensory data, a list of active objects, and metadata. The core identifier includes a context ID and a version number. The four-dimensional information integrates the time dimension, the three-dimensional spatial dimension, and the location information of the active objects. The profile of each active object in the list of active objects integrates its basic information, physical attributes, capability attributes, status information, historical information, and relationship information. The data of the SpatiotemporalContextV2 structure is associated with the spatial identifier of the target three-dimensional space to form a structured storage with space as the dimension, thus obtaining the spatial representation data.
9. The spatial representation method based on a four-dimensional multi-sensory spatiotemporal data framework according to claim 8, characterized in that, The formation of structured storage with spatial dimensions includes: A sharded storage mechanism is established for different data modules in the SpatiotemporalContextV2 structure. The sharded storage mechanism performs time sharding of four-dimensional information according to time windows, with each shard corresponding to spatiotemporal data of a fixed duration; it performs type sharding of multi-sensory data according to the perception types of vision, hearing, touch, smell, taste, EEG, ECG, EMG, and EOG; and it performs object sharding of the list of active objects according to the object types of human, digital human, and robot roles. Each shard establishes a global association index through context ID and spatial identifier to achieve fast aggregation of sharded data and cross-shard query.
10. The spatial representation method based on a four-dimensional multi-sensory spatiotemporal data framework according to claim 8, characterized in that, The step of associating the data of the SpatiotemporalContextV2 structure with the spatial identifier of the target three-dimensional space includes: When storing data in the SpatiotemporalContextV2 structure, a version management field and an extended field are added to the data. The version management field is used to record the data structure version and update log. The extended field uses a key-value pair format to reserve an interface for compatibility with future additions of perception dimension data, activity object types, or metadata fields. The storage process establishes multi-dimensional indexes, including a time index built by timestamp buckets, a spatial index built by quadtree partitioning based on spatial bounding boxes, an object index associated with object ID and object type, and a multi-sensory data feature index that builds inverted indexes for key features such as brightness, volume, and temperature. This enables the spatial representation data to support fast fusion queries of multi-dimensional combinations, and the version management field enables forward and backward compatibility and expansion of different versions of data.
11. The spatial representation method based on a four-dimensional multi-sensory spatiotemporal data framework according to claim 1, characterized in that, After storing the spatiotemporal context data within the target three-dimensional space to obtain the spatial representation data of the target three-dimensional space, the spatial representation method based on the four-dimensional multi-sensory spatiotemporal data framework further includes: Receive multi-dimensional query parameters, wherein the query parameters include at least any two of the following combinations: time dimension parameters, spatial dimension parameters, activity object dimension parameters, and multi-sensory data dimension parameters; The query parameters are parsed, and the spatial representation data is filtered according to the corresponding dimensions. This includes filtering data based on time dimension parameters to match timestamps or time windows, filtering data based on spatial dimension parameters to match spatial ranges or spatial types, filtering data based on activity object dimension parameters to match object types or attributes, and filtering data based on multi-sensory data dimension parameters to match perceptual feature thresholds. At least two of these parameters are then used to obtain the corresponding filtering results. After performing an intersection operation on the filtering results of each dimension to achieve multimodal fusion, the fusion results are sorted according to preset rules, and multimodal query results are returned based on the sorting results.
12. A spatial representation system based on a four-dimensional multi-sensory spatiotemporal data framework, characterized in that, Includes multiple data acquisition devices and processing units; Among them, multiple data acquisition devices are set in the target three-dimensional space, and each data acquisition device is used to acquire data in the corresponding modal perception dimension; The processing device is used to execute the spatial representation method based on a four-dimensional multi-sensory spatiotemporal data framework as described in any one of claims 1 to 11.
13. A spatial representation system based on a four-dimensional multi-sensory spatiotemporal data framework, characterized in that, include: The four-dimensional multi-sensory data acquisition module is used to acquire data from acquisition devices with multiple different modal sensing dimensions along multiple different times in the target three-dimensional space, and obtain four-dimensional multi-sensory data along the time dimension in the target three-dimensional space. The active object detection and classification module is used to detect and classify active objects in the target three-dimensional space based on the four-dimensional multi-sensory data, and obtain the classification result of the active objects; The spatiotemporal context construction module is used to perform multimodal fusion based on the classification result of the active object and the four-dimensional multisensory data to obtain the spatiotemporal context data in the target three-dimensional space; The spatial representation module is used to store the spatiotemporal context data within the target three-dimensional space to obtain the spatial representation data of the target three-dimensional space.
14. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the spatial representation method based on a four-dimensional multi-sensory spatiotemporal data framework as described in any one of claims 1 to 11.
15. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the spatial representation method based on a four-dimensional multi-sensory spatiotemporal data framework as described in any one of claims 1 to 11.