Methods, systems and media for storing embodied intelligent multimodal data

CN122672722APending Publication Date: 2026-09-01SHANGHAI QIONCHE INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611161027.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-03
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

其次,造成了严重的网络带宽浪费

Benefits of technology

通过将一次采集会话的整体数据,在云端按数据的语义维度和时序单元进行拆分和组织,形成多个可被独立寻址和访问的数据子集及数据单元,下游任务可以根据实际需求,仅选择性地读取其所需的表、列或帧的子集,避免了对整个大型文件的全量下载。这显著降低了冗余数据传输所造成的网络带宽消耗,缓解了数据访问瓶颈,从而缩短了数据处理管线的整体时延,提升了数据处理效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122672722A_ABST
    Figure CN122672722A_ABST
Patent Text Reader

Abstract

This invention provides a method, system, and medium for storing embodied intelligent multimodal data, comprising: receiving raw multimodal data; splitting the data into multiple independent data subsets in the cloud according to semantic dimensions such as sensor type, and organizing time-series data such as camera data by frame, making each frame an independently addressable and decodeable data unit; storing the converted data subsets into object storage; and responding to downstream task requests by selectively reading a subset of specified tables, columns, or frames from the object storage. This application significantly reduces network bandwidth consumption and shortens data processing latency by transforming coarse-grained file access into fine-grained data access, and can flexibly adapt to the dynamic requirements of different algorithms for data frame rate and resolution, thereby improving data utilization efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and more specifically, to a method, system, and medium for storing embodied intelligent multimodal data. Background Technology

[0002] In fields such as embodied intelligence and autonomous driving, model training and iteration rely on processing massive amounts of real-world, multimodal data. A typical processing flow involves uploading multi-sensor data collected by edge devices such as robots or vehicles, including data from multiple cameras, LiDAR, pose measurements, and gripper states, to a cloud-based object storage service (OSS). This data is then used concurrently by various downstream tasks, such as data visualization, automatic annotation, and model training.

[0003] One existing data storage method typically involves serializing all multimodal data from a single acquisition session and storing it in a single large binary container file. For example, raw data, including multi-camera video, depth information, and pose information, is merged into a container file that can reach several gigabytes in size and stored as a whole object in object storage.

[0004] However, this single-container file storage method has significant drawbacks in cloud application scenarios. First, the reading granularity is too coarse, failing to achieve on-demand retrieval. Downstream tasks, even those requiring only partial data from a single sensor (e.g., a few frames from a camera at a specific time), must download the entire gigabyte container file locally before parsing and data extraction. Second, it results in severe network bandwidth waste. When multiple downstream tasks concurrently read data, repeated and redundant downloads of the same large file quickly make the object storage's outbound bandwidth a performance bottleneck for the entire data processing pipeline, significantly increasing overall data processing latency. Finally, this method cannot flexibly adapt to the dynamic needs of algorithms. In the early stages of algorithm development in fields such as embodied intelligence, data specifications, such as frame rate and resolution, dynamically change with algorithm iterations. The single-container format cannot support flexible access modes such as on-demand frame skipping and on-demand resolution selection. Downstream tasks can only read the entire file and then downsample it themselves, further wasting client computing resources and processing time. Summary of the Invention

[0005] To address the shortcomings of existing technologies, the purpose of this invention is to provide a method, system, and medium for storing embodied intelligent multimodal data.

[0006] A method for storing embodied intelligence multimodal data according to the present invention includes: Receive raw acquisition data containing multimodal data from multiple sensors; The original collected data is transformed, the transformation including: splitting the original collected data into multiple independent data subsets according to a preset semantic dimension; organizing at least one of the data subsets at a preset time-series data unit granularity, wherein each time-series data unit forms a data unit that can be independently addressed and accessed; The transformed subsets of data are stored in object storage. In response to a read request from a downstream task, and based on the read request, selectively read one or more data subsets from the data subset, or one or more data units from the data units, from the object storage.

[0007] Furthermore, the preset semantic dimensions include: the type of sensor data stream and / or the acquisition end of the data source.

[0008] Furthermore, the time-series data unit is a camera data frame, and the independently addressable and accessible data unit contains single-frame image data that can be independently decoded.

[0009] Furthermore, the timing data unit also stores the original frame index corresponding to the original acquisition timing sequence; The method further includes: associating the camera data frame with data from other data subsets that have a temporal sequence corresponding to the original frame index, based on the original frame index.

[0010] Furthermore, selective reading methods include: The data unit is read in a frame skipping manner according to a preset step size; The data units are randomly read according to the specified index list; The data unit is read within a range according to a preset time window.

[0011] Furthermore, the method also includes: For at least one data subset in the plurality of data subsets, multiple accompanying data subsets with different resolutions are pre-generated and stored in the object storage; The selective reading includes: selecting to read the accompanying data subset corresponding to the resolution specified in the reading request.

[0012] A storage system for embodied intelligent multimodal data according to the present invention includes: The data receiving module is used to receive raw acquired data containing multimodal data from multiple sensors; A data conversion module is used to convert the original collected data. The data conversion module is configured to: split the original collected data into multiple independent data subsets according to a preset semantic dimension; and organize at least one of the data subsets at a preset time-series data unit granularity, wherein each time-series data unit forms a data unit that can be independently addressed and accessed. The storage interface module is used to store the multiple independent data subsets transformed by the data transformation module into object storage; The data access module is used to respond to read requests from downstream tasks and, based on the read requests, selectively read one or more data subsets from the data subset or one or more data units from the data units in the object storage.

[0013] Furthermore, the data conversion module includes: Parsing submodule: Decompresses and parses the raw data packet, identifying the type, source acquisition terminal, and metadata of each sensor data stream contained in the raw data packet; The splitting submodule: Based on the semantic dimensions identified by the parser, the mixed data stream is split into multiple independent data subsets; Encoding and Organization Submodule: Aligns the camera data stream with the timeline, then breaks it down into frames and encodes each frame into a self-contained data unit that can be decoded independently. It also generates or associates the necessary metadata for each data unit. Table Generator: Formats all data subsets and data units according to a predefined pattern and generates multiple independent data table files that conform to columnar storage format.

[0014] Furthermore, the data access module performs data access at three levels: reading by table, reading by column, and reading by frame.

[0015] According to the present invention, a computer-readable storage medium is provided, wherein a computer program is stored in the computer-readable storage medium, and the program is executed by a processor to implement the method for storing embodied intelligent multimodal data.

[0016] Compared with the prior art, the present invention has the following beneficial effects: By splitting and organizing the entire data from a single acquisition session in the cloud according to the semantic dimensions and temporal units of the data, multiple independently addressable and accessible data subsets and data units are formed. Downstream tasks can selectively read only the subset of tables, columns, or frames they need, avoiding the need to download the entire large file. This significantly reduces network bandwidth consumption caused by redundant data transmission, alleviates data access bottlenecks, thereby shortening the overall latency of the data processing pipeline and improving data processing efficiency.

[0017] Furthermore, this invention supports flexible reading modes such as on-demand frame skipping and on-demand resolution selection, enabling the same stored data to efficiently serve algorithm tasks at different iteration stages and with different specifications, improving the flexibility and reuse efficiency of data utilization, and realizing an efficient data service mode of "write once, read on demand". Attached Figure Description

[0018] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a schematic diagram of the overall data flow of an embodied intelligent data storage method provided in an embodiment of this application; Figure 2 This is a schematic diagram comparing the storage organization methods of the embodiments of this application and the prior art; Figure 3 A schematic diagram illustrating the on-demand selective reading mode provided in an embodiment of this application; Figure 4 A schematic diagram illustrating the data table structure and cross-table relationships provided in the embodiments of this application; Figure 5 A flowchart illustrating an embodied intelligent data storage method provided in this application embodiment; Figure 6 This is a timing diagram of the signaling interaction when the system performs selective reading in the embodiments of this application. Detailed Implementation

[0019] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.

[0020] Example 1 In one embodiment of this application, a method for storing embodied intelligent multimodal data is provided. This method optimizes the organization of cloud data to support efficient, on-demand, and selective reading by downstream tasks. (Reference) Figure 1This illustrates the overall data flow of an embodied intelligent data storage method provided in this application embodiment. Specifically, the data processing flow originates from the end-side robot body 10, which collects a large amount of multimodal sensor data during task execution. This data is then uploaded to the cloud and processed by the cloud conversion module 20. The core function of the cloud conversion module 20 is to split, transform, and organize the raw, typically mixed data according to the method provided in this application embodiment, and then store the structured results in object storage 30. Correspondingly, various downstream tasks, such as data visualization, single-frame annotation / slicing, automatic labeling, automatic quality inspection, metadata storage, model training, or automatic annotation, can request data services according to their specific needs and read only the required subset of data from object storage 30.

[0021] Reference Figure 5 The diagram illustrates a flowchart of an embodied intelligent data storage method provided in an embodiment of this application. Specifically, the method includes the following steps: Step S101: Receive multimodal acquisition data. In a specific application scenario, an embodied intelligent robot, during a data acquisition session of approximately 2 minutes, may generate about 600MB of raw data using its onboard sensors (e.g., multiple cameras, depth sensors, LiDAR, body pose sensors, gripper status sensors, etc. located at different positions on the robot body). This data can be packaged into one or more compressed files, for example, divided into three independent compressed packages according to the acquisition end (e.g., left arm, right arm, and top of the head), and then uploaded to a designated receiving point in the cloud.

[0022] Step S102: Split the data stream by topic and acquisition end. After receiving the raw acquired data, the cloud conversion module 20 first parses it. This parsing process aims to identify the semantic structure within the data. In one embodiment of this application, the preset semantic dimensions may include the type of sensor data stream (often referred to as a topic in fields such as robot operating systems) and the acquisition end (or side) of the data source. For example, the cloud conversion module 20 may identify "image data stream from the overhead camera," "state data stream from the left gripper," and "pose data stream of the robot body," etc. Based on these identified semantic dimensions, the original mixed data stream is split into multiple independent and logically cohesive data subsets, where each combination of "type × acquisition end" constitutes an independent data subset.

[0023] Step S103: Organize the camera data stream by frame. Among all multimodal data, camera data typically has the largest volume and is accessed most frequently by downstream tasks. To achieve fine-grained access to camera data, this embodiment further organizes this specific subset of data, the camera data stream. Specifically, the basic granularity of organization is a temporal data unit, which in this scenario is a camera data frame. The original video stream is decomposed into a series of independent image frames. Each frame is encoded to form a self-contained single-frame image data that can be independently decoded. For example, the H.264 video coding standard can be used, but each frame is encoded as an I-frame (i.e., a keyframe). This way, decoding any frame does not depend on the frames before or after it. This processing method ensures that each frame is a data unit that can be independently addressed and accessed.

[0024] Step S104: Write the organized data into independent columnar storage tables. After processing in steps S102 and S103, all independent data subsets are written into their respective persistent storage structures. As a preferred implementation, this embodiment uses a columnar storage format (e.g., Apache Parquet or Lance format) to create the data tables. Each data subset (i.e., each "type × acquisition end" combination) corresponds to an independent data table file. For example, the "image data stream from the overhead camera" is stored in a table named camera_ego.lance, and the "pose data stream of the robot body" is stored in a table named pose_ego.lance.

[0025] For camera data tables organized frame-by-frame, their structure is particularly critical. (Reference) Figure 4 This demonstrates an example structure for a camera table. Each row in the table corresponds to a data unit, namely a camera frame. Each row contains at least the following columns: a globally unique frame index (frame_index) used for addressing within the table; a high-precision timestamp (timestamp_ns); and a column containing independently decodeable single-frame image data (e.g., rgb_h264). In this way, each frame in the camera data can be uniquely located in the table.

[0026] All generated columnar storage table files are ultimately written to object storage 30 in the cloud, organized, for example, according to a path structure of "session ID / table name.lance". This contrasts with the background technique of packaging all data into a single container file 110 that can reach several gigabytes in size (e.g., ...). Figure 2 Compared to the prior art 100 shown on the left, the method proposed in the embodiments of this application (such as...) Figure 2The present invention 200 shown on the right decomposes the data into multiple smaller, more semantically meaningful independent tables, such as camera table 210, low-resolution video table 220, and other topic tables 230, etc.

[0027] Step S105: Respond to the read request from the downstream task. When a downstream task needs to access data, it no longer needs to download the entire session's data packet. Instead, it sends a specific read request to a data access interface (described in detail in subsequent embodiments), which specifies the range of data it needs to access.

[0028] Step S106: Selectively read a specified subset of data from object storage. The data access interface parses the request and translates it into low-level read operations on one or more specific files (i.e., data tables) in object storage 30. Because the data has been organized into independently addressable units, the system can precisely read only the data specified in the request.

[0029] Taking a single-frame annotation task as an example, this task requires examining the image of frame 100. The request specifies reading the row in the `camera_ego.lance` table where `frame_index` is 100. The system quickly locates the physical location and byte range of that row in the object storage file using metadata in columnar storage format, and then initiates only a single read request for that small range of data. Compared to downloading the entire approximately 2GB single container file, the actual data read volume for this task can be reduced to tens of KB, achieving a significant reduction of approximately four to five orders of magnitude.

[0030] Let's take another data visualization task as an example. This task requires previewing the robot's first-person perspective in video format and simultaneously displaying its pose. This task could simply request to read a pre-generated low-resolution video table (such as...). Figure 2 The low-resolution video table 220 and pose_ego.lance table are downloaded instead of the high-resolution camera table 210, which also greatly reduces the amount of data transmitted over the network.

[0031] In summary, the method provided in this embodiment transforms coarse-grained file-level access into fine-grained row and column-level data access. This not only significantly reduces network bandwidth consumption caused by redundant data downloads and alleviates the egress bandwidth bottleneck of object storage, but also enables multiple downstream tasks to access data efficiently and concurrently, thereby shortening the latency of the entire data processing pipeline.

[0032] Example 2 As an optional implementation, this embodiment further elaborates on the diverse and efficient on-demand selective read modes supported by this application, based on the foregoing embodiments. These read modes enable downstream applications to flexibly and accurately obtain the required data according to their complex business logic, without having to implement cumbersome data parsing and filtering logic on the client side. This function is mainly implemented through a data access module, which is responsible for parsing the read requests from the upper-layer application and converting them into specific operations on the underlying columnar storage table.

[0033] refer to Figure 3 This diagram illustrates the on-demand selective reading mode provided in an embodiment of this application. The granularity of data access can be divided into three levels: table-by-table reading, column-by-column reading, and frame-by-frame reading.

[0034] Reading by table is the most basic selective reading. When downstream tasks are only concerned with data from specific sensors, such as an algorithm that only analyzes the robot's motion trajectory, it only needs to read the pose table (pose_ego.lance) and does not need to access the massive camera table at all.

[0035] Reading by column offers greater precision. Once the table to be read is determined, if the task only needs a portion of the information—for example, a quality control task used to calculate camera exposure times—it can request only the `timestamp_ns` and `exposure_time` columns (if they exist) from the camera table, without needing to read the `rgb_h264` column containing image data. Column-oriented storage formats inherently support this column pruning operation, returning only the necessary column data at the storage end, further reducing network traffic.

[0036] Frame-by-frame reading is a core advantage of this application's embodiments. It provides multiple row-level filtering and selection methods for data tables organized by time-series data units (such as camera frames). The following description illustrates the specific working process: 1. Step-by-step frame skipping read: This mode corresponds to Figure 3The example illustrates frame skipping. In the iterative process of embodied intelligence algorithms, it's often necessary to test model performance using data at different frame rates. Assume the original camera data is acquired and stored at 30fps, while a model training task currently needs to process data at 10fps. When initiating a read request, this task can specify a step size of 3. Upon receiving this request, the data access module converts it into a filtering query on the camera table, with the query condition expressed as `frame_index % 3 == 0`. The underlying data reading library efficiently performs this filtering operation, reading only about one-third of the frame data that meets the condition from object storage and returning it to the downstream task. This eliminates the need to download all 30fps data and then downsample it on the client side, achieving efficient frame skipping on the server side. Similarly, by setting different step sizes, data streams with arbitrarily lower frame rates can be easily simulated.

[0037] 2. Random Read by Index List: This mode corresponds to the random access method in frame-by-frame reading. In scenarios such as data quality inspection or sparse annotation, tasks may need to inspect or process a set of discontinuous specific frames. For example, a quality inspector marks the images in frames 10, 50, 152, and 300 as problematic and requiring further analysis. In this case, the downstream task can submit a list containing the indices of these frames [10, 50, 152, 300] as a read request. The data access module will call the random row access interface provided by the underlying library. This interface can directly retrieve these specific rows of data from object storage based on the index list, thereby avoiding full table scans or large-scale data reads.

[0038] 3. Read by Time Window Range: This mode corresponds to the interval scanning method in frame-by-frame reading. When performing video slicing, event analysis, or problem backtracking, tasks typically need to read all data within a specific time period. For example, an analysis task might need to view all robot actions between the 10th and 15th second. Downstream tasks can submit a start timestamp (t_start = 10s) and an end timestamp (t_end = 15s) as request parameters. The data access module first converts these two timestamps into the corresponding frame index range. If the data is 30fps, the corresponding frame index range is approximately frame_index>= 300 and frame_index<450. Then, the data access module uses a range scanner or equivalent filtering conditions to efficiently read this continuous data block.

[0039] refer to Figure 6This demonstrates the signaling interactions during the selective read process described above. Regardless of the read mode, the process is similar: the downstream task sends a read request with specific filtering conditions (such as step size, index list, time range, etc.) to the data access module. The data access module parses the request and sends one or more precise, possibly byte-range-specific, underlying GET requests to the object storage. The object storage returns only the requested data blocks. Finally, the data access module integrates these data blocks, if necessary, and returns them to the downstream task.

[0040] By providing these diverse reading modes, this embodiment significantly enhances the flexibility and efficiency of data usage, enabling downstream applications to accurately retrieve data based on their inherent logic without incurring heavy parsing, filtering, and downsampling tasks on the client side. This not only reduces the complexity of application development and operational overhead but also further amplifies the beneficial effects of bandwidth savings.

[0041] Example 3 This embodiment, based on the foregoing embodiments, further demonstrates how to support two crucial functions in embodied intelligence applications: on-demand resolution selection and multi-sensor data synchronization, by pre-generating an accompanying data table and utilizing association keys.

[0042] First, regarding on-demand resolution selection. Different downstream tasks have different image resolution requirements. For example, an application for quick previews or lightweight visualization interfaces doesn't need to load full-resolution images; using low-resolution images will suffice, resulting in faster loading speeds and lower network overhead. Conversely, a task for detailed automatic annotation or model training requires the highest resolution images to capture all details.

[0043] To meet these diverse needs, the method provided in this embodiment, when performing step S104 (writing the organized data into an independent columnar storage table), in addition to generating a full-resolution camera table (e.g., camera_ego.lance with a resolution of 1280x720, such as...), Figure 2 In the camera table 210, one or more accompanying data subsets with different resolutions are pre-generated for the camera data stream and stored in object storage 30. For example, a low-resolution video table (video_ego_640.lance, e.g., 640x360) can be generated simultaneously. Figure 2 (See Table 220 for low-resolution video). This process can be accomplished by downsampling and recoding the original full-resolution image.

[0044] When a downstream task initiates a read request, it can specify the required resolution in the request. The data access module will then select to read the corresponding subset of accompanying data based on the specified resolution. For example, a lightweight visualization application might directly specify to read the `video_ego_640.lance` table when making a request. The system will then read and return only the data from object storage for this table, which is much smaller than the full-resolution table, thus achieving on-demand selective reading based on resolution (e.g., ...). Figure 3 (The algorithm reads data by resolution). Algorithms requiring finer analysis request to read the `camera_ego.lance` table at full resolution. This "write-once, select resolution as needed" model achieves a good balance between storage cost and read efficiency, allowing the same raw data to serve multiple different types of tasks in the most efficient way.

[0045] Secondly, regarding cross-table association and synchronization of multi-sensor data. Embodied intelligence tasks often require combining information from different sensors for analysis, such as overlaying robot pose data onto camera images or aligning depth information with color images. Understandably, in the raw data acquisition, the sampling frequencies and timestamps of different sensors may not be entirely consistent. Furthermore, the method in this application's embodiments splits data from different sensors into different physical tables, making maintaining their time synchronization a problem that needs to be addressed.

[0046] This embodiment solves this problem by preserving the association key during the data organization phase. See again... Figure 4 The diagram details the data table structure and cross-table relationships. When organizing camera data streams by frame, in addition to generating aligned and re-encoded new data (e.g., uniformly set to 30fps) and assigning a new, consecutive frame index `frame_index` to each frame in the camera table, an additional crucial field must be stored in each row: the original frame index `orig_frame_index` corresponding to the original acquisition timing. This `orig_frame_index` records the original sequence number of the image frame at the time of acquisition or a closely associated timestamp. For other subsets of sensor data (e.g., pose data), when organized into the pose table, their frame index `frame_index` is directly used or aligned to the original acquisition timing.

[0047] In this way, the `orig_frame_index` field forms the join key for the cross-table join. When a task needs to synchronize pose data with camera images, its workflow may include: 1. The task first reads one or more target frames from the camera table, for example, reading the row with frame_index 50. 2. From this row of data, the task retrieves the value of the orig_frame_index field, assuming the value is 158. 3. Next, the task uses this value 158 as a query condition to search for the row with frame_index 158 in the corresponding pose table. 4. The acquired pose data (e.g., records containing position (pos_xyz) and pose quaternion) is pose data that is precisely synchronized with the camera image with frame_index 50 at the original acquisition time.

[0048] Understandably, this association process can be completed automatically by the data access module, remaining transparent to upper-layer applications. The application only needs to request "pose data synchronized with a specific camera frame," and the data access module will perform the aforementioned read and lookup operations.

[0049] Therefore, this embodiment ensures that even if multimodal data is physically split into different tables for storage, their logical time synchronization can still be accurately reconstructed and maintained. This guarantees data availability and integrity, enabling the smooth implementation of complex algorithms based on multi-sensor fusion, thereby greatly enhancing the value of the stored data.

[0050] Example 4 This embodiment describes a system for storing embodied intelligence data at the system level. This system is used to execute the methods described in any of the foregoing embodiments. The system can be deployed on a cloud server cluster, serving as a stable, efficient, and scalable data infrastructure to provide robust data support for the development and iteration of upper-layer embodied intelligence applications.

[0051] The storage system for embodied intelligent data may include the following core modules: 1. Data Receiving Module: This module serves as the system's entry point, responsible for communicating with end-side devices (such as the end-side robot body 10). Its main function is to receive raw multimodal data uploaded from one or more end-side devices. This data is typically in the form of compressed packages (e.g., tar files), containing raw data streams from multiple sensors and multiple acquisition ends during a single acquisition session. The data receiving module must possess high throughput and high reliability to handle the continuous uploading of large-scale data.

[0052] 2. Data Conversion Module: This module is connected to the data receiving module and is a key execution unit for implementing the core technical solution of this application. After receiving the raw data, this module is configured to execute a series of conversion and organization logics. Specifically, it can be divided into several sub-components, including: Parser: Responsible for decompressing and parsing the raw data packets, identifying the type of each sensor data stream contained therein, the source acquisition terminal, and their respective metadata.

[0053] Stream splitter: Based on the semantic dimensions (i.e., topic and acquisition end) identified by the parser, the mixed data stream is split into multiple independent data subsets.

[0054] Encoder and Organizer: Performs special processing on specific subsets of data (especially camera data). For example, it time-aligns the camera data stream, then breaks it down into frames, and calls an encoder (e.g., an H.264-based encoder) to encode each frame into a self-contained data unit that can be independently decoded. Simultaneously, it generates or associates necessary metadata for each data unit, such as a new frame index (frame_index), the original frame index (orig_frame_index), and a timestamp (timestamp_ns).

[0055] Table Generator: Formats all processed data subsets and data units according to a predefined pattern and generates multiple independent data table files that conform to a specific columnar storage format (such as Lance or Parquet).

[0056] exist Figure 1 In this context, the data conversion module's functionality is represented by the cloud-based conversion module 20, which performs... Figure 5 Steps S102, S103, and S104 in the process.

[0057] 3. Storage Interface Module: This module acts as a bridge between the system and the underlying persistent storage, connecting with the data transformation module and object storage 30. On one hand, it is responsible for efficiently and reliably writing all data table files generated by the data transformation module into object storage 30 for persistence. On the other hand, when data needs to be read, it receives underlying read instructions from upper-level modules, translates them into specific API calls to object storage 30 (e.g., GET requests with byte ranges), and then returns the retrieved data blocks.

[0058] 4. Data Access Module: This module serves as the unified service interface for the entire system to downstream tasks. For example... Figure 6 As shown, this module can be represented as a data access module, and its main responsibilities include: Receive and parse high-level, declarative read requests from downstream tasks. These requests can be very flexible, such as "retrieve the camera_ego table for images and corresponding pose data at 5fps between 10 and 15 seconds".

[0059] Translate high-level requests into a specific, low-level sequence of read operations on one or more data tables. This includes determining which tables and columns need to be accessed, and what filtering conditions (such as by step, by index list, by range, etc.) to apply, as described in Example 2.

[0060] Perform cross-table join logic as described in Example 3. When the request involves multi-sensor synchronization, this module automatically handles lookups and matches based on the join key (such as orig_frame_index).

[0061] The instruction storage interface module performs the final low-level read operation.

[0062] After obtaining the data, some integration or deserialization operations may be required to return the data to the requester in a format that meets the requirements of the downstream task.

[0063] The system's workflow can be summarized as follows: An edge robot continuously uploads and collects data. The data receiving module and data conversion module form an automated data processing pipeline, converting the incoming raw data into a structured, on-demand optimized format in real-time or near real-time and storing it in object storage. Simultaneously, multiple downstream applications, such as visualization tools, annotation platforms, and model training jobs deployed on different computing clusters, can efficiently retrieve their respective required data subsets from the same converted data source in parallel and on demand by calling the interfaces provided by the data access module, without interfering with each other.

[0064] It is understandable that the system architecture provided in this embodiment decouples data storage from computing tasks. By adding a one-time conversion overhead when writing data, it gains extremely high efficiency and flexibility for countless subsequent read operations, thereby constructing a cloud data infrastructure that can effectively support large-scale embodied intelligence research and development.

[0065] Furthermore, this application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it can implement the method for storing embodied intelligent data described in any of the above embodiments. The storage medium can be any electronic, magnetic, optical, or other physical device capable of storing program code, such as a read-only memory, random access memory, flash memory, hard disk, or optical disk.

[0066] Those skilled in the art will understand that, besides implementing the system and its various devices, modules, and units provided by this invention in the form of purely computer-readable program code, the same functions can be achieved entirely through logical programming of the method steps, making the system and its various devices, modules, and units of this invention function in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, the system and its various devices, modules, and units provided by this invention can be considered as a hardware component, and the devices, modules, and units included therein for implementing various functions can also be considered as structures within the hardware component; alternatively, the devices, modules, and units for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0067] Specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.

Claims

1. A method for storing embodied intelligent multimodal data, characterized in that, include: Receive raw acquisition data containing multimodal data from multiple sensors; The original collected data is transformed, the transformation including: splitting the original collected data into multiple independent data subsets according to a preset semantic dimension; organizing at least one of the data subsets at a preset time-series data unit granularity, wherein each time-series data unit forms a data unit that can be independently addressed and accessed; The transformed subsets of data are stored in object storage. In response to a read request from a downstream task, and based on the read request, selectively read one or more data subsets from the data subset, or one or more data units from the data units, from the object storage.

2. The method for storing embodied intelligent multimodal data according to claim 1, characterized in that, The preset semantic dimensions include: the type of sensor data stream and / or the acquisition end of the data source.

3. The method for storing embodied intelligent multimodal data according to claim 1, characterized in that, The time-series data unit is a camera data frame, and the data unit that can be independently addressed and accessed contains single-frame image data that can be independently decoded.

4. The method for storing embodied intelligent multimodal data according to claim 3, characterized in that, The timing data unit also stores the original frame index corresponding to the original acquisition timing sequence; The method further includes: associating the camera data frame with data from other data subsets that have a temporal sequence corresponding to the original frame index, based on the original frame index.

5. The method for storing embodied intelligent multimodal data according to claim 1, characterized in that, Selective reading methods include: The data unit is read in a frame skipping manner according to a preset step size; The data units are randomly read according to the specified index list; The data unit is read within a range according to a preset time window.

6. The method for storing embodied intelligent multimodal data according to claim 1, characterized in that, The method further includes: For at least one data subset in the plurality of data subsets, multiple accompanying data subsets with different resolutions are pre-generated and stored in the object storage; The selective reading includes: selecting to read the accompanying data subset corresponding to the resolution specified in the reading request.

7. A storage system for embodied intelligent multimodal data, characterized in that, include: The data receiving module is used to receive raw acquired data containing multimodal data from multiple sensors; A data conversion module is used to convert the original collected data. The data conversion module is configured to: split the original collected data into multiple independent data subsets according to a preset semantic dimension; and organize at least one of the data subsets at a preset time-series data unit granularity, wherein each time-series data unit forms a data unit that can be independently addressed and accessed. The storage interface module is used to store the multiple independent data subsets transformed by the data transformation module into object storage; The data access module is used to respond to read requests from downstream tasks and, based on the read requests, selectively read one or more data subsets from the data subset or one or more data units from the data units in the object storage.

8. The storage system for embodied intelligent multimodal data according to claim 7, characterized in that, The data conversion module includes: Parsing submodule: Decompresses and parses the raw data packet, identifying the type, source acquisition terminal, and metadata of each sensor data stream contained in the raw data packet; The splitting submodule: Based on the semantic dimensions identified by the parser, the mixed data stream is split into multiple independent data subsets; Encoding and Organization Submodule: Aligns the camera data stream with the timeline, then breaks it down into frames and encodes each frame into a self-contained data unit that can be decoded independently. It also generates or associates the necessary metadata for each data unit. Table Generator: Formats all data subsets and data units according to a predefined pattern and generates multiple independent data table files that conform to columnar storage format.

9. The storage system for embodied intelligent multimodal data according to claim 7, characterized in that, The data access module performs data access at three levels: reading by table, reading by column, and reading by frame.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, is used to implement the method for storing embodied intelligent multimodal data as described in any one of claims 1 to 6.