Cross-framework data transmission method and system based on unified representation

Through a unified representation of cross-frame data transmission method, the DataFrame instance is used to perform folder-level operations and data views in two-dimensional table form, which solves the problem of inefficient transmission of multi-source heterogeneous data, and realizes seamless transmission and efficient loading of scientific data across computing frameworks.

CN120263780APending Publication Date: 2025-07-04COMP NETWORK INFORMATION CENT CHINESE ACADEMY OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510171822.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing scientific data workflow systems are unable to effectively handle the complex nested structures of multi-source heterogeneous data and their interoperability requirements across computing frameworks, resulting in inefficient data sharing and collaborative analysis, and frequent format conversions increase storage and transmission costs.

Method used

Through a cross-frame data transmission method with unified representation, data set operations are used to use DataFrame instances, support folder-level operations, and seamless data transmission is achieved through a unified two-dimensional table of data views and data content, providing an interoperable interface between Python and R, reducing serialization and deserialization overhead.

Benefits of technology

It realizes seamless transmission of scientific data across computing frameworks, improves the scalability and flexibility of data loading, reduces operation time overhead, provides a user-friendly operation experience, and reduces the overhead of inter-process communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263780A_ABST
    Figure CN120263780A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-frame data transmission method and system based on unified representation, and belongs to the technical field of cross-frame data transmission. The method comprises the following steps: acquiring a data set of scientific data and a URL list of the data set in a public database through a unique ID of the data set, and constructing a tree data view in a dictionary form according to the URL list of the data set; according to a data view request of a consumer, the tree-shaped data view is converted into a DataFrame instance, and the DataFrame instance is sent to the consumer, so that when the consumer carries out data set operation on the DataFrame instance, the DataFrame instance will return a new DataFrame instance. According to the method, the problems of description and sharing of multi-source heterogeneous data can be solved, so that seamless transmission of scientific data across computing frameworks is realized, an efficient and compatible basic data support environment is provided for scientific research, and researchers can access and process multidisciplinary and multi-type scientific data resources in a scientific data workflow system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cross-frame data transmission, and in particular, to a cross-frame data transmission method and system based on unified representation. Background Art

[0002] With the development of scientific research towards multi-disciplinary collaboration and big data-driven directions, data resources in different fields are presenting the characteristics of multi-source heterogeneity and distributed storage. These scientific data often have problems such as complex structures, diverse formats, and inconsistent semantics, which limit the efficiency of cross-field research and the efficient utilization of data resources. For example, frontier research such as climate change, astronomical observations, and biofabrication requires the integration of various data in fields such as meteorology, astronomy, and biology. However, the lack of unified description and interaction standards between the data leads to huge challenges in data sharing and collaborative analysis.

[0003] Existing scientific data workflow systems mostly focus on task scheduling and workflow management, lacking support for the representation, unified supply, and on-demand transmission of complex scientific data. For example, Apache Airflow supports the reading of structured data (CSV, JSON), semi-structured data (XML, YAML), and unstructured data (TXT, audio files, picture files) in common data sources such as AWS S3, HDFS, and SQL by implementing hook (data request and loading) and operator (data processing) interfaces for different data sources, encapsulating the data content in its unique XCom unified data format or directly opening it using the corresponding file system of cloud storage (such as S3FileSystem). Since this workflow is entirely written in Python, only Python native data structures are usually allowed for transmission and processing in components, making it unable to effectively handle the complex nested structure of scientific data sets and their interoperability requirements across computing frameworks. The open-source platform Galaxy based on the Web also supports scientific data analysis in multiple scientific fields. In this framework, Galaxy also writes classes for different data types using Python. When specifically requesting data content, Galaxy loads the data in the corresponding data format from data sources such as local, URL, and cloud storage based on the temporary file mechanism according to the meta-information. When the data needs to be presented to the user, the data content will be written as HTML and provided to the Web page through an HTTP request. Galaxy also optimizes the data transfer process through caching, multi-threading, and task scheduling strategies. However, using Python native data structures as the carrier for data processing and transmission has relatively large serialization and deserialization overheads. At the same time, to support new file types, the framework needs to be continuously updated for specific file formats, resulting in insufficient scalability and flexibility. More prominently, the data transfer between different computing frameworks often depends on the way of data landing on disk. This method not only increases the storage and transmission costs but also leads to frequent format conversions and performance losses, further reducing the efficiency of cross-framework data collaboration. Summary of the Invention

[0004] In view of the above problems, the present invention proposes a cross-framework data transmission method and system based on unified representation, aiming to solve the description and sharing problems of multi-source heterogeneous data through a unified representation model, thereby realizing seamless transmission of scientific data across computing frameworks and providing an efficient and compatible basic data support environment for scientific research, enabling researchers to access and process multi-disciplinary and multi-type scientific data resources in a scientific data workflow system.

[0005] To achieve the above object, the technical solution of the present invention includes the following content.

[0006] A cross-framework data transmission method based on unified representation, the method comprising:

[0007] Obtain the dataset of scientific data and the list of URLs of the dataset in the public database through the unique ID of the dataset, and construct a tree-shaped data view in the form of a dictionary based on the list of URLs of the dataset; wherein, the dataset is in the form of a folder or a compressed package in the public database;

[0008] Based on the data view request of the consumer, convert the tree-shaped data view into a DataFrame instance, and send the DataFrame instance to the consumer, so that when the consumer performs dataset operations on the DataFrame instance, the DataFrame instance will return a new DataFrame instance; wherein, the new DataFrame instance contains data content or a data view after operations.

[0009] Furthermore, the DataFrame instance is a two-dimensional table with folders and / or files as rows and dictionary elements as columns; wherein, the dictionary elements include: file name, file type, and depth.

[0010] Furthermore, the dataset operations include: file-level dataset operations and folder-level dataset operations; wherein, the folder-level dataset operations include: trigger loading operations and non-trigger loading operations.

[0011] Furthermore, the trigger loading operation includes: the operation of opening one or more unloaded files;

[0012] When the consumer performs the operation of opening one or more unloaded files on the DataFrame instance, the DataFrame instance will return a new DataFrame instance, including:

[0013] The DataFrame instance finds the list of URLs of the unloaded file according to the filtered file name and the mapping between the file name and the URL;

[0014] Initiate an http request based on the list of URLs of the unloaded file to download the unloaded file and cache it in memory; wherein, the file cached in memory is a binary format file;

[0015] Construct and return a new DataFrame instance with files as rows and file contents as columns according to the content of the binary format file.

[0016] Furthermore, the trigger loading operation includes: the operation of opening a file in a format supported by the DataFrame instance that has been loaded in the data content;

[0017] When the consumer operates on the DataFrame instance to open a file in a format supported by the DataFrame instance in the loaded data content, the DataFrame instance will return a new DataFrame instance, including:

[0018] The DataFrame instance finds the URL list of the file to be opened according to the filtered file name and the mapping between the file name and the URL.

[0019] Initiate an http request based on the URL list of the file to be opened to download the file to be opened and cache it in memory; among them, the file cached in memory is a binary format file.

[0020] Construct and return a new DataFrame instance whose rows, columns are equivalent to the rows and columns of the file content according to the content of the binary format file.

[0021] Furthermore, the non-triggered loading operations include: opening a folder operation, filtering files operation, merging data sets operation.

[0022] When the consumer operates on the DataFrame instance for data set operations, the DataFrame instance will return a new DataFrame instance, including:

[0023] When the consumer operates on the DataFrame instance to open a folder, the DataFrame instance finds the sub-files and sub-folders under the folder according to the folder name in the data view, and constructs and returns a new DataFrame instance with the sub-files and sub-folders as the data view.

[0024] When the consumer operates on the DataFrame instance to filter files, the DataFrame instance finds all files and folders that meet the requirements according to the regular expression input by the consumer in the data view, and constructs and returns a new DataFrame instance with the matching results as the data view.

[0025] When the consumer operates on two DataFrame instances to merge data sets, the first DataFrame instance to be merged constructs and returns a new DataFrame instance with the merge of the data views of the two DataFrame instances as the data view.

[0026] Furthermore, when performing cross-frame data transfer between python and R, the method further includes: defining a unified interoperability interface between python and R; among them, the interoperability interface is based on the rpc protocol and the ipc protocol, and the interoperability interface includes:

[0027] The constructor is used to initialize a DataFrame with configuration information;

[0028] get_schema is used to obtain the data view of a dataset;

[0029] open is used to open the folder or file of a dataset;

[0030] flat_open is used to open all the specified files or files in the data view, and expand and flatten the nested interfaces;

[0031] filter is used to perform regular matching and filtering on folders and files;

[0032] concat is used for row concatenation of DataFrames.

[0033] A cross-frame data transmission system based on unified representation, the system includes:

[0034] A data view construction module, which is used to obtain the dataset of scientific data and the URL list of the dataset in the public database through the unique ID of the dataset, and construct a tree-shaped data view in the form of a dictionary according to the URL list of the dataset; wherein, the dataset is in the form of a folder or a compressed package in the public database;

[0035] A data transmission module, which is used to convert the tree-shaped data view into a DataFrame instance based on the data view request of the consumer, and send the DataFrame instance to the consumer, so that when the consumer performs dataset operations on the DataFrame instance, the DataFrame instance will return a new DataFrame instance; wherein, the new DataFrame instance contains data content or the data view after operations.

[0036] An electronic device, the electronic device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the cross-frame data transmission method based on unified representation described in any one of the above is implemented.

[0037] A computer-readable storage medium, characterized in that computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by a processor, the cross-frame data transmission method based on unified representation described in any one of the above is implemented.

[0038] Compared with the prior art, the present invention has at least the following beneficial effects.

[0039] 1. This representation method is applicable to all data formats, that is, all data can be loaded as a DataFrame without special support for the format, and has higher scalability.

[0040] 2. This characterization method is applicable to datasets with any folder structure, meeting more data loading requirements.

[0041] 3. This characterization method supports folder-level dataset operations by ensuring the consistency of the results of triggered loading and non-triggered loading operations, enabling users to load data more conveniently according to their needs.

[0042] 4. By distinguishing between triggered loading and non-triggered loading operations, this characterization method realizes lazy loading, that is, data is not requested during non-triggered loading operations, and data is requested only during triggered loading operations, reducing the time overhead of operations.

[0043] 5. The unified two-dimensional table form of the data view and data content achieves consistency in data presentation, providing a more user-friendly operation experience.

[0044] 6. By adopting the well-known arrow technology, columnar storage and in-memory transmission of data are realized, making the communication overhead between processes running in different languages lower. Additionally, since DataFrame uses binary streams as the underlying data format, separating the transport layer and the computing layer further reduces the serialization and deserialization overhead during inter-process communication. Brief Description of the Drawings

[0045] Figure 1 It is a flowchart of the cross-frame data transmission method based on unified characterization.

[0046] Figure 2 It is a design diagram of DataFrame for scientific data.

[0047] Figure 3 It is a schematic diagram of cross-language interoperability of DataFrame.

[0048] Figure 4 It is a programming flowchart when users use DataFrame.

[0049] Figure 5 It is a flowchart for constructing a data view.

[0050] Figure 6 It is a schematic diagram of the principle of the open operation of DataFrame. Detailed Implementation Manner

[0051] The present invention will be further described in detail below with reference to the accompanying drawings. The examples given are only used to explain the present invention and are not intended to limit the scope of the present invention.

[0052] The cross-frame data transmission method based on unified characterization of the present invention, as Figure 1 shown, includes the following steps 1 to 2.

[0053] Step 1: Obtain the dataset of scientific data and the list of URLs of this dataset in the public database through the unique ID of the dataset, and construct a tree-shaped data view in the form of a dictionary based on the list of URLs of the dataset.

[0054] Since datasets are usually published and shared in the form of folders or compressed packages, the present invention first represents multi-source heterogeneous scientific data resources with a multi-level data view and then supplies it to consumers. Consumers can be the client that trains and persists the data, or a certain data processing component in the workflow system.

[0055] Step 1.1: Obtain the dataset of scientific data through the unique ID of the dataset.

[0056] The dataset of scientific data in the present invention is in the form of a folder or compressed package in the public database and in the form of a list of URLs externally.

[0057] Step 1.2: Create an empty instance of a DataFrame that does not contain valid information according to the ID of this dataset.

[0058] For the unified representation of data content, combined with the data abstraction of DataFrame widely recognized in the field of data transmission and processing, where Figure 2 is the design diagram of DataFrame for scientific data.

[0059] Step 1.3: Construct the data view of this DataFrame instance based on the list of URLs of the dataset.

[0060] The data view of the present invention is a file directory tree at the abstract level and is actually a two-dimensional table with files as rows, file names, file types, and depths as columns. Specifically, when constructing the data view of the DataFrame instance, the content of scientific data files is loaded in the form of a two-dimensional table, that is, a table with named rows and columns with files as rows and file content as columns, and the internal data format is unified as binary stream. When specifically processing the data, combined with the metadata information of the data view and the corresponding data loader, the data can be restored to the original analyzable file. This enables this representation to support any data format without manual update.

[0061] Step 2: Based on the data view request of the consumer, convert this tree-shaped data view into a DataFrame instance and send this DataFrame instance to this consumer, so that when the consumer performs dataset operations on this DataFrame instance, this DataFrame instance will return a new DataFrame instance.

[0062] Different from the data abstraction of general DataFrames, the DataFrame for scientific data supports folder-level dataset operations. Specifically, users can operate on this DataFrame in a way similar to operating on a file system, that is, opening folders, opening files, returning to the parent directory, and constructing a completely new DataFrame for each operation to achieve the nesting of DataFrames, that is, a DataFrame is generated by a DataFrame and can only generate DataFrames. These operations are all based on the operations on the data view. This enables users not to care about the organization method of the dataset stored on the physical cluster, but only need to program facing the DataFrame, thus realizing a DataFrame instance based on the data view and data content organization, and obtaining a unified representation of the dataset.

[0063] To be compatible with folder-level operations on the data content-level DataFrame, three levels of data in the DataFrame need to be designed: folders, files, and data. The folder and file levels are supported by the data view, and the file and data levels are supported by the data content. Among them, the folder-level dataset operations include triggered loading operations and non-triggered loading operations. Triggered loading operations include: opening one or more unloaded files, or opening files in the data content that are in a format supported by the DataFrame. Non-triggered loading operations include: all operations other than triggered loading operations such as opening folders, filtering files, and merging datasets.

[0064] For opening one or more unloaded files, the DataFrame finds the URL list of the files to be loaded according to the filtered file names and the mapping between file names and URLs, constructs a two-dimensional table of data content with files as rows and file contents as columns, that is, the data content, as a new DataFrame instance of the attribute, and loads the scientific data into this two-dimensional table, where the file format in the two-dimensional table is a binary stream.

[0065] For opening files that have been loaded in the data content, if the files in the data content are in a format supported by the DataFrame (common text formats), the DataFrame will load the binary files into the corresponding format content through the relevant loader and can further open the file. This DataFrame will return a new DataFrame instance with the data content of each row and column of the file as the rows and columns of the attribute. At this time, each row and column is equivalent to the rows and columns in the file.

[0066] When performing a non-triggered load operation on a DataFrame, the DataFrame will return a new DataFrame instance that has the data view after the operation as an attribute. For non-triggered operations, the DataFrame clones its data view, and the operation will act directly on the data view. For example, a filtering operation will filter the data view Figure 2 on the dimension table and retain the rows whose file names meet the requirements. By achieving consistency in the results of triggered and non-triggered load operations, that is, both returning new DataFrame instances with new data views or data contents as attributes, the DataFrame supports folder-level data operations.

[0067] To meet the programming languages commonly used in different disciplines, a unified scientific data representation should be able to be used seamlessly across different programming languages. For this purpose, the science-oriented DataFrame provides a unified interoperability interface between Python and R. The interoperability schematic diagram is as follows Figure 3 shown. The interoperability interface is based on the rpc and ipc protocols and includes: a constructor for initializing the DataFrame through configuration information, get_schema for obtaining the data view of the dataset, open for opening the folder or file of the dataset, corresponding to the operations of opening a single file in the data content, opening a folder, and opening an unloaded file, flat_open for opening all the specified or view files and expanding and flattening the nested interfaces, corresponding to opening one or more unloaded files in the triggered operation. filter is used for regular matching and filtering of folders and files, and concat is used for row concatenation of DataFrames.

[0068] Next, a specific implementation will be used to illustrate the cross-framework data transmission method based on unified representation of the present invention.

[0069] The interoperability of this implementation will be illustrated by taking the specific needs of users as an example. The programming flowchart is as follows Figure 4 . The specific needs of a certain user are described as follows: The user needs to load and process the dataset of a certain public scientific database, and requires filtering specific folders and files according to the file organization structure of the dataset, and finally loading the specific file content.

[0070] First, the user inputs the ID of the required scientific dataset and specifies the data source, and creates a DataFrame instance by instantiating an object. The user can also, according to their own needs, specify whether streaming loading is required when loading the data. When creating a DataFrame instance, the constructor will automatically assign a UID as a unique identifier for the instance and automatically start a client for rpc communication in the background.

[0071] For all programming of DataFrame, users only need to operate on the DataFrame instance (hereinafter referred to as df). For example, when obtaining a view of the data set, users run df.get_schema() in the Python environment and df$get_schema() in the R environment. After calling this method, as Figure 5 shown, df will further call the communication client to send a data view request. After receiving the request, the server assembles the data URL list from the scientific data set ID and data source and requests the data from the public data set via HTTP. After receiving the URL list, the list is passed to the URL parser to construct the data view. Specifically, the URL is split by specific characters to form a list of file paths, and the nested relationship of folders is represented by a dictionary. To reduce the serialization-deserialization overhead during transmission, the dictionary-form file directory tree will be converted into a two-dimensional table in columnar storage with files as rows, file names, file types, and depths as columns before transmission. This two-dimensional table is the data view of this df. After the client receives the data view, the content is returned to df and presented to the user. Due to the orderliness of Python dictionaries and two-dimensional tables, directly traversing and printing this two-dimensional table in sequence can obtain the correct file directory organization, reducing the overhead of repeatedly searching for subfolders when correctly reconstructing the nested structure.

[0072] Then, users can perform free open and return operations based on the data view. Using the open command, users can open any folder and file in the data view. The principle of the open operation is as Figure 6As shown, when the user opens a folder, df calls the filter_index method to filter the sub-files in the folder through the depth column in the data view and construct a new data view. Based on this data view, a new sub-df instance is instantiated and returned to the user, thus forming an abstract nested structure. When the user opens a file, df directly requests the data entity from the server through the client. After the server requests the data, it performs preprocessing according to the user's needs and transmits it to the client as required. df assembles the data entity received by the client into data content that can be presented to the user, that is, in the form of a two-dimensional table. Depending on the situation, the content can be a binary stream or other data types. Using the flat_open command, the user can pass in multiple paths, and df will directly load the data content of all files under the paths. If no parameters are passed in, df will open all files in the data view and load them as a two-dimensional table in a tiled form, with each row representing the file content as a binary stream. Using the back command, df will retrieve the parent df according to the UID of the recorded parent df through the dictionary that records the correspondence between the UID and df, thus returning to the previous directory. Using the filter operation. If the operation is currently on a folder, the user can pass in a regular expression to filter the data whose file names in the dataset meet the requirements. For example, entering "\.jpg$" means matching files with the file suffix jpg, so that all jpg format files in the dataset can be filtered out in one step. If the operation is on the file content, the user can directly filter out the lines that meet the requirements.

[0073] The user can also perform further operations on the data content. For example, the select operation allows the user to select certain columns in the file. Concat can add rows to the data, and combined with the filter operation, it can merge any part of the data content of two dfs. These two operations require the row data types to match. After performing basic data operations, the user can also obtain the data in the required format through interfaces such as to_pandas and to_numpy for further processing.

[0074] The DataFrame interface specification for scientific data and the corresponding function summary are shown in Table 1.

[0075] Table 1 DataFrame Interface Specification for Scientific Data and the Corresponding Functions

[0076] get_schema Obtain the data view of the dataset open Open the specified file (folder) flat_open Open all sub-files in the specified path back Return to the parent directory filter Filter files (folders) / data content select Select columns of data content concat Add rows to the data content to_pandas(python) Return the pandas dataframe of the data content to_numpy(python) Return the numpy array of the data content rename Rename the specified folder (name) move Move the file (folder) to the specified path delete Delete the specified file (folder)

[0077] Although specific embodiments of the present invention are disclosed for illustrative purposes, which are intended to assist in understanding the content of the present invention and implementing it accordingly, those skilled in the art can understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the best embodiments, and the scope of protection claimed by the present invention shall be subject to the scope defined by the claims.

Claims

1. A cross-frame data transmission method based on unified representation, characterized in that The method includes: Obtaining a dataset of scientific data and a list of URLs of the dataset in a public database through a unique dataset ID, and constructing a tree-shaped data view in the form of a dictionary based on the list of URLs of the dataset; wherein, the dataset is in the form of a folder or a compressed package in the public database; Based on a consumer's data view request, converting the tree-shaped data view into a DataFrame instance and sending the DataFrame instance to the consumer, so that when the consumer performs dataset operations on the DataFrame instance, the DataFrame instance will return a new DataFrame instance; wherein, the new DataFrame instance contains data content or a data view after operations.

2. The method according to claim 1, wherein The DataFrame instance is a two-dimensional table with folders and / or files as rows and dictionary elements as columns; wherein, the dictionary elements include: file name, file type, and depth.

3. The method according to claim 2, characterized in that, The dataset operations include: file-level dataset operations and folder-level dataset operations; wherein, the folder-level dataset operations include: triggered loading operations and non-triggered loading operations.

4. The method according to claim 3, characterized in that, The triggered loading operation includes: an operation of opening one or more unloaded files; When the consumer performs an operation of opening one or more unloaded files on the DataFrame instance, the DataFrame instance will return a new DataFrame instance, including: The DataFrame instance finds the list of URLs of the unloaded file according to the filtered file name and the mapping between the file name and the URL; Initiating an HTTP request based on the list of URLs of the unloaded file to download the unloaded file and caching it in memory; wherein, the file cached in memory is a binary format file; Constructing and returning a new DataFrame instance with files as rows and file content as columns according to the content of the binary format file.

5. The method according to claim 3, characterized in that, The triggered loading operation includes: an operation of opening a file in a format supported by the DataFrame instance that has been loaded in the data content; When the consumer performs an operation of opening a file in a format supported by the DataFrame instance that has been loaded in the data content on the DataFrame instance, the DataFrame instance will return a new DataFrame instance, including: The DataFrame instance finds the list of URLs of the file to be opened according to the filtered file name and the mapping between the file name and the URL; Initiating an HTTP request based on the list of URLs of the file to be opened to download the file to be opened and caching it in memory; wherein, the file cached in memory is a binary format file; Constructing and returning a new DataFrame instance with rows and columns equivalent to the rows and columns of the file content according to the content of the binary format file.

6. The method according to claim 3, characterized in that, The non-triggered loading operation includes: an operation of opening a folder, an operation of filtering files, and an operation of merging datasets; When the consumer performs dataset operations on this DataFrame instance, this DataFrame instance will return a new DataFrame instance, including: When the consumer performs an operation to open a folder on this DataFrame instance, this DataFrame instance searches for sub-files and sub-folders under the folder according to the folder name in the data view, and constructs and returns a new DataFrame instance with the sub-files and sub-folders as the data view; When the consumer performs a file filtering operation on this DataFrame instance, this DataFrame instance searches for all files and folders that meet the requirements according to the regular expression input by the consumer in the data view, and constructs and returns a new DataFrame instance with the matching results as the data view; When the consumer performs a dataset merging operation on two DataFrame instances, the first DataFrame instance to be merged constructs and returns a new DataFrame instance with the merge of the data views of the two DataFrame instances as the data view.

7. The method according to claim 1, wherein When performing cross-frame data transfer between Python and R, the method further includes: defining a unified interoperability interface between Python and R; wherein, the interoperability interface is based on the rpc protocol and the ipc protocol, and the interoperability interface includes: The constructor is used to initialize the DataFrame through configuration information; get_schema is used to obtain the data view of the dataset; open is used to open the folder or file of the dataset; flat_open is used to open the specified or all files in the data view, and expand and flatten the nested interfaces; filter is used to perform regular matching and filtering on folders and files; concat is used for row concatenation of the DataFrame.

8. A cross-frame data transmission system based on unified representation, characterized in that, The system includes: A data view construction module, which is used to obtain the dataset of scientific data and the URL list of this dataset in the public database through the unique ID of the dataset, and construct a tree-shaped data view in dictionary form according to the URL list of the dataset; wherein, the dataset is in the form of a folder or a compressed package in the public database; A data transfer module, which is used to convert the tree-shaped data view into a DataFrame instance based on the data view request of the consumer, and send the DataFrame instance to the consumer, so that when the consumer performs dataset operations on this DataFrame instance, this DataFrame instance will return a new DataFrame instance; wherein, the new DataFrame instance contains data content or the data view after operations.

9. An electronic device, characterized in that, The electronic device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, it implements the cross-frame data transfer method based on unified representation according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, Computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by a processor, the cross-frame data transmission method based on unified representation described in any one of claims 1-7 is implemented.