File processing method and apparatus based on time series database

CN122817320APending Publication Date: 2026-09-25TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611029822.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-10
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]本发明提供一种基于时序数据库的文件处理方法及装置,用以解决如何提供一种方案,既能够保持标准文件系统接口的兼容,又能够利用时序数据库统一管理能力,实现文件的高效处理的问题

Benefits of technology

[0019]本发明还提供一种非暂态计算机可读存储介质,其上存储有计算机程序,该计算机程序被处理器执行时实现如上述任一种所述基于时序数据库的文件处理方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817320A_ABST
    Figure CN122817320A_ABST
Patent Text Reader

Abstract

The application provides a file processing method and device based on a time series database, and relates to the technical field of computers.The method comprises the following steps: establishing a storage model in the time series database, wherein the storage model is used for storing files corresponding to different time points of different logical paths; receiving a calling request of a target directory corresponding file from an upper application, wherein the calling request comprises a target logical path, a calling operation and an access path; and processing the files corresponding to different time points of the target logical path in the time series database based on the target operation and the access path in the time series database converted from the calling operation. The storage model established in the time series database is used for storing the files corresponding to different time points of different logical paths, ensuring the compatibility of the standard file system interface, and the files stored in the time series database are used to process the files corresponding to different time points of the target logical path in the time series database, improving the processing efficiency of the files.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to a file processing method and apparatus based on a time-series database. Background Technology

[0002] Traditional file systems typically manage directory entries, inodes, and data blocks directly on local disks or distributed storage, providing mature Portable Operating System Interface (POSIX) or POSIX-like file access interfaces to upper-layer applications. Database systems, on the other hand, excel at unified persistence, version indexing, auditing, querying, disaster recovery, and centralized operation and maintenance.

[0003] In scenarios requiring both standard file access and unified database management capabilities, the file system and database system are often disconnected: if upper-layer applications directly access the database, they need to be modified into database clients; if they continue to access the file system, it is difficult to utilize the database's indexing, querying, auditing, and time-based capabilities. The separation between standard file interfaces and database storage interfaces makes it difficult to use unified database management capabilities without altering the upper-layer applications' file access habits.

[0004] Therefore, how to provide a solution that can maintain compatibility with standard file system interfaces while leveraging the unified management capabilities of time-series databases to achieve efficient file processing is an urgent problem to be solved. Summary of the Invention

[0005] This invention provides a file processing method and apparatus based on a time-series database, which addresses the problem of how to provide a solution that can maintain compatibility with standard file system interfaces while leveraging the unified management capabilities of a time-series database to achieve efficient file processing.

[0006] This invention provides a file processing method based on a time-series database, comprising: A storage model is established in the time-series database, which is used to store files at different time points corresponding to different logical paths. Receive a call request from an upper-layer application for a file corresponding to a target directory. The call request includes the target logical path, the call operation, and the access path. Transform the invocation operation into the target operation in the time series database; Based on the target operation and the access path, the files at different time points corresponding to the target logical path in the time series database are processed.

[0007] According to the present invention, a file processing method based on a time-series database is provided, wherein the processing of files at different time points corresponding to the target logical path in the time-series database based on the target operation and the access path includes any one of the following: When the access path is the first logical path, based on the target operation, the file at the latest time point among different time points corresponding to the target logical path in the time series database is processed; When the access path is a second logical path and the second logical path carries a timestamp, based on the target operation, the files corresponding to the timestamps at different time points corresponding to the target logical path in the time series database are processed.

[0008] According to a file processing method based on a time-series database provided by the present invention, the step of processing the file at the latest time point among different time points corresponding to the target logical path in the time-series database based on the target operation includes at least one of the following: If the target operation includes writing a file, the file to be written is written to the file at the latest time point among the different time points corresponding to the target logical path in the time series database; If the target operation includes reading a file, the target byte range of the file at the latest time point among the different time points corresponding to the target logical path in the time series database is read; If the target operation includes closing a file, close the file being written to in the time series database, and update at least one of the following: the write file block corresponding to the file to be written, the write end marker, the size of the file being written to in the time series database, the modification time, and the access status.

[0009] According to a file processing method based on a time-series database provided by the present invention, the step of writing the file to be written to the file at the latest time point among different time points corresponding to the target logical path in the time-series database includes: Based on at least one of the file offset, byte content, object identifier, end marker, and file size of the file to be written included in the write request, the file to be written is written in the form of file blocks to the binary storage area pointed to by the object content field or object pointer of the latest time point of the file corresponding to the target logical path in the time series database.

[0010] According to a file processing method based on a time-series database provided by the present invention, the step of reading the target byte range of the file at the latest time point among different time points corresponding to the target logical path in the time-series database includes: Based on the file offset and read length included in the read request, the target byte range in the file of the latest time point among different time points corresponding to the target logical path in the time series database is read; the target byte range is determined based on the file offset and the read length.

[0011] According to the file processing method based on time-series database provided by the present invention, for files at different time points corresponding to any logical path, the files at different time points are recorded as a version sequence ordered by time.

[0012] According to a file processing method based on a time-series database provided by the present invention, the method further includes: When a user opens a file in read-write mode, a new time-version file is generated for the file based on the file's logical path; Perform the operation corresponding to the read / write mode in the new time version file.

[0013] According to a file processing method based on a time-series database provided by the present invention, the new time version file is displayed with a normal filename, and the file is displayed with a hidden filename carrying a timestamp.

[0014] According to a file processing method based on a time-series database provided by the present invention, the method further includes: When reading the file corresponding to the hidden filename, the logical path and the timestamp are parsed from the hidden filename; Based on the logical path and the timestamp, the file corresponding to the hidden filename is determined.

[0015] According to a file processing method based on a time-series database provided by the present invention, the method further includes: When reading the file corresponding to the ordinary filename, select the latest time version file and use the latest time version file as the file corresponding to the ordinary filename.

[0016] According to a file processing method based on a time-series database provided by the present invention, the storage model includes at least one of the following: a time primary key, a logical path, a parent directory, a file name, a directory identifier, a file size, a permission bit, a user identifier, a group identifier, a number of links, a creation time, a modification time, an access time, and an object content field; the logical path represents a file path visible to the user, the parent directory is used for directory traversal, the time primary key is used to distinguish different file versions under the same logical path, and the object content field is used to store the original bytes of the file or an object pointer pointing to the original bytes of the file.

[0017] The present invention also provides a file processing apparatus based on a time-series database, comprising: A module is established to create a storage model in the time-series database. The storage model is used to store files at different time points corresponding to different logical paths. The receiving module is used to receive a call request from the upper-layer application for a file corresponding to the target directory. The call request includes the target logical path, the call operation, and the access path. The conversion module is used to convert the call operation into the target operation in the time series database; The processing module is used to process files at different time points corresponding to the target logical path in the time-series database based on the target operation and the access path.

[0018] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the file processing method based on a time-series database as described above.

[0019] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the file processing method based on a time-series database as described above.

[0020] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the file processing method based on a time-series database as described above.

[0021] The present invention provides a file processing method and apparatus based on a time-series database. This method establishes a storage model in the time-series database to store files at different times corresponding to different logical paths. It receives a call request from an upper-layer application for a file corresponding to a target directory. The call request includes a target logical path, a call operation, and an access path. The call operation is converted into a target operation in the time-series database. Based on the target operation and the access path, the files at different times corresponding to the target logical path in the time-series database are processed. By storing files at different times corresponding to different logical paths through the storage model established in the time-series database, compatibility with standard file system interfaces is ensured. Simultaneously, by utilizing the files at different times corresponding to different logical paths stored in the time-series database, and based on the target operation converted from the call operation and the access path, the processing of files at different times corresponding to the target logical path in the time-series database is achieved, improving file processing efficiency. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0023] Figure 1 This is one of the flowcharts of the file processing method based on a time-series database provided by the present invention.

[0024] Figure 2 This is a schematic diagram of the system architecture of the file processing method based on a time-series database provided by the present invention.

[0025] Figure 3 This is a schematic diagram illustrating the file version sequence at different time points provided by the present invention, opening files in read-write mode, and reading files corresponding to hidden filenames and ordinary filenames.

[0026] Figure 4 This is the second flowchart of the file processing method based on a time-series database provided by the present invention.

[0027] Figure 5 This is the third flowchart of the file processing method based on a time-series database provided by the present invention.

[0028] Figure 6 This is a schematic diagram illustrating the overwrite versioning and historical version hiding of files provided by the present invention.

[0029] Figure 7 This is a schematic diagram of the structure of the file processing device based on a time-series database provided by the present invention.

[0030] Figure 8 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0032] To facilitate a clear understanding of the various embodiments of this application, relevant background knowledge will be introduced first.

[0033] As the scale of scientific, engineering, and business archives continues to grow, files are no longer just static objects that can be written and read in their entirety. Many applications need to preserve historical states as files are constantly updated, supporting write recovery, audit trails, and comparisons between different versions. At the same time, reading large files often only requires the target byte range or a partial object within the file. If the entire file is still transferred or the entire object is overwritten, it will result in significant network, disk, and memory overhead.

[0034] Therefore, the existing technologies generally have the following shortcomings: (1) Directory metadata, object content and version information are usually maintained in a scattered manner, and additional consistency processing is required between file attributes, object content and historical status; (2) Large file writing often adopts whole file writing or whole object overwriting, which cannot make full use of the offset reading and writing capabilities of object fields, and the cost of failure recovery and partial update is high; (3) When implementing the file system based on the database, high-frequency operations such as getting attributes (getattr), reading directories (readdir), opening (open), reading (read) are prone to generating a large number of connection establishment, metadata query and object reading overhead; (4) Overwrite usually destroys old content. If historical backtracking or auditing is required, a backup, snapshot or archiving mechanism needs to be established separately, the system implementation is complex and the access entry is not unified.

[0035] Therefore, it is urgent to solve the following problems: how to maintain compatibility with standard file system interfaces, utilize the time-series database's time primary key and query capabilities to uniformly carry file system metadata, object content, and historical versions; how to support offset-based streaming writing and reading of object content; and how to reduce the access overhead of the database's backend file system through version navigation, context management, and multi-level caching.

[0036] The following is combined Figures 1 to 6 This invention describes a file processing method based on a time-series database.

[0037] Figure 1 This is one of the flowcharts illustrating the file processing method based on a time-series database provided by the present invention, such as... Figure 1 As shown, the method includes steps 101-104.

[0038] Step 101: Establish a storage model in the time series database. The storage model is used to store files at different time points corresponding to different logical paths.

[0039] Specifically, the storage model includes at least one of the following: a time primary key, a logical path, a parent directory, a filename, a directory identifier, a file size, permission bits, a user identifier, a group identifier, a number of links, a creation time, a modification time, an access time, and an object content field, where the object is a file. The logical path represents the user-visible file path, the parent directory is used for directory traversal, the time primary key is used to distinguish different file versions under the same logical path, and the object content field is used to store the original file bytes or an object pointer pointing to the original file bytes.

[0040] In one implementation, the storage model can consist of a metadata directory table and an object content table. The metadata directory table includes a time primary key, logical path, parent directory, filename, directory identifier, file size, permission bits, user identifier, group identifier, number of links, creation time, modification time, and access time. The object content table includes object content fields. In another implementation, the metadata fields and object content fields can be stored together in the same table. The metadata fields include a time primary key, logical path, parent directory, filename, directory identifier, file size, permission bits, user identifier, group identifier, number of links, creation time, modification time, and access time.

[0041] A storage model is established in the time-series database to support the file system. This storage model stores files at different points in time corresponding to different logical paths.

[0042] Step 102: Receive a request from the upper-layer application to call the file corresponding to the target directory. The request includes the target logical path, the call operation, and the access path.

[0043] Specifically, the target directory can be a mounted directory containing files. The invocation operations can include at least one of the following: attribute reading, directory traversal, directory creation, file creation, file opening, file reading, file writing, file closing, truncation, renaming, file deletion, and directory deletion. The access path includes either a first logical path or a second logical path; the first logical path is a regular logical path, and the second logical path is a hidden version logical path.

[0044] The system receives requests from upper-layer applications for accessing files in the target directory via a Filesystem in Userspace (FUSE) or an equivalent userspace file system. The request includes the target logical path, the call operation, and the access path.

[0045] Step 103: Transform the call operation into the target operation in the time series database.

[0046] Specifically, the target operation can include at least one of the following: query, insert, update, delete, file read, file write, and file close operations. The invoked operation can be transformed into the target operation in the time-series database.

[0047] Therefore, the upper-level command-line tools, scripts, and business programs still operate through ordinary file interfaces, while the lower-level database handles persistence and version management.

[0048] Step 104: Based on the target operation and the access path, process the files at different time points corresponding to the target logical path in the time series database.

[0049] Specifically, based on the target operation and access path, files at different time points corresponding to the target logical path in the time series database can be processed.

[0050] This invention provides a file processing method based on a time-series database. The method establishes a storage model within the time-series database to store files at different times corresponding to different logical paths. It receives a request from an upper-layer application to access files in a target directory; this request includes the target logical path, a call operation, and an access path. The call operation is then converted into a target operation within the time-series database. Based on the target operation and the access path, the method processes files at different times corresponding to the target logical path in the time-series database. By storing files at different times corresponding to different logical paths using the storage model established in the time-series database, compatibility with standard file system interfaces is ensured. Furthermore, by utilizing the files at different times corresponding to different logical paths stored in the time-series database, and based on the target operation converted from the call operation and the access path, the method achieves the processing of files at different times corresponding to the target logical path in the time-series database, thereby improving file processing efficiency.

[0051] In some embodiments, step 104 above is specifically implemented in any of the following ways: When the access path is a first logical path, the file with the latest time point among different time points corresponding to the target logical path in the time series database is processed based on the target operation; when the access path is a second logical path and the second logical path carries a timestamp, the file corresponding to the timestamp among different time points corresponding to the target logical path in the time series database is processed based on the target operation.

[0052] Specifically, when the access path is the first logical path, such as accessing via a normal logical path, based on the target operation, the file at the latest time point among different time points corresponding to the target logical path in the time series database is processed.

[0053] Alternatively, if the access path is a second logical path and the second logical path carries a timestamp, such as accessing with a hidden version logical path, the files corresponding to the timestamps at different time points corresponding to the target logical path in the time series database are processed based on the target operation.

[0054] In some embodiments, processing the file at the latest time point among different time points corresponding to the target logical path in the time series database based on the target operation includes at least one of the following: If the target operation includes writing a file, the file to be written is written to the file at the latest time point among the different time points corresponding to the target logical path in the time series database; if the target operation includes reading a file, the target byte range is read from the file at the latest time point among the different time points corresponding to the target logical path in the time series database; if the target operation includes closing a file, the file to be written in the time series database is closed, and at least one of the following is updated: the write file block corresponding to the file to be written, the write end marker, the size of the file to be written in the time series database, the modification time, and the access status.

[0055] Specifically, when the target operation includes writing a file, the file to be written can be written to the file at the latest time point among different time points corresponding to the target logical path in the time series database; when the target operation includes reading a file, the target byte range in the file at the latest time point among different time points corresponding to the target logical path in the time series database can be read; when the target operation includes closing a file, the file to be written in the time series database can be closed, and at least one of the following can be updated: the write file block corresponding to the file to be written, the write end marker, the size of the file to be written in the time series database, the modification time, and the access status.

[0056] In some embodiments, writing the file to be written to the file at the latest time point among the different time points corresponding to the target logical path in the time-series database includes: Based on at least one of the file offset, byte content, object identifier, end marker, and file size of the file to be written included in the write request, the file to be written is written in the form of file blocks to the binary storage area pointed to by the object content field or object pointer of the latest time point of the file corresponding to the target logical path in the time series database.

[0057] Specifically, when the target operation includes writing to a file, a write request can be constructed based on at least one of the following: file offset, byte content, object identifier, end marker, and file size of the file to be written.

[0058] Based on at least one of the following included in the write request: file offset, byte content, object identifier, end marker, and file size of the file to be written, the file to be written is written in the form of file blocks to the binary storage area pointed to by the object content field or object pointer of the file at the latest time point of the target logical path in the time series database.

[0059] In some embodiments, reading the target byte range from the file at the latest time point among different time points corresponding to the target logical path in the time-series database includes: Based on the file offset and read length included in the read request, the target byte range in the file of the latest time point among different time points corresponding to the target logical path in the time series database is read; the target byte range is determined based on the file offset and the read length.

[0060] Specifically, when the target operation includes reading a file, a read request can be constructed based on the file offset and read length of the file to be read. The file to be read is the latest file among different time points corresponding to the target logical path in the time series database.

[0061] Based on the file offset and read length included in the read request, the target byte range is determined. Then, the read interface is called to read the target byte range from the file at the latest time point in the time series database corresponding to the target logical path, i.e., only the target byte range is returned. For example, the read interface is READ_OBJECT(data,off,len) or an equivalent partial read interface.

[0062] In some embodiments, based on the target operation, the files corresponding to the timestamps at different time points corresponding to the target logical path in the time series database are processed, including: When the target operation includes writing a file, based on at least one of the file offset, byte content, object identifier, end marker and file size of the file to be written included in the write request, the file to be written is written in the form of file blocks to the binary storage area pointed to by the object content field or object pointer of the file corresponding to the timestamp at different time points of the target logical path in the time series database. If the target operation includes reading a file, the target byte range in the file corresponding to the timestamp at different time points in the time series database is read based on the file offset and read length included in the read request; the target byte range is determined based on the file offset and the read length. If the target operation includes closing a file, close the file being written to in the time series database, and update at least one of the following: the write file block corresponding to the file to be written, the write end marker, the size of the file being written to in the time series database, the modification time, and the access status.

[0063] In some embodiments, for files at different times corresponding to any logical path, the files at different times are recorded as a version sequence ordered by time. When accessing a normal logical path, the system resolves it to the latest time version of that logical path; when accessing a hidden version path with timestamps, the system resolves it to the version at the specified time. In this way, the same file path can be represented as a normal file, or, when needed, as a navigable collection of historical versions.

[0064] In some embodiments, the method further includes: When a user opens a file in read-write mode, a new time-version file is generated based on the file's logical path; the operation corresponding to the read-write mode is then performed in the new time-version file.

[0065] Specifically, when a user opens a file in read-write mode, the current version of the file is not directly destroyed. Instead, a new time-version file is generated based on the file's logical path. The corresponding read-write operation is then performed in the new time-version file.

[0066] In some embodiments, the new time-version file (i.e., the latest version file) is displayed with a normal filename, and the file (i.e., the historical version file) is displayed with a hidden filename carrying a timestamp.

[0067] In some embodiments, the method further includes: When reading the file corresponding to the hidden filename, the logical path and the timestamp are parsed from the hidden filename; based on the logical path and the timestamp, the file corresponding to the hidden filename is determined.

[0068] Specifically, when reading a file corresponding to a hidden filename, the logical path and timestamp are parsed from the hidden filename; based on the logical path and timestamp, the file corresponding to the hidden filename can be determined. That is, when reading a hidden file, the logical path and timestamp are parsed from the hidden filename to locate the specified historical version. This mechanism enables overwrites to have historical readability, audit trails, and write error recovery capabilities.

[0069] In some embodiments, the method further includes: When reading the file corresponding to the ordinary filename, select the latest time version file and use the latest time version file as the file corresponding to the ordinary filename.

[0070] Specifically, when reading a file corresponding to a regular filename, the latest time version file can be selected and used as the file corresponding to the regular filename.

[0071] In some embodiments, a session pool for the time-series database can also be established, and session borrowing and returning can be managed through session leases. Session borrowing and returning refer to each file processing operation, such as file reading and file writing.

[0072] In some embodiments, each open file corresponds to an open file context, which centrally stores the file path, version time, current length, blocks to be read / written, read cache, write status, and closing / closing information.

[0073] In some embodiments, inode write-back cache, directory sub-item cache, raw file cache, and open context read block cache can also be maintained to merge attribute updates, accelerate directory traversal, hit small files or repeated reads of entire files, and optimize sequential reads and partial repeated reads. Each cache uses logical path and time version as key dimensions to avoid confusion between different versions of content.

[0074] In some embodiments, when a time-series database is deployed in a cluster, the partitioning, replication, consensus protocol, pre-write log, fault detection, and replication reconstruction mechanisms of the time-series database are utilized to ensure reliable storage of metadata and object content.

[0075] In some embodiments, for object content fields, methods such as direct disk write, object pointer, and consensus copying of object fields can be used to reduce the WAL amplification and repeated read / write overhead of large files.

[0076] Figure 2 This is a schematic diagram of the system architecture of the file processing method based on a time-series database provided by the present invention, as shown below. Figure 2 As shown, the system includes client applications, FUSE mount points, a file system core process, and a time-series database cluster. The client applications include command-line tools, scripts, a graphical file manager, and general business applications. The file system core process includes a time-series database connection module, a version management module, an object read / write module, an open file context management module, a metadata caching module, and an object content caching module. The system uses Apache IoTDB as the time-series database.

[0077] Upon system startup, the system reads configuration files, initializes logs and session pools, and automatically creates a time-series database and related tables, such as a metadata directory table and an object content table. If the root directory does not exist, a root directory record is created. The metadata directory table, which can be named `inode`, includes at least the following fields: time key (`time`), logical path (`path`), parent directory (`parent`), filename (`name`), directory identifier (`is_dir`), file size (`size`), permission bits (`mode`), user identifier (`uid`), group identifier (`gid`), number of links (`nlink`), creation time (`ctime`), modification time (`mtime`), and access time (`atime`). The object content table, which can be named `blob`, includes at least the following fields: time key (`time`), logical path (`path`), metadata (`data`), file size (`size`), and object content. In another implementation, the object content fields can be directly set in the `inode` table. The time key field serves as the creation time of the file version, and the logical path field serves as the user-visible logical path; both jointly identify a unique version.

[0078] After the client performs standard file system operations on the mount point, FUSE forwards the file system request to the file system kernel process, which then converts the request into access to the time-series database table, object fields (i.e., object content fields), and version records.

[0079] Figure 3 This is a schematic diagram illustrating the file version sequence at different time points, opening files in read-write mode, and reading files corresponding to hidden filenames and ordinary filenames provided by this invention, such as... Figure 3 As shown, for any logical path p, the files V(t) at different time points t corresponding to any logical path p are recorded as a time-ordered version sequence F(p) = {V(t1), V(t2), V(t3), ..., {V(t...}}. n Each time version file includes the version time t, file system metadata, binary content (object content), and may also include other content such as logical paths.

[0080] When a user opens a file in read-write mode, for example, writing to the file corresponding to logical path p, does not directly destroy the current version of the file. Instead, based on the file's logical path, a new timestamp file V(t) is generated for that file. n+1The operation corresponding to read / write mode is performed in the new time-version file. The new time-version file (i.e., the latest version file) is displayed with a normal filename, while the file (i.e., the historical version file) is displayed with a hidden filename carrying a timestamp. Thus, file updates are expressed as appending a new version, rather than a destructive overwrite of an old version.

[0081] When reading a file corresponding to a hidden filename, i.e., accessing it via a hidden version path, such as / .p_time.format, the logical path and timestamp are parsed from the hidden filename. Based on the logical path and timestamp, the file corresponding to the hidden filename can be determined. In other words, when reading a hidden file, the logical path and timestamp are parsed from the hidden filename to locate the specified historical version. This mechanism enables overwrites to have historical readability, audit trails, and write error recovery capabilities.

[0082] When reading files corresponding to ordinary filenames, i.e., accessing them using ordinary version paths, such as / p, logical paths are parsed from hidden filenames; the latest time version file corresponding to the logical path can be identified as the file corresponding to the hidden filename.

[0083] Figure 4 This is the second flowchart of the file processing method based on a time-series database provided by the present invention, as shown below. Figure 4 As shown, requests for file system calls to the target directory from upper-layer applications are received via FUSE or an equivalent user-space file system. Operation types are distributed based on the upper-layer application's calls to the corresponding files in the target directory. For example, the calls may include attribute reading (getattr), directory traversal (readdir), directory creation (mkdir), file creation (creat), file opening (open), file reading (read), file writing (write), file closing (release), renaming (rename), file deletion (unlink), and directory deletion (rmdir). These calls are then translated into target operations in the time-series database.

[0084] When the system receives a `getattr` call, it determines the latest or specified version based on path resolution rules and returns the file type, size, permissions, and time attributes. When the system receives a `readdir` call, it queries the child items based on the parent directory field and organizes the latest version into the directory list with ordinary filenames, while organizing historical versions into hidden filenames according to configuration. When the system receives a `mkdir` or `create` call, it inserts directory or file records into the metadata table. When the system receives an `open` call, it creates an open file context based on the open flag and determines whether the current version or the new version is used for this access. When the system receives a `read` call, it reads the object content based on the version time, path, offset, and length in the context. When the system receives a `write` call, it temporarily stores the byte block in the context or writes it to the object field. When the system receives a `release` call, it refreshes the block to be written, writes the end marker, and updates the final size and time attributes. When the system receives a `rename` call, it updates the path and parent directory of the target record; if the target is a directory, it recursively updates the subpaths. When the system receives an `unlink` or `rmdir` call, it deletes visible records according to file or directory rules or retains historical versions according to a policy. Based on the above, the mapping from standard file system operations to multi-version object storage semantics is completed.

[0085] Figure 5 This is the third flowchart of the file processing method based on a time-series database provided by the present invention, as shown below. Figure 5 As shown, the system retrieves object content read / write requests and distributes them according to request type. Request types include write request / close process or read process. Upon receiving a write request, the system first reads the current file length, current version time, and pending write status from the open file context. It then determines whether the write is sequential. For sequential append writes, the system adds the object blocks to be written to the batch write queue and writes them uniformly to the object fields when the size threshold, time threshold, or file closure is reached. For non-sequential writes, the system writes to the corresponding position according to the request offset and updates the file length and modification time in the context. Each object write request includes the target path, version time, file offset, byte content, byte length, end marker, and file size. When the file is closed, the system writes the end marker, refreshes all pending write blocks, and writes the final file size back to the metadata record.

[0086] Upon receiving a read request, the system first checks if the open context read block cache hits the target offset range. If it misses, it checks if the original file cache already contains the complete content or a reusable fragment of that version of the file. If it still misses, it calls READ_OBJECT(data,off,len) or an equivalent partial read interface to read bytes of the specified offset and length from the object content field. After the read is complete, the system writes the returned content to the read block cache, and can merge the access time update into the inode and write it back to the cache with delayed flushing, before returning the read content.

[0087] The above steps complete the offset writing, reading, and closing of object content, avoiding the need to perform full file transfer for each read and reducing the database access overhead of sequential and repeated reads.

[0088] Figure 6 This is a schematic diagram illustrating the overwrite versioning and historical version hiding of files provided by the present invention, such as... Figure 6 As shown, different time points under the same logical path correspond to different file versions. For example, files V(t1)-V(t4) at time points t1-t4, where V(t3) is the original latest version file and V(t4) is the new version file. Files at different time points are recorded as a version sequence ordered by time. Each time version file includes version time, logical path, file system metadata, and binary object content.

[0089] When a user opens a regular file in read-only mode (e.g., `open(sample.dat, O_RDONLY)`), the system selects the original, latest time-version file (V(t3)). When a user opens an existing file in write mode (e.g., `open(sample.dat, O_WRONLY / O_RDWR)`), a new time-version file is generated based on the file's logical path. This involves copying the necessary metadata and generating a new time record as the new version (V(t4)). Subsequent writes are all written to the new version object. During directory traversal, for example, if the directory list includes `sample.dat`, `.sample_1713600000000.dat`, and `.sample_1713700000000.dat`, the latest time-version file is displayed with a regular filename, such as `sample.dat`; historical versions are displayed with hidden filenames carrying timestamps, such as `.sample_1713600000000.dat`. When a user reads a hidden filename (i.e., `read(.sample_1713700000000.dat)`), the system parses the timestamp and logical path from the hidden filename, uses the logical path to reconstruct the original logical path `sample.dat`, and uses the logical path and the timestamp to locate the specified historical version (V(t2)). Example of a hidden filename: `sample_1713600000000.dat`.

[0090] In some embodiments, the system may also maintain the following four types of caches: First, an inode write-back cache is used to merge updates of attributes such as file size, access time, modification time, and permissions, and is flushed to the database periodically or uniformly after a closing event; Second, a directory sub-item cache is used to cache query results from a parent directory to a list of sub-items, and is invalidated or incrementally updated after operations such as creation, deletion, and renaming; Third, a raw file cache is used to cache the complete content of small files or hot files, using logical path and version time as cache keys; Fourth, an open context read block cache is used to cache recently read blocks within a specified open file to optimize sequential reads and partial repeated reads. All of the above caches include a version dimension to avoid content confusion between historical versions and the latest version.

[0091] In some embodiments, the system is deployed using a time-series database cluster. Metadata and object content are managed through database partitioning, replication, and consensus protocols. When object content fields are stored as large objects or object pointers, the system can directly write the binary content to the local file system of the data nodes, and the database records the object pointers, time versions, and consistency status. For large file writes, the system can reduce the duplication of the pre-write log for the complete binary content, performing log protection only on necessary metadata, object pointers, and consistency information, thereby reducing disk I / O amplification.

[0092] In the above embodiments, this invention uniformly carries multi-version file system metadata and object content through a metadata directory table and an object content table in a time-series database. It maintains POSIX file interface compatibility through FUSE, improves large file processing efficiency through object content offset read / write, enhances write security and historical traceability through an overwrite versioning mechanism, and reduces the access overhead of the database backend file system through session pools, open file contexts, and multi-level caching. Compared with existing ordinary file systems, simple object storage, or simple database backup mechanisms, this invention can simultaneously provide current file access, historical version navigation, and object-level local read / write capabilities within the same file system namespace.

[0093] The file processing apparatus based on a time-series database provided by the present invention is described below. The file processing apparatus based on a time-series database described below and the file processing method based on a time-series database described above can be referred to in correspondence with each other.

[0094] Figure 7 This is a schematic diagram of the structure of the file processing device based on a time-series database provided by the present invention, as shown below. Figure 7 As shown, the file processing device 700 based on a time-series database includes: an establishment module 701, a receiving module 702, a conversion module 703, and a processing module 704; wherein, Module 701 is used to establish a storage model in the time series database. The storage model is used to store files at different time points corresponding to different logical paths. The receiving module 702 is used to receive a call request from an upper-layer application for a file corresponding to a target directory. The call request includes the target logical path, the call operation, and the access path. The conversion module 703 is used to convert the call operation into the target operation in the time series database; The processing module 704 is used to process files at different time points corresponding to the target logical path in the time series database based on the target operation and the access path.

[0095] The file processing device based on a time-series database provided by this invention establishes a storage model in the time-series database to store files at different times corresponding to different logical paths; receives a call request from an upper-layer application for a file corresponding to a target directory, the call request including the target logical path, the call operation, and the access path; converts the call operation into a target operation in the time-series database; and processes the files at different times corresponding to the target logical path in the time-series database based on the target operation and the access path. By storing files at different times corresponding to different logical paths through the storage model established in the time-series database, compatibility with standard file system interfaces is ensured. Simultaneously, by utilizing the files at different times corresponding to different logical paths stored in the time-series database, and based on the target operation and access path converted from the call operation, the device processes files at different times corresponding to the target logical path in the time-series database, thereby improving file processing efficiency.

[0096] In some embodiments, the processing module 704 is specifically used for any of the following: When the access path is the first logical path, based on the target operation, the file at the latest time point among different time points corresponding to the target logical path in the time series database is processed; When the access path is a second logical path and the second logical path carries a timestamp, based on the target operation, the files corresponding to the timestamps at different time points corresponding to the target logical path in the time series database are processed.

[0097] In some embodiments, the processing module 704 is further configured to perform at least one of the following: If the target operation includes writing a file, the file to be written is written to the file at the latest time point among the different time points corresponding to the target logical path in the time series database; If the target operation includes reading a file, the target byte range of the file at the latest time point among the different time points corresponding to the target logical path in the time series database is read; If the target operation includes closing a file, close the file being written to in the time series database, and update at least one of the following: the write file block corresponding to the file to be written, the write end marker, the size of the file being written to in the time series database, the modification time, and the access status.

[0098] In some embodiments, the processing module 704 is further configured to: Based on at least one of the file offset, byte content, object identifier, end marker, and file size of the file to be written included in the write request, the file to be written is written in the form of file blocks to the binary storage area pointed to by the object content field or object pointer of the latest time point of the file corresponding to the target logical path in the time series database.

[0099] In some embodiments, the processing module 704 is further configured to: Based on the file offset and read length included in the read request, the target byte range in the file of the latest time point among different time points corresponding to the target logical path in the time series database is read; the target byte range is determined based on the file offset and the read length.

[0100] In some embodiments, for files at different times corresponding to any logical path, the files at different times are recorded as a version sequence ordered by time.

[0101] In some embodiments, the time-series database-based file processing apparatus 700 further includes: The generation module is used to generate a new time version file for a file based on the logical path of the file when the user opens the file in read-write mode. The execution module is used to perform the operations corresponding to the read / write mode in the new time version file.

[0102] In some embodiments, the new time-version file is displayed with a normal filename, while the file is displayed with a hidden filename carrying a timestamp.

[0103] In some embodiments, the time-series database-based file processing apparatus 700 further includes: The first parsing module is used to parse the logical path and the timestamp from the hidden file name when reading the file corresponding to the hidden file name; The first determining module is used to determine the file corresponding to the hidden filename based on the logical path and the timestamp.

[0104] In some embodiments, the time-series database-based file processing apparatus 700 further includes: The second parsing module is used to select the latest time version file when reading the file corresponding to the ordinary file name, and to use the latest time version file as the file corresponding to the ordinary file name.

[0105] In some embodiments, the storage model includes at least one of the following: a time primary key, a logical path, a parent directory, a file name, a directory flag, a file size, a permission bit, a user identifier, a group identifier, a number of links, a creation time, a modification time, an access time, and an object content field; the parent directory is used for directory traversal, the time primary key is used to distinguish different file versions under the same logical path, and the object content field is used to store the original bytes of the file or an object pointer pointing to the original bytes of the file.

[0106] Figure 8 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 8 As shown, the electronic device may include a processor 810, a communications interface 820, a memory 830, and a communication bus 840, wherein the processor 810, communications interface 820, and memory 830 communicate with each other via the communication bus 840. The processor 810 can call logical instructions in the memory 830 to execute a file processing method based on a time-series database. This method includes: establishing a storage model in the time-series database, the storage model being used to store files at different times corresponding to different logical paths; receiving a call request from an upper-layer application for a file corresponding to a target directory, the call request including a target logical path, a call operation, and an access path; converting the call operation into a target operation in the time-series database; and processing the files at different times corresponding to the target logical path in the time-series database based on the target operation and the access path.

[0107] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0108] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the file processing method based on a time-series database provided by the above methods. The method includes: establishing a storage model in the time-series database, the storage model being used to store files at different times corresponding to different logical paths; receiving a call request from an upper-layer application for a file corresponding to a target directory, the call request including a target logical path, a call operation, and an access path; converting the call operation into a target operation in the time-series database; and processing the files at different times corresponding to the target logical path in the time-series database based on the target operation and the access path.

[0109] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program is implemented to perform the file processing method based on a time-series database provided by the above methods. The method includes: establishing a storage model in the time-series database, the storage model being used to store files at different times corresponding to different logical paths; receiving a call request from an upper-layer application for a file corresponding to a target directory, the call request including a target logical path, a call operation, and an access path; converting the call operation into a target operation in the time-series database; and processing the files at different times corresponding to the target logical path in the time-series database based on the target operation and the access path.

[0110] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0111] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0112] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A file processing method based on a time-series database, characterized in that, include: A storage model is established in the time-series database, which is used to store files at different time points corresponding to different logical paths. Receive a call request from an upper-layer application for a file corresponding to a target directory. The call request includes the target logical path, the call operation, and the access path. Transform the invocation operation into the target operation in the time series database; Based on the target operation and the access path, the files at different time points corresponding to the target logical path in the time series database are processed.

2. The file processing method based on a time-series database according to claim 1, characterized in that, The step of processing files at different time points corresponding to the target logical path in the time-series database based on the target operation and the access path includes any one of the following: When the access path is the first logical path, based on the target operation, the file at the latest time point among different time points corresponding to the target logical path in the time series database is processed; When the access path is a second logical path and the second logical path carries a timestamp, based on the target operation, the files corresponding to the timestamps at different time points corresponding to the target logical path in the time series database are processed.

3. The file processing method based on a time-series database according to claim 2, characterized in that, The step of processing the file at the latest time point among different time points corresponding to the target logical path in the time series database based on the target operation includes at least one of the following: If the target operation includes writing a file, the file to be written is written to the file at the latest time point among the different time points corresponding to the target logical path in the time series database; If the target operation includes reading a file, the target byte range of the file at the latest time point among the different time points corresponding to the target logical path in the time series database is read; If the target operation includes closing a file, close the file being written to in the time series database, and update at least one of the following: the write file block corresponding to the file to be written, the write end marker, the size of the file being written to in the time series database, the modification time, and the access status.

4. The file processing method based on a time-series database according to claim 3, characterized in that, The step of writing the file to be written to the latest time point among the different time points corresponding to the target logical path in the time-series database includes: Based on at least one of the file offset, byte content, object identifier, end marker, and file size of the file to be written included in the write request, the file to be written is written in the form of file blocks to the binary storage area pointed to by the object content field or object pointer of the latest time point of the file corresponding to the target logical path in the time series database.

5. The file processing method based on a time-series database according to claim 3, characterized in that, The step of reading the target byte range from the file at the latest time point among different time points corresponding to the target logical path in the time-series database includes: Based on the file offset and read length included in the read request, the target byte range in the file of the latest time point among different time points corresponding to the target logical path in the time series database is read; the target byte range is determined based on the file offset and the read length.

6. The file processing method based on a time-series database according to claim 1, characterized in that, For any logical path, files at different times are recorded as a version sequence ordered by time.

7. The file processing method based on a time-series database according to claim 6, characterized in that, The method further includes: When a user opens a file in read-write mode, a new time-version file is generated for the file based on the file's logical path; Perform the operation corresponding to the read / write mode in the new time version file.

8. The file processing method based on a time-series database according to claim 7, characterized in that, The new time-version file is displayed with a normal filename, while the file is displayed with a hidden filename that carries a timestamp.

9. The file processing method based on a time-series database according to claim 8, characterized in that, The method further includes: When reading the file corresponding to the hidden filename, the logical path and the timestamp are parsed from the hidden filename; Based on the logical path and the timestamp, the file corresponding to the hidden filename is determined.

10. The file processing method based on a time-series database according to claim 8, characterized in that, The method further includes: When reading the file corresponding to the ordinary filename, select the latest time version file and use the latest time version file as the file corresponding to the ordinary filename.

11. The file processing method based on a time-series database according to any one of claims 1-10, characterized in that, The storage model includes at least one of the following: a time primary key, a logical path, a parent directory, a filename, a directory identifier, a file size, a permission bit, a user identifier, a group identifier, a number of links, a creation time, a modification time, an access time, and an object content field; the logical path represents the file path visible to the user, the parent directory is used for directory traversal, the time primary key is used to distinguish different file versions under the same logical path, and the object content field is used to store the original bytes of the file or an object pointer pointing to the original bytes of the file.

12. A file processing device based on a time-series database, characterized in that, include: A module is established to create a storage model in the time-series database. The storage model is used to store files at different time points corresponding to different logical paths. The receiving module is used to receive a call request from the upper-layer application for a file corresponding to the target directory. The call request includes the target logical path, the call operation, and the access path. The conversion module is used to convert the call operation into the target operation in the time series database; The processing module is used to process files at different time points corresponding to the target logical path in the time-series database based on the target operation and the access path.