Large file shared storage method, device and equipment based on ES system

Through streaming writing and compression encoding, large files are processed and uploaded to the rolling index of the ES system in batches, the problem of increasing the amount of data stored in large files is solved, efficient storage and transmission is achieved, operation and maintenance costs are reduced, and data reading speed is improved.

CN119917477APending Publication Date: 2025-05-02BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411998427.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

In the data asset management platform system, with the increase in the number of large files, the data becomes larger and larger, and a shared storage solution that can reduce the amount of data stored in large files is needed.

Method used

By responding to the file write instruction, the data of the target large file is written into the local empty text file in a streaming manner, and the local text file is obtained, and then uploaded to the scroll index of the ES system in batches.

Benefits of technology

It realizes the reduction of large file storage data, improves transmission efficiency, reduces the operation and maintenance costs of storage devices, and improves data reading speed through the cache mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119917477A_ABST
    Figure CN119917477A_ABST
Patent Text Reader

Abstract

The invention provides a large file shared storage method, device and equipment based on an ES system, and relates to big data, information flow and data processing in the computer technology, in particular to the field of the large file shared storage method, device and equipment based on the ES system. According to the specific implementation scheme, data in a target large file indicated by a file write-in instruction is written into a local empty text file in a streaming write-in mode, and an initial text file is obtained; and performing compression coding processing on the initial text file to obtain a local text file, and uploading data in the local text file to the rolling index of the ES system in batches. According to the technical scheme, the local text file obtained after the initial text file is compressed and coded is uploaded to the rolling index of the ES system in batches, the purpose of shared storage of reducing the storage data volume of a large file is achieved, and the effects of improving the transmission efficiency and reducing the operation and maintenance cost of storage equipment are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to big data, information flow and data processing in computer technology, and in particular to a shared storage method, device and equipment for large files based on an ES system. Background Art

[0002] The Data Asset Management Platform (DAMP) system includes many large files. The data volume of the large files is large, and as the number of large files increases, the data in the DAMP system becomes larger and larger. Therefore, it is necessary to store the large files in the DAMP system in other systems or devices to share the large files.

[0003] Furthermore, there is an urgent need for a shared storage solution that can reduce the storage data volume of large files. Summary of the invention

[0004] The present disclosure provides a method, apparatus and device for shared storage of large files based on an ES system for reducing the storage data volume of large files.

[0005] According to a first aspect of the present disclosure, a shared storage method for large files based on an ES system is provided, comprising:

[0006] In response to a file write instruction, the data in the target large file indicated by the file write instruction is written into a local empty text file in a streaming write manner to obtain an initial text file; and the initial text file is compressed and encoded to obtain a local text file; wherein the file write instruction is used to instruct the target large file to be stored;

[0007] The data in the local text file is uploaded to the rolling index of the ES system in batches.

[0008] According to a second aspect of the present disclosure, a shared storage device for large files based on an ES system is provided, comprising:

[0009] A writing unit, for responding to a file writing instruction, writing the data in the target large file indicated by the file writing instruction into a local empty text file in a streaming writing manner to obtain an initial text file;

[0010] A processing unit, used for performing compression encoding processing on the initial text file to obtain a local text file; wherein the file writing instruction is used to instruct to store the target large file;

[0011] The storage unit is used to upload the data in the local text file to the rolling index of the ES system in batches.

[0012] According to a third aspect of the present disclosure, there is provided an electronic device, including:

[0013] at least one processor; and

[0014] a memory communicatively connected to the at least one processor; wherein,

[0015] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect.

[0016] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method described in the first aspect.

[0017] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising: a computer program, wherein the computer program is stored in a readable storage medium, at least one processor of an electronic device can read the computer program from the readable storage medium, and the at least one processor executes the computer program so that the electronic device executes the method described in the first aspect.

[0018] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0020] Figure 1 is a scene graph according to an embodiment of the present disclosure;

[0021] Figure 2 is a schematic diagram according to a first embodiment of the present disclosure;

[0022] Figure 3 is a schematic diagram according to a second embodiment of the present disclosure;

[0023] Figure 4 is a schematic diagram according to a third embodiment of the present disclosure;

[0024] Figure 5 is a schematic diagram according to a third embodiment of the present disclosure;

[0025] Figure 6 is a schematic diagram according to a third embodiment of the present disclosure;

[0026] Figure 7is a schematic diagram according to a fourth embodiment of the present disclosure;

[0027] Figure 8 is a schematic diagram according to a fifth embodiment of the present disclosure;

[0028] Fig. 9 A schematic block diagram of an example electronic device 900 that may be used to implement embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0029] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0030] Figure 1 It is a scene diagram according to an embodiment of the present disclosure. In today's data-intensive application environment, it is crucial to effectively process and store large amounts of data. As the amount of data continues to increase, traditional batch processing methods may lead to performance bottlenecks. When the amount of data reaches a certain scale, the amount of data that the local device 101 transmits to the shared storage system 102 in a single transmission will become very large, which will not only consume a large amount of network bandwidth, but also may cause an increase in delays in the data transmission process, thereby affecting the response speed and efficiency of the entire system.

[0031] The present invention provides a shared storage method, device and equipment for large files based on the ES system, which is applied to big data, information flow and data processing in computer technology, so as to achieve the purpose of shared storage that reduces the storage data volume of large files, realizes the effect of improving transmission efficiency and reducing the operation and maintenance cost of storage equipment.

[0032] In order to enable readers to more deeply understand the implementation principle of the present disclosure, the following Figure 2-Figure 8 right Figure 2 The illustrated embodiment is further refined.

[0033] Figure 2 is a schematic diagram according to the first embodiment of the present disclosure, such as Figure 2 As shown, the present disclosure provides a shared storage method for large files based on an ES system, the method comprising:

[0034] 201. In response to a file write instruction, write the data in the target large file indicated by the file write instruction into a local empty text file in a streaming write manner to obtain an initial text file; and perform compression encoding processing on the initial text file to obtain a local text file; wherein the file write instruction is used to instruct the target large file to be stored.

[0035] Exemplarily, the execution subject of this embodiment may be a local storage system.

[0036] When the system receives a file write instruction, the instruction may include information about the target large file, such as the file path, the file size, etc. The function of this instruction is to instruct the system to start processing the specified large file.

[0037] The system will create an empty local text file. Then, the system will read and write the data in the target large file block by block into this empty text file in a streaming writing manner. Streaming writing means that the data will be read and written in batches instead of loading the entire file into the memory at one time. This method can effectively reduce memory usage and improve writing efficiency. In the end, the system will get an initial text file containing the data of the target large file.

[0038] Next, the system will compress and encode the initial text file. Compression encoding can reduce the file size, thereby saving storage space and improving the efficiency of data transmission. The compressed file is a local text file. Compared with the original target large file, this file is smaller in size and more suitable for storage and transmission. For example: including GZIP, BZIP2, etc.

[0039] It should be noted that uploading all the data at once may cause excessive network or server load. Therefore, uploading the data in batches can effectively alleviate the network or server load.

[0040] The storage problem of large files is effectively handled through two steps: streaming writing and compression encoding. Streaming writing reduces memory usage and improves writing efficiency; while compression encoding reduces file size and saves storage space.

[0041] 202. Upload the data in the local text file to the rolling index of the ES system in batches.

[0042] For example, ES system: Elasticsearch is a distributed search and analysis engine suitable for processing large-scale data. In Elasticsearch, rolling index is a strategy for managing the index life cycle, which can automatically create new indexes and delete old indexes to maintain cluster performance and effective use of storage. ElasticSearch systems are generally retrieval systems with the advantages of high availability, lightweight, and low operation and maintenance costs.

[0043] Exemplarily, aliases are provided for read operations and write operations respectively so that the rolling process can be seamlessly switched. Upload The rolling index uploaded to the ES system is equivalent to a write operation.

[0044] Write operation: When the scrolling conditions are met, a new index will be created and the writeAlias ​​will be switched from the old index to the new index;

[0045] Read operation: Use the readAlias ​​alias defined in the rolling index template.

[0046] Exemplarily, the rolling condition may include: using a timed task to regularly detect the following three factors of the rolling index, and rolling the index if one of the conditions is met:

[0047] Condition 1: When the index size reaches 50G;

[0048] Condition 2: When the number of docment documents reaches 10 million, the index can be rolled to a new index;

[0049] Condition 3: The index was created more than 7 days ago.

[0050] Exemplarily, a set of basic APIs are provided to operate the large file shared storage system, including: getBufferWriter, getBufferReader, open, lostFileNames, exists, mkdir, isDirectory, deleteFile, etc.

[0051] For example, when performing a write operation, first determine the rolling index to be written, then write the large file in a streaming manner to the pod local file xxfile (equivalent to the initial text file), compress the streaming into zip, perform base64 encoding, and store it as the pod local text file - xxendcodedFile (equivalent to the local text file). Divide the text file - xxendcodedFile into multiple chunks according to the preset size (100mb) and upload them to the rolling index.

[0052] Based on the above characteristics, the read data is divided into multiple batches. Each batch contains a certain amount of data, which can be set according to the system performance and network bandwidth. The purpose of batching is to avoid network congestion or excessive server pressure caused by sending a large amount of data at one time.

[0053] In this embodiment, the target large file is streamed into a text file to obtain an initial text file; then the initial text file is compressed and encoded to obtain a local file. Then, the data in the local text file is uploaded to the rolling index of the ES system in batches. Since the large file is streamed into the text file and compressed and encoded, the amount of data in the shared storage of the large file is reduced. And for the convenience of query, the initial text file is cached, and then when the user or system needs to access the data, it can be directly obtained from the cache without having to read from the original file every time, thereby improving the data reading speed.

[0054] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0055] Figure 3 is a schematic diagram according to the second embodiment of the present disclosure, such as Figure 3 As shown, the present disclosure provides a shared storage method for large files based on an ES system, the method comprising:

[0056] 301. Perform compression encoding processing on the initial text file to obtain a local text file.

[0057] In one example, step 301 includes:

[0058] The first step of step 301 is to compress the initial text file in the streaming compression mode to obtain the initial text file in the target compression format;

[0059] The second step of step 301 is to encode the initial text file in the target compression format based on the Base64 encoding method to obtain the local text file.

[0060] Exemplarily, the file is compressed by a compression algorithm (such as gzip, bzip2, etc.). The compressed binary data is converted into Base64 encoding format. Base64 is an encoding method based on 64 printable characters to represent binary data, which is commonly used to transmit binary data in a text environment. The Base64-encoded data is saved as a local text file, that is, the final local text file.

[0061] Through these two steps, the initial text file is compressed and encoded into a local text file that can be safely transmitted and stored in a text environment.

[0062] 302. Upload the data in the local text file to the rolling index of the ES system in batches.

[0063] In one example, step 302 includes:

[0064] The first step of step 302 is to read data corresponding to the preset data reading value from the local text file based on the preset data reading value; wherein the preset data reading value indicates the data amount of the data read from the local text file each time.

[0065] The second step of step 302 is to upload the read data to the rolling index of the ES system.

[0066] Exemplarily, the preset data reading value indicates the amount of data read from the local text file each time. For example, it can be preset to read 100MB or 1MB of data each time.

[0067] According to the preset reading value, read the corresponding amount of data from the local text file. This step ensures that too much data will not be read at one time, causing memory overflow, and the amount of data read each time can be adjusted as needed.

[0068] Upload the read data to the rolling index of the Elasticsearch system. Rolling index is a strategy for managing the index lifecycle, which can automatically create new indexes and delete old indexes to maintain cluster performance and effective storage utilization. Since direct one-time uploading may cause excessive network or server load, the data is uploaded in batches.

[0069] In one example, the second step of step 302 includes:

[0070] Determine an index name corresponding to the read data in the ES system; and store the read data in a rolling index corresponding to the determined index name.

[0071] Exemplarily, an index name corresponding to the read data is determined according to business logic or data characteristics. For example, the index can be named according to a timestamp, a data type, or other identifier. The read data is uploaded to a rolling index corresponding to the determined index name.

[0072] Rolling indexes can significantly improve query performance by storing data shards in different indexes. Each index contains data within a time range, so when querying data for a specific time period, you only need to access the relevant index instead of scanning the entire data set. Using rolling indexes can make data maintenance tasks such as data archiving, compression, and deletion simpler and more efficient. You can archive or delete old indexes regularly to ensure that the system only retains necessary data.

[0073] 303. If it is determined that the local text file has a first preset identifier, it is determined that the local text file needs to be cached; wherein the first preset identifier indicates that the local text file is a cacheable file.

[0074] In order to facilitate understanding of the above step 303, Figure 4 Provide explanations, Figure 4 is a schematic diagram according to a third embodiment of the present disclosure.

[0075] 401. First, the large file is stored and written locally in a streaming manner. If it is determined that the local text file has a first preset identifier, it is determined that the local text file needs to be cached.

[0076] For large files that need to be stored, the system uses streaming writing to save them locally. This method can effectively reduce memory usage and increase file processing speed.

[0077] In one example, the first step of step 303 includes:

[0078] If it is determined that the local text file does not have the first preset identifier, it is determined that there is no need to cache the local text file, and the initial text file and the local text file are deleted.

[0079] 304. Cache the initial text file corresponding to the local text file; and delete the local text file.

[0080] Exemplarily, the first preset identifier indicates that the local text file is a cacheable file. For example, it can be judged by the file name, file header information or other metadata. If the local text file has the first preset identifier, cache processing is performed, that is, the local text file is cached for subsequent quick access. If the local text file does not have the first preset identifier, the initial text file and the local text file are deleted, that is, files that are no longer needed are cleaned up to save storage space of the local storage device.

[0081] When the text file - xxendcodedFile is written to the rolling index, if the file needs to be cached, xxcontentFile is stored and xxendcodedFile is deleted. If caching is not required, xxfile and xxendcodedFile are deleted from the local computer.

[0082] When performing a read operation, first determine whether the file to be read is a local cache file. If so, directly obtain the xxcontentFile file from the local stream.

[0083] 305. In response to a file read instruction, if it is determined that the target large file indicated by the file read instruction has a second preset identifier, based on the initial text file corresponding to the target large file in the local cache area, read the data in the initial text file corresponding to the target large file; wherein the file read instruction is used to instruct to read the data of the target large file; and the second preset identifier indicates that the target large file is a cacheable file.

[0084] In one example, the first step of step 305 includes:

[0085] If it is determined that the local cache area contains data of the target large file indicated by the file read instruction, it is determined that the initial text file corresponding to the target large file is cached in the local cache area, and the data in the initial text file corresponding to the target large file is read from the local cache area in a streaming reading manner.

[0086] like Figure 4 As shown, when reading a large file, it is first determined whether a cache is required. The determination process is based on whether the target large file indicated by the file reading instruction has a second preset identifier.

[0087] 402. When reading a large file, if it is determined that the target large file has a second preset identifier, which means that the file is cacheable, the system will first check whether the local cache area has stored the initial text file of the file. If the data of the file exists in the local cache area, the system will directly obtain the data from the cache area in a streaming reading manner, thereby improving the efficiency of data reading and reducing direct access to the original large file.

[0088] For example, the system receives a file reading instruction, which specifies a target large file to be read. The system checks whether the target large file has a second preset identifier. The second preset identifier indicates that the file can be cached.

[0089] If the file data is not available in the cache, the file needs to be read from a remote storage system, such as an ES system. If the system confirms that the initial text file corresponding to the target large file is cached in the local cache, the system reads the data in the initial text file from the cache in a streaming manner.

[0090] The read data is returned to the requester to complete the file reading operation.

[0091] By using cache, the system can avoid frequently reading data from disk or remote servers. Data in the cache can be accessed quickly, significantly improving the speed and efficiency of data reading. Reading data directly from the cache can reduce I / O operations to the underlying storage system. This not only improves data access speed, but also reduces the burden on the storage system, helping to improve the performance and stability of the entire system.

[0092] When the system detects that the target large file has the second preset identifier, it can directly read the data from the cache without parsing and processing the original file again. This simplifies the data processing process and reduces the possibility of errors.

[0093] In one example, if it is determined that the local cache does not contain the data of the target large file indicated by the file read instruction, the local text file corresponding to the target large file is obtained from the ES system; the obtained local text file is decoded and decompressed to obtain an initial text file corresponding to the obtained local text file, so as to feed back the data in the initial text file.

[0094] 306. In response to the file reading instruction, if it is determined that the target large file indicated by the file reading instruction does not have a second preset identifier, a local text file corresponding to the target large file is obtained from the ES system; the second preset identifier indicates that the target large file is a cacheable file;

[0095] Exemplarily, if the target large file does not have the second preset identifier, the system will obtain the local text file corresponding to the target large file from the ES system. This process ensures that necessary data can be read even without cache.

[0096] 307. Decode and decompress the acquired local text file to obtain an initial text file corresponding to the acquired local text file, and feed back the data in the initial text file.

[0097] Figure 5 is a schematic diagram according to a fourth embodiment of the present disclosure, Figure 5 As shown:

[0098] 501. Local text files obtained from the ES system need to be decoded

[0099] 502. After decoding and decompression, an initial text file corresponding to the local text file is obtained.

[0100] Exemplarily, the local text file obtained from the ES system needs to be decoded. Decoding is the process of restoring the encoded data to a readable format to ensure that the file content can be correctly parsed and understood. Decompression is the decompression operation on the obtained local text file. Through the decompression step, the data in the compressed file is released and restored to its original state, which is convenient for subsequent data processing and use. After decoding and decompression, the initial text file corresponding to the local text file is obtained. The system finally feeds back the data in the initial text file, completes the entire file reading process, and ensures the accuracy and integrity of the data.

[0101] Figure 6 is a schematic diagram according to the third embodiment of the present disclosure, Figure 6 As shown, the present disclosure provides a shared storage device 600 for large files based on an ES system, including:

[0102] The writing unit 601 is used to respond to the file writing instruction and write the data in the target large file indicated by the file writing instruction into a local empty text file in a streaming writing manner to obtain an initial text file.

[0103] The processing unit 602 is used to perform compression encoding processing on the initial text file to obtain a local text file; wherein the file write instruction is used to instruct to store the target large file.

[0104] The storage unit 603 is used to upload the data in the local text file to the rolling index of the ES system in batches.

[0105] The device of this embodiment can execute the technical solution in the above method. Its specific implementation process and technical principles are the same and will not be repeated here.

[0106] Figure 7 is a schematic diagram according to a fourth embodiment of the present disclosure, Figure 7 As shown, the present disclosure provides a shared storage device 700 for large files based on an ES system, including:

[0107] The writing unit 701 is used to respond to the file writing instruction and write the data in the target large file indicated by the file writing instruction into a local empty text file in a streaming writing manner to obtain an initial text file.

[0108] The processing unit 702 is used to perform compression encoding processing on the initial text file to obtain a local text file; wherein the file write instruction is used to instruct to store the target large file.

[0109] The storage unit 703 is used to upload the data in the local text file to the rolling index of the ES system in batches.

[0110] In one example, the processing unit 702 includes:

[0111] The compression module 7021 is used to compress the initial text file in a streaming compression manner to obtain the initial text file in a target compression format.

[0112] The encoding module 7022 is used to encode the initial text file in the target compression format based on the Base64 encoding method to obtain a local text file.

[0113] In one example, the storage unit 703 includes:

[0114] The reading module 7031 is used to read data corresponding to a preset data reading value from a local text file based on the preset data reading value; wherein the preset data reading value indicates the data size of the data read from the local text file each time.

[0115] The storage module 7032 is used to upload the read data to the rolling index of the ES system.

[0116] In one example, the storage module 7032 is specifically used for:

[0117] Determine an index name corresponding to the read data in the ES system; and store the read data in a rolling index corresponding to the determined index name.

[0118] In one example, the shared storage device 700 for large files based on the ES system further includes:

[0119] The first cache unit 704 is used to determine that the local text file needs to be cached if it is determined that the local text file has a first preset identifier; wherein the first preset identifier indicates that the local text file is a cacheable file; cache the initial text file corresponding to the local text file; and delete the local text file.

[0120] The first deleting unit 705 is configured to determine that there is no need to cache the local text file if it is determined that the local text file does not have the first preset identifier, and delete the initial text file and the local text file.

[0121] In one example, the shared storage device 700 for large files based on the ES system further includes:

[0122] The first reading unit 706 is used to respond to the file reading instruction. If it is determined that the target large file indicated by the file reading instruction has a second preset identifier, the data in the initial text file corresponding to the target large file in the local cache area is read.

[0123] The file read instruction is used to instruct to read data of the target large file; and the second preset identifier indicates that the target large file is a cacheable file.

[0124] In one example, when the first reading unit 706 reads data in the initial text file corresponding to the target large file based on the initial text file corresponding to the target large file in the local cache area, it is specifically used to:

[0125] If it is determined that the local cache area contains data of the target large file indicated by the file read instruction, it is determined that the initial text file corresponding to the target large file is cached in the local cache area, and the data in the initial text file corresponding to the target large file is read from the local cache area in a streaming reading manner.

[0126] In one example, the first reading unit 706 is further configured to:

[0127] If it is determined that the local cache area does not contain the data of the target large file indicated by the file read instruction, the local text file corresponding to the target large file is obtained from the ES system; the obtained local text file is decoded and decompressed to obtain an initial text file corresponding to the obtained local text file, so as to feed back the data in the initial text file.

[0128] In one example, the shared storage device 700 for large files based on the ES system further includes:

[0129] The second reading unit 707 is used to respond to the file reading instruction; if it is determined that the target large file indicated by the file reading instruction does not have the second preset identifier, the local text file corresponding to the target large file is obtained from the ES system; the second preset identifier indicates that the target large file is a cacheable file; the obtained local text file is decoded and decompressed to obtain an initial text file corresponding to the obtained local text file, so as to feed back the data in the initial text file.

[0130] In one example, the shared storage device 700 for large files based on the ES system further includes:

[0131] The second deleting unit is used to delete the initial text file and the local text file when it is determined that the local text file is not to be cached.

[0132] The device of this embodiment can execute the technical solution in the above method. Its specific implementation process and technical principles are the same and will not be repeated here.

[0133] Figure 8 is a schematic diagram according to a fifth embodiment of the present disclosure, Figure 8 As shown, the electronic device 800 in this embodiment may include: a processor 801 and a memory 802 .

[0134] The memory 802 is used to store programs; the memory 802 may include volatile memory (English: volatile memory), such as random-access memory (English: random-access memory, abbreviated: RAM), such as static random-access memory (English: static random-access memory, abbreviated: SRAM), double data rate synchronous dynamic random access memory (English: Double Data Rate Synchronous Dynamic Random Access Memory, abbreviated: DDR SDRAM), etc.; the memory may also include non-volatile memory (English: non-volatile memory), such as flash memory (English: flash memory). The memory 802 is used to store computer programs (such as applications, functional modules, etc. that implement the above method), computer instructions, etc., and the above computer programs, computer instructions, etc. can be partitioned and stored in one or more memories 802. And the above computer programs, computer instructions, data, etc. can be called by the processor 801.

[0135] The above-mentioned computer programs, computer instructions, etc. may be stored in partitions in one or more memories 802 . And the above-mentioned computer programs, computer instructions, etc. may be called by the processor 801 .

[0136] The processor 801 is used to execute the computer program stored in the memory 802 to implement each step of the method involved in the above embodiment.

[0137] For details, please refer to the relevant description in the previous method embodiment.

[0138] The processor 801 and the memory 802 may be independent structures or integrated structures. When the processor 801 and the memory 802 are independent structures, the memory 802 and the processor 801 may be coupled and connected via a bus 803 .

[0139] The electronic device of this embodiment can execute the technical solution in the above method, and its specific implementation process and technical principle are the same, which will not be repeated here.

[0140] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0141] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, which includes: a computer program, the computer program is stored in a readable storage medium, at least one processor of an electronic device can read the computer program from the readable storage medium, and at least one processor executes the computer program so that the electronic device executes the solution provided by any of the above embodiments.

[0142] Fig. 9 A schematic block diagram of an example electronic device 900 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0143] like Fig. 9 As shown, the device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0144] A number of components in the device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0145] The computing unit 901 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 901 performs the various methods and processes described above, such as a shared storage method for large files based on an ES system. For example, in some embodiments, a shared storage method for large files based on an ES system may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 900 via ROM 902 and / or a communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of the shared storage method for large files based on the ES system described above may be executed. Alternatively, in other embodiments, the computing unit 901 may be configured to execute a shared storage method for large files based on an ES system in any other appropriate manner (for example, by means of firmware).

[0146] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0147] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0148] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0149] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0150] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0151] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The server may also be a server of a distributed system, or a server combined with a blockchain.

[0152] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0153] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A shared storage method for large files based on an ES system, comprising: In response to a file write instruction, the data in the target large file indicated by the file write instruction is written into a local empty text file in a streaming write manner to obtain an initial text file; and the initial text file is compressed and encoded to obtain a local text file; wherein the file write instruction is used to instruct the target large file to be stored; The data in the local text file is uploaded to the rolling index of the ES system in batches.

2. The method according to claim 1, wherein: The initial text file is compressed and encoded to obtain a local text file, including: Compressing the initial text file in a streaming compression manner to obtain an initial text file in a target compression format; Based on the Base64 encoding method, the initial text file in the target compression format is encoded to obtain the local text file.

3. The method according to claim 1, wherein: Upload the data in the local text file to the rolling index of the ES system in batches, including: Based on a preset data reading value, reading data corresponding to the preset data reading value from the local text file; wherein the preset data reading value indicates the data amount of the data read from the local text file each time; The read data is uploaded to the rolling index of the ES system.

4. The method according to claim 3, wherein: Uploading the read data to the rolling index of the ES system includes: Determine an index name corresponding to the read data in the ES system; and store the read data in a rolling index corresponding to the determined index name.

5. The method according to any one of claims 1 to 4, further comprising: If it is determined that the local text file has a first preset identifier, it is determined that the local text file needs to be cached; wherein the first preset identifier indicates that the local text file is a cacheable file; The initial text file corresponding to the local text file is cached; and the local text file is deleted.

6. The method according to claim 5, further comprising: If it is determined that the local text file does not have the first preset identifier, it is determined that there is no need to cache the local text file, and the initial text file and the local text file are deleted.

7. The method according to any one of claims 1 to 6, further comprising: In response to the file reading instruction, if it is determined that the target large file indicated by the file reading instruction has a second preset identifier, based on the initial text file corresponding to the target large file in the local cache area, read the data in the initial text file corresponding to the target large file; The file reading instruction is used to instruct reading data of a target large file; and the second preset identifier indicates that the target large file is a cacheable file.

8. The method according to claim 7, based on the initial text file corresponding to the target large file in the local cache area, reading the data in the initial text file corresponding to the target large file comprises: If it is determined that the local cache area contains data of the target large file indicated by the file read instruction, it is determined that the initial text file corresponding to the target large file is cached in the local cache area, and the data in the initial text file corresponding to the target large file is read from the local cache area in a streaming reading manner.

9. The method according to claim 8, further comprising: If it is determined that the local cache area does not have the data of the target large file indicated by the file read instruction, obtaining a local text file corresponding to the target large file from the ES system; The acquired local text file is decoded and decompressed to obtain an initial text file corresponding to the acquired local text file, so as to feed back the data in the initial text file.

10. The method according to any one of claims 1 to 9, further comprising: In response to a file read instruction; If it is determined that the target large file indicated by the file reading instruction does not have a second preset identifier, a local text file corresponding to the target large file is obtained from the ES system; the second preset identifier indicates that the target large file is a cacheable file; The acquired local text file is decoded and decompressed to obtain an initial text file corresponding to the acquired local text file, so as to feed back the data in the initial text file.

11. The method according to any one of claims 1 to 10, further comprising: When it is determined not to cache the local text file, the initial text file and the local text file are deleted.

12. A shared storage device for large files based on an ES system, comprising: A writing unit, for responding to a file writing instruction, writing the data in the target large file indicated by the file writing instruction into a local empty text file in a streaming writing manner to obtain an initial text file; A processing unit, used for performing compression encoding processing on the initial text file to obtain a local text file; wherein the file writing instruction is used to instruct to store the target large file; The storage unit is used to upload the data in the local text file to the rolling index of the ES system in batches.

13. The device according to claim 12, wherein the processing unit comprises: A compression module, used for compressing the initial text file in a streaming compression manner to obtain an initial text file in a target compression format; The encoding module is used to encode the initial text file in the target compression format based on the Base64 encoding method to obtain the local text file.

14. The device according to claim 12, wherein: The storage unit comprises: A reading module, configured to read data corresponding to a preset data reading value from the local text file based on the preset data reading value; wherein the preset data reading value indicates the data amount of the data read from the local text file each time; The storage module is used to upload the read data to the rolling index of the ES system.

15. The device according to claim 14, wherein: The storage module is specifically used for: Determine an index name corresponding to the read data in the ES system; and store the read data in a rolling index corresponding to the determined index name.

16. The apparatus according to any one of claims 12 to 15, further comprising: The first cache unit is used to determine that the local text file needs to be cached if it is determined that the local text file has a first preset identifier; wherein the first preset identifier indicates that the local text file is a cacheable file; cache the initial text file corresponding to the local text file; and delete the local text file.

17. The apparatus according to claim 16, further comprising: The first deleting unit is configured to determine that it is not necessary to cache the local text file if it is determined that the local text file does not have the first preset identifier, and to delete the initial text file and the local text file.

18. The apparatus according to any one of claims 12 to 17, further comprising: A first reading unit is used for responding to a file reading instruction, and if it is determined that the target large file indicated by the file reading instruction has a second preset identifier, reading data in the initial text file corresponding to the target large file based on the initial text file corresponding to the target large file in the local cache area; The file reading instruction is used to instruct reading data of a target large file; and the second preset identifier indicates that the target large file is a cacheable file.

19. The device according to claim 18, when the first reading unit reads the data in the initial text file corresponding to the target large file based on the initial text file corresponding to the target large file in the local cache area, is specifically used to: If it is determined that the local cache area contains data of the target large file indicated by the file read instruction, it is determined that the initial text file corresponding to the target large file is cached in the local cache area, and the data in the initial text file corresponding to the target large file is read from the local cache area in a streaming reading manner.

20. The device according to claim 19, wherein the first reading unit is further configured to: If it is determined that the local cache area does not have the data of the target large file indicated by the file read instruction, obtaining a local text file corresponding to the target large file from the ES system; The acquired local text file is decoded and decompressed to obtain an initial text file corresponding to the acquired local text file, so as to feed back the data in the initial text file.

21. The apparatus according to any one of claims 12 to 20, further comprising: A second reading unit, used to respond to a file reading instruction; If it is determined that the target large file indicated by the file reading instruction does not have a second preset identifier, a local text file corresponding to the target large file is obtained from the ES system; the second preset identifier indicates that the target large file is a cacheable file; the obtained local text file is decoded and decompressed to obtain an initial text file corresponding to the obtained local text file, so as to feed back the data in the initial text file.

22. The apparatus according to any one of claims 12 to 21, further comprising: The second deleting unit is configured to delete the initial text file and the local text file when it is determined that the local text file is not to be cached.

23. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 11.

24. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-11.

25. A computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 11.