Method and apparatus for processing streaming data, electronic device, and storage medium

CN116431063BActive Publication Date: 2026-09-29NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310220485.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-03
Publication Date
2026-09-29
Estimated Expiration
2043-03-03

AI Technical Summary

Benefits of technology

[0019]本申请实施例所提供的技术方案,当前检查点到来时,先检查该检查点时间间隔对应的存储空间中所存储的流式数据的数据量是否小于数据量阈值,如果小于则需要将该存储空间中的流式数据复制到下一个检查点时间间隔对应的存储空间中。这样,下一个检查点时间间隔对应的存储空间中将存储有两个时间间隔对应的流式数据,从而实现了存储空间的合并,减少了存储系统中“小文件”的数量。同时,还需要暂时保留“小文件”,在检查点时刻更新“小文件”的状态为可读状态,这样是为了在下一个检查点时间间隔内,“小文件”中的数据可以被下游应用读取,避免数据读取延迟。待下一个检查点到来,下一个检查点时间间隔对应的存储空间的流式数据写入过程结束后,再将“小文件”删除。而如果当前检查点时间间隔对应的存储空间中所存储的流式数据的数据量大于或等于数据量阈值时,则说明该检查点时间间隔写入的数据量足够多,该存储空间也为“大文件”,不会对存储空间的管理造成拖累。那么就保留该存储空间即可。这样,在流式数据写入过程中,存储系统就能自动完成“小文件”的合并,存储系统中的“小文件”的数量将大大减少,从而提高存储系统的管理性能。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116431063B_ABST
    Figure CN116431063B_ABST
Patent Text Reader

Abstract

The application provides a processing method and device of streaming data, electronic equipment and storage medium, and applies to the technical field of computers. The method comprises the following steps: when a first checkpoint arrives, detecting the data amount of the streaming data stored in a first storage space in a storage system. When the data amount of the streaming data stored in the first storage space is less than a data amount threshold, copying the streaming data stored in the first storage space to a second storage space of the storage system. Updating the state of the first storage space to a readable state. The method can realize the merging of storage spaces without manual operation, thereby reducing the number of small files generated in the streaming data writing process of the storage system, and further improving the performance of the storage system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, electronic device and storage medium for processing streaming data. Background Technology

[0002] Game logs, or game records, are used to document the game system's operational information during gameplay. Once generated, game logs are temporarily stored in a Kafka queue used for processing streaming data. The Flink processing engine then periodically reads the game logs from the Kafka queue and writes all game logs read within a period to a file in the distributed file system (Hadoop Distributed File System, HDFS) for querying by downstream applications (such as Hive databases).

[0003] When Flink writes game logs to a file in HDFS, the file's state is invisible to downstream tools. At the end of a cycle (checkpoint), the file's state is updated to be visible to downstream tools. Therefore, to ensure timely reading of game logs by downstream tools, a short cycle needs to be set for HDFS. However, a short cycle results in less data being written to each file, potentially creating a large number of small "small files."

[0004] Since each file requires a slot, and HDFS needs to set up an interface for each file, a large number of "small files" will severely waste HDFS resources and impact HDFS performance. Furthermore, downstream tools need to constantly jump from one file to the next when reading data, significantly affecting their reading efficiency. Therefore, how to merge "small files" in HDFS to reduce their number has become an urgent problem to solve. Summary of the Invention

[0005] In view of this, this application provides a method, apparatus, electronic device and storage medium for processing streaming data, thereby solving the problem of poor HDFS performance caused by excessive small files generated by the checkpoint mechanism during the data writing process in the prior art.

[0006] A first aspect of this application provides a method for processing streaming data, the method comprising:

[0007] Streaming data is processed by a streaming engine, and the processed streaming data is written to the storage system according to the checkpoint interval of the streaming engine; the streaming data written within different checkpoint intervals is stored in different storage spaces of the storage system.

[0008] When the first checkpoint arrives, the amount of streaming data stored in the first storage space of the storage system is detected; wherein, the first storage space is used to store the streaming data written by the processing engine during the first checkpoint time interval, and the first checkpoint is the end time of the first checkpoint time interval.

[0009] When the amount of streaming data stored in the first storage space is less than the data amount threshold, the streaming data stored in the first storage space is copied to the second storage space of the storage system; the second storage space is used to store the streaming data written by the streaming processing engine within the second checkpoint time interval, and the second checkpoint time interval is the next checkpoint time interval adjacent to the first checkpoint time interval.

[0010] Update the status of the first storage space to readable.

[0011] A second aspect of this application provides a streaming processing engine that processes streaming data and writes the processed streaming data into a storage system according to the checkpoint time interval of the streaming processing engine; wherein the streaming data written within different checkpoint time intervals is stored in different storage spaces of the storage system; the streaming processing engine includes:

[0012] The detection unit is used to detect the amount of streaming data stored in the first storage space of the storage system when the first checkpoint arrives; wherein, the first storage space is used to store the streaming data written by the streaming processing engine within the first checkpoint time interval, and the first checkpoint is the end time of the first checkpoint time interval.

[0013] The processing unit is used to copy the streaming data stored in the first storage space to the second storage space of the storage system when the amount of streaming data stored in the first storage space is less than the data amount threshold; the second storage space is used to store the streaming data written by the streaming processing engine within the second checkpoint time interval, and the second checkpoint time interval is the next checkpoint time interval adjacent to the first checkpoint time interval.

[0014] The processing unit is also used to update the state of the first storage space to a readable state.

[0015] A third aspect of this application also provides a server, including: a processor and a memory. Wherein:

[0016] The memory stores the instructions that the computer executes.

[0017] The processor executes computer execution instructions, causing the electronic device to perform the streaming data processing method described in the first aspect above.

[0018] A fourth aspect of this application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the streaming data processing method described in the first aspect above.

[0019] The technical solution provided in this application embodiment, when the current checkpoint arrives, first checks whether the amount of streaming data stored in the storage space corresponding to the checkpoint time interval is less than the data volume threshold. If it is less, the streaming data in the storage space needs to be copied to the storage space corresponding to the next checkpoint time interval. In this way, the storage space corresponding to the next checkpoint time interval will store streaming data corresponding to two time intervals, thereby achieving storage space merging and reducing the number of "small files" in the storage system. Simultaneously, it is also necessary to temporarily retain the "small files," updating their status to readable at each checkpoint. This ensures that the data in the "small files" can be read by downstream applications within the next checkpoint time interval, avoiding data read latency. Once the next checkpoint arrives and the streaming data writing process in the storage space corresponding to the next checkpoint time interval is completed, the "small files" are deleted. However, if the amount of streaming data stored in the storage space corresponding to the current checkpoint time interval is greater than or equal to the data volume threshold, it indicates that the amount of data written in that checkpoint time interval is sufficient, and the storage space is also a "large file," which will not hinder storage space management. Therefore, the storage space can be retained. In this way, during the streaming data writing process, the storage system can automatically merge "small files", which will greatly reduce the number of "small files" in the storage system and thus improve the management performance of the storage system. Attached Figure Description

[0020] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 A system architecture diagram of a streaming data storage system provided in this application embodiment;

[0022] Figure 2 A flowchart illustrating a streaming data processing method provided in an embodiment of this application;

[0023] Figure 3 A flowchart illustrating another method for processing streaming data provided in an embodiment of this application;

[0024] Figure 4This is a schematic diagram of the structure of a streaming engine provided in an embodiment of this application;

[0025] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0026] In view of this, this application provides a method, apparatus, electronic device and storage medium for processing streaming data, thereby solving the problem of poor HDFS performance caused by excessive small files generated by the checkpoint mechanism during the data writing process in the prior art.

[0027] To enable those skilled in the art to better understand the technical solutions of this application, the application will be clearly and completely described below with reference to the accompanying drawings of the embodiments. However, this application can be implemented in many other ways different from those described above. Therefore, based on the embodiments provided in this application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this application.

[0028] It should be noted that the terms "first," "second," "third," etc., in the claims, specification, and drawings of this application are used to distinguish similar objects and are not used to describe a specific order or sequence. Such data are interchangeable where appropriate so that the embodiments of this application described herein can be implemented in a sequence other than that shown or described herein. Furthermore, the terms "comprising," "having," and their variations are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses.

[0029] First, the technical background of this application will be explained:

[0030] Game logs, or game records, are used to document the game system's operational information during gameplay. Game logs are continuously generated during game execution, so Kafka, which processes streaming data, can be used to manage them, temporarily storing the logs in a queue. The processing engine Flink then reads the game logs from the Kafka queue and writes them to the HDFS distributed file system for downstream applications to query.

[0031] The process of Flink writing game logs to HDFS is as follows: Flink periodically reads newly generated game logs from Kafka and writes all the game logs read within a period to a file in the HDFS distributed file system. While Flink is writing game logs to the HDFS file, the file's state is invisible to downstream tools. Therefore, Flink needs to periodically write data and then write the game logs generated in different periods to different files. This means that every certain period, new game logs generated within that period are written to a separate file for storage. When that period ends, Flink stops writing data to that file and instead writes data to a new file. This ensures that the game logs can be queried promptly at the end of each period. Understandably, to ensure the timeliness of downstream tools reading game logs, a relatively short period needs to be set for HDFS. However, a short period results in a smaller amount of data written to each file, potentially creating a large number of small "small files".

[0032] Since each file requires a slot, and HDFS needs to set up an interface for each file, a large number of "small files" will severely waste HDFS resources and impact HDFS performance. Furthermore, downstream tools need to constantly jump from one file to the next when reading data, significantly affecting their reading efficiency. Merging multiple "small files" in HDFS would improve HDFS performance. Current technology typically involves system administrators manually calling relevant APIs to merge files once all files in a partition have been written. For example, in a log directory partitioned by date, when no files are written to that partition, a command is called at a predetermined time to merge the "small files" within that partition. This method requires manual merging and can only be performed after all files in the partition have been written, resulting in untimely and extremely poor efficiency. Therefore, how to efficiently merge "small files" in HDFS to reduce their number has become a pressing problem.

[0033] To address the aforementioned technical problems, this application provides a method, apparatus, electronic device, and storage medium for processing streaming data. In an embodiment of this application, when a checkpoint arrives (the end of a cycle), it is necessary to check whether the amount of streaming data stored in the storage space (file) corresponding to that cycle is less than a data volume threshold. If it is less, it indicates that the storage space is a "small file," and the streaming data in the "small file" needs to be copied to the storage space corresponding to the next cycle. Thus, the storage space corresponding to the next cycle will store streaming data generated in two cycles. This allows for timely and effective merging of "small files," reducing the number of "small files" in the storage system. The method, apparatus, terminal, and computer-readable storage medium of this application will be further described in detail below with reference to specific embodiments and accompanying drawings.

[0034] The network architecture in the embodiments of this application will be briefly described below. Figure 1 A system architecture diagram of a streaming data storage system provided in this application embodiment is shown below. Figure 1 As shown, the system includes a Kafka queue, a Flink processing engine, an HDFS distributed file system, and downstream applications.

[0035] Kafka is a leading message middleware, essentially a data storage platform. As open-source software, it boasts advantages such as support for multi-language application development and real-time large-scale message processing. The Kafka system primarily consists of three roles: producer, consumer, and broker. A Kafka cluster comprises multiple broker nodes. When a producer generates new data, it is assigned to a specific topic. A topic is further divided into multiple partitions, which are deployed across multiple brokers. Producers continuously send data to a particular topic; broker nodes temporarily store data from different topics and forward it to consumers, who then use and process the data.

[0036] With the continuous increase in business data volume and the diversification of business needs, many scenarios in the current big data environment use Kafka as a data message queue to relay data messages. The broker server in Kafka receives data messages sent from multiple parties, processes the data, and then writes it to the HDFS distributed file system. Downstream applications can then directly query the data on HDFS.

[0037] In this system, the game system, acting as a producer, continuously generates game logs, which are then temporarily stored in a Kafka queue. The processing engine, Flink, then reads the game logs from the Kafka queue and writes them to files in the HDFS distributed file system. The game log writing process is periodic; Flink periodically reads newly added game logs from the Kafka queue, storing each period's logs in a separate file in HDFS. When a period ends, Flink stops writing to that file and instead starts writing to a newly created file. Furthermore, after each period, the corresponding file is promptly updated to a readable state, making it visible to downstream applications. This allows downstream applications to easily view the game logs in their files.

[0038] Based on the above network architecture Figure 2 This is a flowchart illustrating a streaming data processing method provided in an embodiment of this application. Figure 2 As shown, the method may include the following steps:

[0039] 201. When the first checkpoint arrives, detect the amount of streaming data in the first storage space of the storage system.

[0040] The streaming engine processes streaming data. This includes, for example, data cleaning, data integration, and data transformation. Specifically, data cleaning involves filling in corrupted values ​​in the streaming data, deleting outliers, and correcting problematic data. Data integration combines multiple streaming data sets according to certain rules to obtain composite data. Data transformation converts the received streaming data to conform to the data type requirements of downstream applications. After processing the received streaming data, the streaming engine stores the processed data in a storage system for querying or access by other downstream applications.

[0041] During the processing of streaming data, the streaming engine stores streaming data based on checkpoints. Specifically, the streaming engine periodically writes streaming data to the storage system, and streaming data generated within different checkpoint intervals is written to different storage spaces in the storage system. That is, when a checkpoint arrives, the storage space corresponding to the previous checkpoint interval will stop writing data, and then newly generated streaming data from that checkpoint onwards will be written to a new storage space.

[0042] For example, in Figure 1In the described system architecture scenario, the processing engine Flink reads streaming data (such as game log data) from a Kafka queue and writes the read streaming data to files (storage space) in the HDFS distributed file system. It's understandable that the processing engine Flink periodically writes streaming data to HDFS. The streaming data read in each cycle is independently written to a file in HDFS. When a cycle ends, at the end time (checkpoint), the processing engine Flink stops writing streaming data to the current file and instead writes new streaming data to a new file. Furthermore, HDFS updates the file status of the current file to a readable state at the checkpoint, making it available for downstream applications to query and access. It's understandable that the file is unreadable during the data writing process; it only becomes visible to downstream applications after the writing is complete.

[0043] In this step, the first storage space is used to store the streaming data written to the storage system by the streaming engine within the current checkpoint time interval. The first checkpoint is the end time of the current checkpoint time interval. When the first checkpoint arrives, the streaming engine stops writing streaming data to the first storage space and begins writing data to the second storage space corresponding to the next checkpoint time interval. At this time, it is necessary to detect the amount of data written to the first storage space during the current checkpoint time interval. Based on the amount of streaming data written, it is determined whether the first storage space is a "large file" or a "small file". Understandably, if the first storage space is a "large file", it will not hinder the file management of the storage system. However, if it is a "small file", then storage space merging is required.

[0044] 202. Determine whether the amount of streaming data stored in the first storage space is less than the data amount threshold. If yes, proceed to step 203. If no, proceed to step 205.

[0045] After obtaining the amount of streaming data stored in the first storage space, it needs to be compared with a preset data volume threshold. If the amount of data stored in the first storage space is greater than or equal to the data volume threshold, the first storage space is considered a "large file". If the amount of data stored in the first storage space is less than the data volume threshold, the first storage space is considered a "small file". Since too many "small files" will seriously affect cluster scalability, "small files" need to be merged later, while "large files" can be stored independently in the storage system.

[0046] For example, in Figure 1In the described system architecture scenario, the storage system is the distributed file system HDFS. The data volume threshold can be the amount of data that a data block can store in HDFS, typically 128MB. This means that each file (storage space) in HDFS must occupy at least one block to prevent the creation of too many "small files." Therefore, if the amount of data stored in a file is less than 128MB after a read / write cycle, the file is determined to be a "small file," and then "small files" need to be merged. If it is greater than or equal to 128MB, the file can be stored normally. Understandably, the amount of data written to a file by the Flink processing engine in one cycle is related to the file data generation rate. Some files are generated quickly, so Flink writes more data. Conversely, if file data is generated slowly, Flink writes less data. Therefore, preferably, the data volume threshold can be reasonably planned based on the amount of file data stored historically.

[0047] For example, the data volume of each historical file in multiple historical files in HDFS can be obtained, then the average data volume of these historical files can be calculated, and this average data volume can be used as the data volume threshold. Alternatively, the data stored in multiple historical files can be deduplicated to obtain the data volume of each deduplicated historical file. Then, the average data volume can be calculated based on the deduplicated data volume of each historical file, and this average data volume can be used as the data volume threshold. Another example is adding constraints to determine a minimum data volume. When the calculated average data volume is greater than this minimum data volume, the average data volume is used as the data volume threshold. Conversely, when the calculated average data volume is less than this minimum data volume, the minimum data volume is used as the data volume threshold. It is understood that the data volume threshold can be flexibly set according to storage needs, and there are no specific limitations.

[0048] 203. Copy the streaming data stored in the first storage space to the second storage space.

[0049] If the data stored in the first storage space is less than the data volume threshold, the streaming data stored in the first storage space needs to be copied to the second storage space corresponding to the next checkpoint interval. Then, after the checkpoint arrives, the streaming engine begins writing the newly added streaming data to the second storage space. Thus, the second storage space includes the streaming data written in the previous checkpoint interval and the streaming data written in the next checkpoint interval, both stored in the first storage space. This is equivalent to merging the storage spaces corresponding to two consecutive checkpoint intervals, merging the data written in both checkpoint intervals into one storage space (the second storage space). This reduces the probability of "small files" being generated.

[0050] 204. Update the status of the first storage space to readable.

[0051] After copying, the first storage space still needs to be retained because the second storage space is still undergoing data writing during the second checkpoint interval, making it unreadable during this time. Therefore, the data copied from the first storage space to the second storage space cannot be queried by downstream applications during the second checkpoint interval. Downstream applications should be able to access the streaming data in the first storage space after the first checkpoint. If the first storage space is deleted immediately, downstream applications will have to wait until the second checkpoint to read the corresponding streaming data from the first storage space, causing delays in data retrieval. Therefore, even if data from the storage space corresponding to the previous cycle is copied and backed up in the storage space corresponding to the next cycle, the storage space corresponding to the previous cycle must be retained first. Furthermore, the status of the storage space corresponding to the previous cycle needs to be updated to a readable state so that downstream applications can query the data in the storage space corresponding to the previous cycle promptly and effectively.

[0052] 205. Update the status of the first storage space to readable.

[0053] If the amount of data stored in the first storage space is greater than or equal to the data volume threshold, then the amount of data stored in the first storage space is considered sufficient, and there is no need to merge them. Since "large files" do not burden the management of the storage system, the status of the first storage space can be directly updated to a readable state, and the first storage space can be saved in the storage system.

[0054] 206. Write the newly added streaming data into the second storage space.

[0055] Understandably, data stored in the first storage space is copied to the second storage space. When the next checkpoint interval begins, the streaming engine needs to write the newly added streaming data to the second storage space. That is, the streaming engine stops writing streaming data to the first storage space (whose state has changed) and starts writing data to the second storage space. At this time, the second storage space will include the streaming data written in the previous checkpoint interval, as well as the streaming data written in the current checkpoint interval. Understandably, when the second checkpoint arrives again, data writing to the second storage space will also stop. At this time, the state of the second storage space will be changed to a readable state, and the second storage space will include the streaming data corresponding to the previous checkpoint interval (the streaming data written to the first storage space) and the streaming data written in the next checkpoint interval. Downstream applications can then retrieve the data from the first storage space by querying the second storage space, at which point the first storage space can be deleted. This reduces the number of "small files" in the storage system and improves the file management performance of the storage space. Downstream applications also do not need to frequently jump between different storage spaces when querying data. Therefore, while ensuring the timeliness of reading for downstream applications, the query efficiency of downstream applications is improved.

[0056] Understandably, if the data writing process in the second storage space stops, and the total amount of streaming data in the second storage space is still less than the data volume threshold—meaning the second storage space containing multiple cycles of streaming data is still a "small file"—then all its data will continue to be copied to the storage space corresponding to the next checkpoint time interval, and merged again. Once another checkpoint time interval ends and the status of its corresponding storage space is updated to a readable state, the second storage space will be deleted. This process continues until the storage space corresponding to a certain checkpoint time interval becomes a "large file" after the data writing process is complete; in this case, the "large file" will be retained. When another checkpoint arrives, the data copying process will no longer occur; streaming data will simply be written directly to a new storage space.

[0057] The technical solution provided in this application embodiment, when the current checkpoint arrives, first checks whether the amount of streaming data stored in the storage space corresponding to the checkpoint time interval is less than the data volume threshold. If it is less, the streaming data in the storage space needs to be copied to the storage space corresponding to the next checkpoint time interval. In this way, the storage space corresponding to the next checkpoint time interval will store streaming data corresponding to two time intervals, thereby achieving storage space merging and reducing the number of "small files" in the storage system. Simultaneously, it is also necessary to temporarily retain the "small files," updating their status to readable at each checkpoint. This ensures that the data in the "small files" can be read by downstream applications within the next checkpoint time interval, avoiding data read latency. Once the next checkpoint arrives and the streaming data writing process in the storage space corresponding to the next checkpoint time interval is completed, the "small files" are deleted. However, if the amount of streaming data stored in the storage space corresponding to the current checkpoint time interval is greater than or equal to the data volume threshold, it indicates that the amount of data written in that checkpoint time interval is sufficient, and the storage space is also a "large file," which will not hinder storage space management. Therefore, the storage space can be retained. In this way, during the streaming data writing process, the storage system can automatically merge "small files", which will greatly reduce the number of "small files" in the storage system and thus improve the management performance of the storage system.

[0058] Based on the above embodiments, Figure 3 This is a flowchart illustrating another streaming data processing method provided in this application embodiment, to fully describe the steps that may be involved in the embodiments of this application, such as... Figure 3 As shown, the method for processing streaming data may include the following steps:

[0059] 301. When the first checkpoint arrives, determine the target data in the streaming data stored in the first storage space.

[0060] Understandably, to conserve storage resources, the data copying process cannot be performed too many times. However, in some scenarios, streaming data generation is slow and the amount of data generated is small. Therefore, after multiple copies, the storage space may still contain "small files" even after storing multiple cycles of streaming data. In such cases, it is necessary to terminate the storage space merging process, stop data copying, and store the "small files" in the storage system. If it is necessary to improve the management performance of the storage system later, then after all streaming data has been written, other strategies can be used to merge the storage space again to eliminate the "small files".

[0061] Therefore, when the first checkpoint arrives, it is first necessary to determine whether the first storage space corresponding to the first checkpoint time interval contains the copied target data. Furthermore, based on the state of the target data, it is necessary to determine whether to copy all data from the first storage space into the second storage space corresponding to the next cycle. Understandably, the target data refers to the streaming data copied from the historical storage space corresponding to the historical checkpoint time intervals within the first storage space.

[0062] 302. Determine if the target data is in a preset state. If yes, proceed to step 303. If no, proceed to step 304.

[0063] When the target data is in a preset state, the streaming data stored in the first storage space will not be copied to the second storage space, even if the first storage space is a "small file," it can be stored directly in the storage system. The following section provides a detailed explanation of the process for determining whether to copy streaming data from the first storage space to the second storage space based on several scenarios.

[0064] Scenario 1:

[0065] For example, the data storage records of the target data in the first storage space can be queried to obtain the number of times the target data has been copied. Then, the maximum number of copies is determined. If the maximum number of copies reaches a first preset number, the streaming data of the first storage space will no longer be copied to the second storage space corresponding to the next checkpoint time interval.

[0066] Understandably, the target data in the first storage space may come from multiple historical storage spaces. For example, the file corresponding to period A is file a, and the file corresponding to period B is file b. Period A is the checkpoint interval preceding period B, and period B is the period preceding the first checkpoint interval. After period A ends, storage space a is a "small file." Therefore, the streaming data in storage space a needs to be copied to storage space b. After period B ends, storage space a is still a "small file," but storage space b now contains the streaming data from storage space a and the newly written streaming data from period B. Then, the streaming data in storage space b is copied to the first storage space. At this point, the first storage space contains the data from storage space a, the data written during period B, and the data written during the first checkpoint interval. Therefore, the target data in the first storage space is the data from storage space a and the data written during period B. The data from storage space a is copied twice, and the file written during period B is copied once. Thus, the maximum number of copies of the target data can be determined to be two.

[0067] If the maximum number of copies corresponding to the target data in the first storage space reaches a preset number, it means that after several read cycles (several checkpoint time intervals), the merged storage space is still a "small file". In this case, the merging process is terminated, and the file is directly retained in the storage system.

[0068] Scenario 2:

[0069] For example, the data storage records of the target data in the first storage space can be queried to obtain the number of times the target data has been copied. Then, the average number of times the target data has been copied is determined. If the average number of copies reaches a second preset number, the streaming data stored in the first storage space will no longer be copied to the second storage space corresponding to the next checkpoint time interval.

[0070] Similarly, the target data in the first storage space may come from multiple historical storage spaces. For example, storage space 'a' corresponds to period A, and storage space 'b' corresponds to period B. Period A is the checkpoint interval preceding period B, and period B is the checkpoint interval preceding the first checkpoint interval. After period A ends, 'a' is a "small file." Therefore, the streaming data in 'a' needs to be copied to 'b'. After period B ends, storage space 'b' is still a "small file," but it now contains the data from 'a' and the streaming data written during period B. Then, the data in 'b' is copied to the first storage space, which now contains the data from 'a', the streaming data written during period B, and the streaming data written during the first checkpoint interval. Therefore, the target data in the first storage space consists of the data from file 'a' and the streaming data written during period B. The data from 'a' is copied twice, and the streaming data written during period B is copied once, so the average number of copies of the target data is 1.5.

[0071] If the average number of copies of the target data in the first storage space reaches a preset number, it means that after several cycles, the merged storage space is still filled with "small files." Furthermore, the initial streaming data is no longer time-sensitive, so the merging process is terminated, and the "small files" are simply retained in the storage system.

[0072] Scenario 3:

[0073] For example, the data storage record of the target data in the first storage space can be queried to obtain the initial write time of the target data. Then, the write duration of the target data is determined based on the initial write time and the current time. If the write duration reaches the preset duration, the streaming data of the first storage space will no longer be copied to the second storage space corresponding to the next checkpoint time interval.

[0074] Similarly, the target data in the first storage space may come from multiple historical storage spaces. For example, storage space 'a' corresponds to period A, and storage space 'b' corresponds to period B. Period A is the checkpoint interval preceding period B, and period B is the checkpoint interval preceding the first checkpoint interval. After period A ends, 'a' is a "small file." Therefore, the streaming data in 'a' needs to be copied to 'b'. After period B ends, 'b' is still a "small file," but it now stores the data from 'a' and the streaming data written during period B. Then, the data in 'b' is copied to the first storage space, which now stores the data from 'a', the streaming data written during period B, and the streaming data written during the first checkpoint interval. Therefore, the target data in the first storage space is the data from 'a' and the streaming data written during period B. The write time corresponding to the data from 'a' is the initial write time. At this point, it's necessary to query the initial write time of the data from 'a'. For example, if the write time is 20:30, and the current first checkpoint time is 20:40, then the data from 'a' has been written for 10 minutes.

[0075] If the data corresponding to 'a' has been written for a preset period of time, it means that after several rounds of merging, the merged storage space is still a "small file". Furthermore, the original data is no longer time-sensitive, so the merging process needs to be terminated, and the first storage space, which is still a "small file", should be retained in the storage system.

[0076] 303. Update the status of the first storage space to readable.

[0077] After terminating the merge process, even if the first storage space is still a "small file," it needs to be directly saved in the storage system without copying its data. At this point, the status of the first storage space is directly updated to readable, allowing downstream applications to access its data. Understandably, once the next checkpoint interval begins, any newly written streaming data will be directly saved in the new storage space and no longer associated with the first storage space.

[0078] 304. Detect the amount of streaming data stored in the first storage space of the storage system.

[0079] However, if the target data is not in the aforementioned state—that is, if the data stored in the first storage space consists entirely of "new data" with very short write durations—a storage space merging step is required to avoid storing too many "small files" in the storage system. For example, it is necessary to obtain the amount of streaming data stored in the first storage space, and then determine whether to merge the first storage space with the storage space corresponding to the next checkpoint time interval based on the amount of data in the first storage space.

[0080] 305. Determine whether the amount of streaming data stored in the first storage space is less than the data amount threshold. If yes, proceed to step 306. If no, proceed to step 310.

[0081] After obtaining the amount of data stored in the first storage space, it needs to be compared with a preset data volume threshold. If the amount of data stored in the first storage space is greater than or equal to the threshold, the first storage space is considered a "large file." If the amount of data stored in the first storage space is less than the threshold, the first storage space is considered a "small file." Understandably, a "large file" indicates that the first storage space contains a sufficient amount of data, thus allowing it to be stored independently in the storage system. A "small file," on the other hand, can affect cluster scalability. Therefore, if the first storage space is considered a "small file," it needs to be merged with subsequent storage spaces to reduce the number of "small files" in the storage system.

[0082] For example, in Figure 1 In the system architecture scenario shown, the data volume threshold can be the amount of data that a data block can store in HDFS, typically 128MB. This means that the data volume of a file stored in HDFS must occupy at least one block to prevent the generation of too many "small files." Therefore, if the amount of data stored in a file is less than 128MB after a read / write cycle, the file is determined to be a "small file," and it needs to be merged with the file corresponding to the next read / write cycle. If it is greater than or equal to 128MB, the file can be stored normally.

[0083] 306. Copy the streaming data stored in the first storage space to the second storage space.

[0084] The specific merging method is as follows: before receiving streaming data at the next checkpoint interval, the streaming data in the first storage space corresponding to the previous checkpoint interval is copied to the second storage space corresponding to the next checkpoint interval. Then, newly written streaming data is stored in the second storage space. Thus, the second storage space includes the data written in the previous checkpoint interval stored in the first storage space and the newly written data in the next checkpoint interval. This is equivalent to merging the storage spaces corresponding to two consecutive checkpoint intervals, merging the streaming data written in both checkpoint intervals into one storage space (the second storage space). This reduces the probability of "small files" being generated.

[0085] 307. Update the status of the first storage space to readable.

[0086] After copying the data from the first storage space, the first storage space still needs to be retained for the next checkpoint interval. This is because the second storage space is undergoing data writing during the next checkpoint interval and is therefore unreadable. Consequently, the data from the first storage space copied to the second storage space is invisible to downstream applications during the next checkpoint interval. Downstream applications should be able to access the data from the first storage space after the previous checkpoint interval ends, but if the first storage space is deleted immediately, downstream applications will have to wait until the next checkpoint interval to read the data, causing delays in data retrieval. Therefore, when a checkpoint arrives, even if the data from the previous checkpoint interval is copied to the storage space corresponding to the next checkpoint interval, the storage space corresponding to the previous checkpoint interval must be retained first. Furthermore, the status of this storage space needs to be updated to a readable state so that downstream applications can query the data from the storage space corresponding to the previous checkpoint interval promptly and effectively during the next checkpoint interval.

[0087] 308. Write the streaming data into the second storage space during the second checkpoint time interval.

[0088] Understandably, when the next checkpoint interval begins, the streaming engine needs to write the newly added streaming data to the second storage space. That is, it stops writing streaming data to the first storage space (where the state has changed) and instead begins writing data to the second storage space. At this time, the second storage space will contain data written during both the first and second checkpoint intervals. Understandably, when the second checkpoint arrives, the data writing process to the second storage space will also stop.

[0089] 309. When the second checkpoint arrives, delete the first storage space.

[0090] After the second checkpoint interval ends, the state of the second storage space will be changed to a readable state. The data corresponding to the first checkpoint interval (streaming data in the first storage space) and the data written during the second checkpoint interval, both contained in the second storage space, will be visible to downstream applications. Therefore, downstream applications querying the second storage space will also retrieve the data from the first storage space. At this point, the first storage space needs to be deleted. In this way, the storage system only needs to store the second storage space, reducing the number of "small files" in the storage system and improving its management performance. Downstream applications also no longer need to frequently switch between different storage spaces when querying data, improving their query efficiency.

[0091] Understandably, if the amount of data in the second storage space still needs to be evaluated after the second checkpoint interval, and if the total amount of data in the second storage space is still less than the data volume threshold (meaning the second storage space containing data from multiple periods is still considered a "small file"), then all its data will be copied to the storage space corresponding to the next checkpoint interval, and merged again. Once the next checkpoint interval ends and the status of the storage space corresponding to that interval is updated to a readable state, the second storage space will be deleted, until a storage space becomes a "large file" after the data writing process is complete. In that case, the "large file" will be retained, and the data copying process will cease.

[0092] 310. Update the status of the first storage space to readable.

[0093] If the amount of data stored in the first storage space is greater than or equal to the data volume threshold, then the first storage space is considered a "large file." That is, the amount of data stored in the first storage space is sufficient, and merging is unnecessary. Since "large files" do not burden the storage system's management, the status of the first storage space can be directly updated to readable, and the first storage space can be saved in HDFS.

[0094] 311. During the second checkpoint time interval, stream data is written to the second storage space.

[0095] Finally, when the next cycle begins, the streaming engine needs to write the new streaming data to the new second storage space. That is, it stops writing streaming data to the first storage space where the state has changed, and starts writing data to the second storage space.

[0096] In this embodiment, the decision to merge storage spaces is based on the data writing status within those spaces. If data in a storage space has been copied multiple times or the writing time is long, even if the storage space still contains "small files," there's no need to copy the data again; instead, the "small files" can be saved in the storage system promptly. This conserves storage resources. Conversely, if data in a storage space hasn't been copied multiple times or the writing time is short, then file merging is necessary based on the amount of data stored in that space, reducing the number of "small files" in the storage system. This improves management efficiency while ensuring data readability. During the data writing process, the storage system can automatically merge "small files," significantly reducing their number and improving file management performance.

[0097] Based on the above method embodiments, Figure 4This is a schematic diagram of the structure of a streaming engine provided in an embodiment of this application, as shown below. Figure 4 As shown, this streaming processing engine processes streaming data and writes the processed streaming data to the storage system according to the checkpoint interval of the streaming processing engine. The streaming data written within different checkpoint intervals is stored in different storage spaces of the storage system. The streaming processing engine includes:

[0098] The detection unit 401 is used to detect the amount of streaming data stored in the first storage space of the storage system when the first checkpoint arrives. The first storage space is used to store the streaming data written by the streaming processing engine during the first checkpoint time interval. The first checkpoint is the end time of the first checkpoint time interval.

[0099] Processing unit 402 is configured to copy the streaming data stored in the first storage space to a second storage space in the storage system when the amount of streaming data stored in the first storage space is less than a data volume threshold. The second storage space is used to store the streaming data written by the streaming processing engine within a second checkpoint time interval. The second checkpoint time interval is the next checkpoint time interval adjacent to the first checkpoint time interval.

[0100] The processing unit 402 is also used to update the state of the first storage space to a readable state.

[0101] In an optional implementation, processing unit 402 is further configured to delete the first storage space when the second checkpoint arrives. The second checkpoint is the end time of the second checkpoint time interval. The state of the second storage space is updated to a readable state.

[0102] In an alternative implementation, the streaming engine also includes a storage unit 403.

[0103] The processing unit 402 is further configured to directly update the state of the first storage space to a readable state when the amount of streaming data stored in the first storage space is greater than or equal to the data amount threshold.

[0104] Storage unit 403 is also used to store newly written streaming data from the streaming engine in the second storage space, starting from the first checkpoint.

[0105] In an optional implementation, processing unit 402 is further configured to retain the first storage space when the second checkpoint arrives. The second checkpoint is the end time of the second checkpoint time interval. The state of the second storage space is updated to a readable state.

[0106] In an optional implementation, the streaming engine further includes a determination unit 404.

[0107] The determining unit 404 is used to determine the target data in the streaming data stored in the first storage space when the first checkpoint arrives. The target data is streaming data copied from the historical storage space corresponding to the historical checkpoint time interval.

[0108] The processing unit 402 is also configured to not copy the streaming data in the first storage space to the second storage space when the state of the target data is a preset state.

[0109] In one optional implementation, the processing unit 402 is specifically used to query data storage records and obtain the maximum number of copies corresponding to the target data. When the maximum number of copies corresponding to the target data reaches a first preset number, all streaming data stored in the first storage space is not copied to the second storage space.

[0110] In one optional implementation, the processing unit 402 is specifically used to query data storage records and obtain the average number of copies corresponding to the target data. When the average number of copies corresponding to the target data reaches a second preset number, all streaming data stored in the first storage space is not copied to the second storage space.

[0111] In an optional implementation, the processing unit 402 is specifically used to query the data storage record and obtain the initial write time corresponding to the target data.

[0112] The determination unit 404 is also used to determine the writing duration corresponding to the target data based on the initial writing time and the first checkpoint.

[0113] The processing unit 402 is specifically used to prevent copying all streaming data stored in the first storage space to the second storage space when the writing time reaches a preset time.

[0114] In an optional implementation, the processing unit 402 is further configured to directly update the state of the first storage space to a readable state at the first checkpoint, and store the streaming data newly written by the streaming engine in the second storage space.

[0115] In an optional implementation, the determining unit 404 is further configured to obtain the data volume corresponding to multiple historical storage spaces stored in the storage system. Based on the data volume corresponding to the multiple historical storage spaces, an average data volume is determined. A data volume threshold is then determined based on the average data volume.

[0116] The technical solution provided in this application embodiment, when the current checkpoint arrives, first checks whether the amount of streaming data stored in the storage space corresponding to the checkpoint time interval is less than the data volume threshold. If it is less, the streaming data in the storage space needs to be copied to the storage space corresponding to the next checkpoint time interval. In this way, the storage space corresponding to the next checkpoint time interval will store streaming data corresponding to two time intervals, thereby achieving storage space merging and reducing the number of "small files" in the storage system. Simultaneously, it is also necessary to temporarily retain the "small files," updating their status to readable at each checkpoint. This ensures that the data in the "small files" can be read by downstream applications within the next checkpoint time interval, avoiding data read latency. Once the next checkpoint arrives and the streaming data writing process in the storage space corresponding to the next checkpoint time interval is completed, the "small files" are deleted. However, if the amount of streaming data stored in the storage space corresponding to the current checkpoint time interval is greater than or equal to the data volume threshold, it indicates that the amount of data written in that checkpoint time interval is sufficient, and the storage space is also a "large file," which will not hinder storage space management. Therefore, the storage space can be retained. In this way, during the streaming data writing process, the storage system can automatically merge "small files", which will greatly reduce the number of "small files" in the storage system and thus improve the management performance of the storage system.

[0117] It should be noted that the division of the various modules in the above device is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, these modules can be implemented entirely in software via processing element calls; they can be fully implemented in hardware; or some modules can be implemented in software via processing element calls, while others are implemented in hardware. Additionally, these modules can be integrated together or implemented independently. The processing element mentioned here can be an integrated circuit with signal processing capabilities. During implementation, each step of the above method or each of the above modules can be completed through integrated logic circuits in the hardware of the processor element or through software instructions.

[0118] The following describes an electronic device provided by an embodiment of this application. Please refer to [link / reference]. Figure 5 , Figure 5 This is a schematic diagram of an electronic device provided in an embodiment of this application. The electronic device 800 may be equipped with... Figure 4 The streaming engine described in the corresponding embodiment is used to implement Figures 1 to 3The functions correspond to those in the embodiments. Specifically, the electronic device 800 includes: a receiver 801, a transmitter 802, a processor 803, and a memory 804 (wherein the number of processors 803 in the execution device 800 can be one or more). Figure 5 (Taking a processor as an example), the processor 803 may include an application processor 8031 ​​and a communication processor 8032. In some embodiments of this application, the receiver 801, transmitter 802, processor 803, and memory 804 may be connected via a bus or other means.

[0119] Memory 804 may include read-only memory and random access memory, and provides instructions and data to processor 803. A portion of memory 804 may also include non-volatile random access memory (NVRAM). Memory 804 stores processor and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations.

[0120] The processor 803 controls the operation of the execution device. In specific applications, the various components of the execution device are coupled together through a bus system, which may include not only the data bus, but also power buses, control buses, and status signal buses. However, for clarity, all buses in the diagram are referred to as the bus system.

[0121] The methods disclosed in the embodiments of this application can be applied to or implemented by processor 803. Processor 803 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above methods can be completed by integrated logic circuits in the hardware of processor 803 or by instructions in software form. Processor 803 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and may further include application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. Processor 803 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 804, and processor 803 reads the information from memory 804 and, in conjunction with its hardware, completes the steps of the above method.

[0122] Receiver 801 can be used to receive input digital or character information, and to generate signal inputs related to the settings and function control of the execution device. Transmitter 802 can be used to output digital or character information through the first interface; transmitter 802 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; transmitter 802 may also include a display device such as a display screen.

[0123] In this embodiment of the application application, the application processor 8031 ​​in the processor 803 is used to execute... Figures 1 to 3 The corresponding embodiment describes a method for processing streaming data. It should be noted that the specific manner in which the application processor 8031 ​​executes each step differs from that in this application. Figures 1 to 3 The various method embodiments are based on the same concept, and the technical effects they bring are the same as those in this application. Figures 1 to 3 The corresponding method embodiments are the same, and for details, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.

[0124] This application also provides a chip for executing instructions, which is used to execute the technical solution of the streaming data processing method in the above embodiments.

[0125] This application also provides a computer-readable storage medium storing computer instructions. When these computer instructions are executed on a server, the server performs the technical solution of the streaming data processing method described in the above embodiments.

[0126] This application also provides a computer program product, including a computer program, which, when executed by a processor, is used to perform the technical solution of the file data storage method described in the above embodiments.

[0127] The aforementioned computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to general-purpose or special-purpose servers.

[0128] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

[0129] Although this application discloses preferred embodiments as described above, it is not intended to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of this application. Therefore, the scope of protection of this application should be determined by the scope defined in the claims of this application.

Claims

1. A method for processing streaming data, characterized in that, The method includes: Streaming data is processed by a streaming processing engine, and the processed streaming data is written to a storage system according to the checkpoint time interval of the streaming processing engine; wherein, the streaming data written within different checkpoint time intervals is stored in different storage spaces of the storage system. When the first checkpoint arrives, the amount of streaming data stored in the first storage space of the storage system is detected; wherein, the first storage space is used to store the streaming data written by the streaming processing engine within the first checkpoint time interval; the first checkpoint is the end time of the first checkpoint time interval. When the amount of streaming data stored in the first storage space is less than the data amount threshold, the streaming data stored in the first storage space is copied to the second storage space of the storage system; the second storage space is used to store the streaming data written by the streaming processing engine within the second checkpoint time interval; the second checkpoint time interval is the next checkpoint time interval adjacent to the first checkpoint time interval; Update the status of the first storage space to readable.

2. The method according to claim 1, characterized in that, After copying the streaming data stored in the first storage space to the second storage space of the storage system, the method further includes: When the second checkpoint arrives, the first storage space is deleted; the second checkpoint is the end time of the second checkpoint time interval. Update the status of the second storage space to readable.

3. The method according to claim 1, characterized in that, The method further includes: When the amount of streaming data stored in the first storage space is greater than or equal to the data amount threshold, the state of the first storage space is directly updated to the readable state. Starting from the first checkpoint, the newly written streaming data by the streaming engine is stored in the second storage space.

4. The method according to claim 3, characterized in that, The method further includes: When the second checkpoint arrives, the first storage space is retained; the second checkpoint is the end time of the second checkpoint time interval. Update the status of the second storage space to readable.

5. The method according to any one of claims 1 to 4, characterized in that, Before detecting the amount of streaming data stored in the first storage space of the storage system, the method further includes: When the first checkpoint arrives, the target data in the streaming data stored in the first storage space is determined; the target data is streaming data copied from the historical storage space corresponding to the historical checkpoint time interval. When the target data is in a preset state, the streaming data in the first storage space will not be copied to the second storage space.

6. The method according to claim 5, characterized in that, The step of not copying streaming data from the first storage space to the second storage space when the target data is in a preset state includes: Query the data storage records to obtain the maximum number of copies corresponding to the target data; When the maximum number of copies corresponding to the target data reaches the first preset number, then all streaming data stored in the first storage space will not be copied to the second storage space.

7. The method according to claim 5, characterized in that, The step of not copying streaming data from the first storage space to the second storage space when the target data is in a preset state includes: Query the data storage records to obtain the average number of copies corresponding to the target data; When the average number of copies corresponding to the target data reaches the second preset number, then all streaming data stored in the first storage space will not be copied to the second storage space.

8. The method according to claim 5, characterized in that, The step of not copying streaming data from the first storage space to the second storage space when the target data is in a preset state includes: Query the data storage record to obtain the initial write time corresponding to the target data; Based on the initial write time and the first checkpoint, determine the write duration corresponding to the target data; When the written duration reaches the preset duration, all streaming data stored in the first storage space will not be copied to the second storage space.

9. The method according to claim 1, characterized in that, The method further includes: At the first checkpoint, the state of the first storage space is directly updated to the readable state, and the newly written streaming data by the streaming engine is stored in the second storage space.

10. The method according to claim 9, characterized in that, The method further includes: Obtain the amount of data corresponding to multiple historical storage spaces stored in the storage system; Determine the average data volume based on the data volume corresponding to the multiple historical storage spaces; The data volume threshold is determined based on the average data volume.

11. A streaming processing engine, characterized in that, The streaming processing engine processes streaming data and writes the processed streaming data into the storage system according to the checkpoint time interval of the streaming processing engine; wherein, the streaming data written within different checkpoint time intervals is stored in different storage spaces of the storage system; the streaming processing engine includes: A detection unit is used to detect the amount of streaming data stored in the first storage space of the storage system when the first checkpoint arrives; wherein, the first storage space is used to store the streaming data written by the streaming processing engine within the first checkpoint time interval; the first checkpoint is the end time of the first checkpoint time interval. The processing unit is configured to copy the streaming data stored in the first storage space to the second storage space of the storage system when the amount of streaming data stored in the first storage space is less than the data amount threshold; the second storage space is used to store the streaming data written by the streaming processing engine within the second checkpoint time interval; the second checkpoint time interval is the next checkpoint time interval adjacent to the first checkpoint time interval. The processing unit is further configured to update the state of the first storage space to a readable state.

12. An electronic device, characterized in that, include: Processor, memory, and computer program instructions stored in said memory and executable on the processor; When the processor executes the computer program instructions, it implements the streaming data processing method as described in any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the streaming data processing method as described in any one of claims 1 to 10.

Citation Information

Patent Citations

  • Opportunistic use of streams for storing data on a solid state device

    CN110312986A

  • Streaming data distribution method and system

    CN113612832A