File merging method and device, equipment and medium
By merging the result data sets in offline computing applications of e-commerce platforms, the high I/O overhead and database connection pressure caused by a large number of small files is solved, and more efficient data reading and storage system stability is achieved.
Patent Information
- Application Number
- CN202510147496.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-06-03
AI Technical Summary
In offline computing applications in the e-commerce field, the generated result data sets are usually stored in the form of a large number of small files, resulting in frequent file system operations, reducing data reading efficiency, and creating high connection pressure on the database, which may cause service crashes.
By capturing the result data set to be stored in the data storage system, the average storage size of a single partition is calculated, and whether to perform file merging operations based on the preset size threshold is determined to adjust the number of partitions and storage size, and the merged result data set is stored in the data storage system.
It significantly reduces the I/O overhead during file reading, improves data retrieval and loading efficiency, reduces the connection pressure of the database, ensures the stability of the data storage system, and significantly reduces the burden on the data storage system during peak hours of e-commerce platforms.
Smart Images

Figure CN120085804A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of e-commerce technologies, and in particular, to a file merging method, apparatus, device, and medium thereof. Background Art
[0002] In the e-commerce field, with the rapid growth of data scales such as user data, commodity data, and transaction data, the processing requirements for massive data are continuously increasing. To cope with these requirements, e-commerce platforms usually adopt an offline computing framework to batch process this massive data and generate corresponding result data sets. Through distributed computing and storage technologies, the offline computing framework can efficiently complete large-scale data analysis and processing tasks.
[0003] However, in actual offline computing applications, there are still some limitations. First of all, the result data sets generated by the offline computing framework are usually stored in a partitioned form, and each partition corresponds to an independent file. Due to the high parallelism of computing tasks and uneven data distribution, the computing results often appear in the form of a large number of small files. If these small files need to be read subsequently, it is necessary to frequently open files, search for metadata, and perform I / O operations, thereby incurring high time costs, reducing the data reading efficiency, and seriously hindering the improvement of overall performance.
[0004] Secondly, each file of the result data set establishes a connection with the database to write the result data set into the database with high parallelism for offline computing tasks. However, the database usually has certain restrictions on the number of simultaneous connections. Excessive connection requests are likely to overload the database, resulting in response delays or even service crashes. Especially during peak hours on e-commerce platforms, a large number of write requests may cause huge pressure on the database, thereby affecting the stability and availability of the entire database.
[0005] Therefore, how to effectively reduce the number of files stored in the data storage system has become a key problem that urgently needs to be solved in the e-commerce field. Summary of the Invention
[0006] The primary objective of this application is to solve at least one of the above problems and provide a file merging method, apparatus, device, and medium thereof.
[0007] To meet the various objectives of this application, the following technical solutions are adopted in this application:
[0008] A file merging method provided to meet one of the objectives of this application includes the following steps:
[0009] Capture a result data set composed of files in multiple partitions to be stored in a data storage system, and obtain the number of partitions and the total storage size corresponding to the result data set;
[0010] Determine the average storage size of a single partition according to the number of partitions and the total storage size;
[0011] When the average storage size is less than a preset size threshold, perform a file merging operation on the result data set to adjust the number of partitions and the storage size of each partition, so that the result data set contains files corresponding to multiple partitions, and there is no more than one file with a storage size less than the size threshold in each partition, and store the merged result data set in the data storage system;
[0012] When the average storage size is greater than a preset size threshold, store the result data set in the data storage system.
[0013] On the other hand, a file merging device provided to meet one of the purposes of the present application includes:
[0014] A data set capture module, configured to capture a result data set composed of files in multiple partitions to be stored in a data storage system, and obtain the number of partitions and the total storage size corresponding to the result data set;
[0015] An average storage calculation module, configured to determine the average storage size of a single partition according to the number of partitions and the total storage size;
[0016] A file merging module, configured to, when the average storage size is less than a preset size threshold, perform a file merging operation on the result data set to adjust the number of partitions and the storage size of each partition, so that the result data set contains files corresponding to multiple partitions, and there is no more than one file with a storage size less than the size threshold in each partition, and store the merged result data set in the data storage system;
[0017] A direct storage module, configured to store the result data set in the data storage system when the average storage size is greater than a preset size threshold.
[0018] On the other hand, a computer device provided to meet one of the purposes of the present application includes a central processing unit and a memory, and the central processing unit is configured to call and run a computer program stored in the memory to execute the steps of the file merging method described in the present application.
[0019] On the other hand, a computer program product provided to meet another purpose of the present application includes computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the method described in any embodiment of the present application are implemented.
[0020] The technical solution of the present application has many advantages, including but not limited to the following aspects:
[0021] This application optimizes the storage structure of the data storage system by capturing the result data set to be stored in the data storage system, calculating the average storage size of a single partition, and determining whether to perform a file merging operation according to a preset size threshold. When there are too many small files, the small files in the result data set are merged to reduce the number of small files.
[0022] Through the merging operation of small files, this application significantly reduces the I / O overhead during file reading, avoids frequent file system operations, thereby greatly improving the data retrieval and loading efficiency, making the overall data processing process smoother and faster. Secondly, by reducing the number of files, the connection pressure on the database is effectively reduced, avoiding problems such as the exhaustion of the connection pool of the data storage system or service crashes caused by high-concurrency writes, ensuring the stability of the data storage system and the stable operation of the business system. Especially during the peak hours of the e-commerce platform, this application can significantly reduce the burden on the data storage system and improve the overall performance of the system. In addition, eliminating the small file problem can also optimize the utilization efficiency of storage resources, reduce the storage cost on the cloud, and further improve the resource utilization rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The above and / or additional aspects and advantages of this application will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0024] Figure 1 is a flowchart of a typical embodiment of the file merging method of this application;
[0025] Figure 2 is a flowchart of the process of adjusting the size threshold in the embodiment of this application;
[0026] Figure 3 is a flowchart of the specific process of adjusting the size threshold in the embodiment of this application;
[0027] Figure 4 is a flowchart of the process of storing the result data set in a distributed file system in the embodiment of this application;
[0028] Figure 5 is a flowchart of the process of storing the result data set in a database in the embodiment of this application;
[0029] Figure 6 is a flowchart of the process of storing the associated partition quantity and the total storage size in the embodiment of this application;
[0030] Figure 7 is a flowchart of the process of verifying the result data set before and after merging in the embodiment of this application;
[0031] Figure 8 is a schematic block diagram of the principle of the file merging device of this application;
[0032] Figure 9 It is a schematic structural diagram of a computer device adopted by this application. Specific implementation manners
[0033] A file merging method of this application can be programmed as a computer program product and deployed to run in a server. For example, in an exemplary application scenario of this application, it can be deployed in the server of an e-commerce platform. Among them, the e-commerce platform can be an e-commerce platform that provides open independent station services. An independent station refers to a new type of official website (website) based on the SaaS technology platform, with an independent domain name, private content, data, and rights, having independent operation sovereignty and operation entity responsibility, supported by social cloud computing capabilities, and capable of independently and freely docking third-party software tools, publicity and promotion media, and channels. In the business scenario of an e-commerce platform, the e-commerce platform usually needs to execute a large number of offline tasks. These tasks usually rely on an offline computing framework (such as Apache Spark) to perform distributed computing on massive data and generate corresponding result data sets. Due to the characteristics of distributed computing, the generated result data sets often consist of files in multiple partitions, and there may be problems with uneven storage sizes among these files. Especially in the case of too many small files, it will seriously affect the performance and query efficiency of the data storage system. The data storage system includes but is not limited to a distributed file system and a database.
[0034] The file merging method of this application is used to process the result data set generated by the offline computing framework. By performing file merging operations, the data storage structure is optimized to improve the performance of the data storage system. In the corresponding embodiment, when the offline computing framework completes distributed computing and generates a result data set, the file merging method of this application will capture the result data set and obtain the corresponding number of partitions and the total storage size. Among them, the number of partitions reflects the number of parallel task units into which the data set is divided during the distributed computing process, which is usually determined by the parallelism of the computing task, the data sharding rule, or user configuration parameters, while the total storage size represents the storage space occupied by the entire result data set. Then, according to the captured number of partitions and the total storage size, calculate the average storage size of a single partition, that is, the average size of a single file (in the embodiments of this application, a partition is regarded as a file), and compare the average storage size with a preset size threshold. If the average storage size is less than the preset size threshold, it indicates that there may be a large number of small files in the current result data set, and file merging operations need to be performed. The merged result data set is stored in the data storage system. In some embodiments, the data storage system is the distributed file system HDFS. In some embodiments, the data storage system is a database, which does not affect the embodiment of the creative spirit of this application.
[0035] A file merging method of the present application can be programmed as a computer program product and implemented by running on a server. For example, in an exemplary application scenario of the present application, it can be implemented by deploying on the server of an e-commerce customer service platform.
[0036] Please refer to Figure 1 , in a typical embodiment of the file merging method of the present application, the following steps are included:
[0037] Step S5100: Capture a result data set composed of files in multiple partitions to be stored in a data storage system, and obtain the number of partitions corresponding to the result data set and the total storage size;
[0038] In the business scenario of an e-commerce platform, the e-commerce platform needs to execute a large number of offline tasks, such as commodity data analysis, user behavior analysis, order statistics, etc. These offline tasks usually rely on an offline computing framework (such as Apache Spark, Apache Hive) to achieve efficient data processing. The offline computing framework divides the input data into multiple partitions and executes computing tasks in parallel on multiple computing nodes, and finally generates a result data set composed of files in multiple partitions. Each partition file contains partial computing result data. In one embodiment, the result data set can be stored in a temporary storage area associated with the number of partitions and the total storage size. When implementing the present application, the number of partitions and the total storage size of the result data set can be directly obtained. In another embodiment, the partition files of the result data set can be accessed through the API of the distributed file system to obtain its metadata information (such as the number of partitions and the size of each partition file), and the sizes of each partition file are accumulated to obtain the total storage size of the result data set. In this way, the number of partitions and the total storage size of the result data set can be dynamically captured.
[0039] In this step, intercept during the stage of writing the computed result data into the data storage system, and capture the result data set to be stored by calling the API of the offline computing framework or the file system interface. The result data set is composed of files in multiple partitions. The number of partitions and the total storage size of the result data set can be obtained by using the data operation tools provided by the offline computing framework. Among them, the total storage size can be obtained by obtaining the size of each partition file and accumulating the sizes of all partition files. The total storage size reflects the storage space occupied by the entire result data set and is an important indicator for evaluating whether the data storage structure is reasonable.
[0040] Step S5200: Determine the average storage size of a single partition according to the number of partitions and the total storage size;
[0041] In one embodiment, the total storage size is divided by the number of partitions to obtain the average storage size of a single partition. For example, if the total storage size is 100 GB and the number of partitions is 10, the average storage size of a single partition is 10 GB. In practical applications, the distribution of data across partitions may not be completely uniform. Some partitions may contain more data while others may contain less. Therefore, the average storage size is only a statistical metric that reflects the average amount of data in each partition file and is used to evaluate the rationality of the overall data storage structure.
[0042] In another embodiment, the average storage size of a single partition is determined by weighted average. First, the size of each partition file is obtained and its product with a preset weight value is calculated. The weight value can be dynamically adjusted according to the data distribution characteristics of the partition (such as data skew degree, partition access frequency, etc.). For example, a higher weight can be assigned to a partition with a larger amount of data, and a lower weight can be assigned to a partition with a smaller amount of data. Then, the weighted sizes of all partition files are accumulated to obtain the weighted total storage size. Finally, the weighted total storage size is divided by the number of partitions to obtain the weighted average storage size. The weighted average storage size can more accurately reflect the actual situation of data distribution and provide a more precise basis for subsequent file merging operations.
[0043] By calculating the average storage size, it is possible to quickly identify whether there is a problem of too many small files in the result dataset, providing a key basis for subsequent file merging operations, thereby optimizing the data storage structure.
[0044] Step S5300: When the average storage size is less than a preset size threshold, perform a file merging operation on the result dataset to adjust the number of partitions and the storage size of each partition, so that the result dataset contains files corresponding to multiple partitions, and there is no more than one file with a storage size less than the size threshold in each partition, and store the merged result dataset in the data storage system;
[0045] Compare the average storage size with a preset size threshold, where the size threshold is set by those skilled in the art as needed. It should be noted that in one embodiment, the size threshold is dynamically adjusted and not fixed. When it is determined that the average storage size is less than the preset size threshold, it can be determined that it is better to perform file merging adjustment on this result dataset before storing it in the data storage system.
[0046] Specifically, when the average storage size is less than a preset size threshold, the result data set is merged to adjust the number of partitions and the storage size of each partition. Finally, the merged result data set is stored in the data storage system. Among them, a distributed file system or a database can be used as the data storage system. Taking the distributed file system as an example, first, according to the preset size threshold and the total storage size of the result data set, calculate the number of files to be merged and the target size of each file after merging. Then traverse all the partition files in the result data set, and merge multiple small files into larger files according to the target size to reduce the number of small files and optimize the storage structure. During the merging process, ensure that the size of each merged file is as close as possible to the target size, while avoiding generating too many small files. Finally, write the merged result data set into the distributed file system (such as HDFS). For the distributed file system, the merged files are written to the specified directory through the file system interface; for the database, the merged data is stored in the specified table through the database connection. Through this step, ensure that there is no more than one file with a storage size less than the preset size threshold for each partition, thereby improving the performance and query efficiency of the data storage system.
[0047] Step S5400: When the average storage size is greater than the preset size threshold, store the result data set in the data storage system.
[0048] Compare the average storage size with the preset size threshold. If the average storage size is greater than the preset threshold, it indicates that the size distribution of the partition files in the result data set is relatively reasonable and no file merging operation is required. The partition files in the result data set can be directly stored in the data storage system as they are through the interface of the data storage system (such as the file writing interface of the distributed file system or the data insertion interface of the database). Through this step, ensure that the result data set is stored in the data storage system with the optimal storage structure, avoiding unnecessary file merging operations, thereby improving the efficiency of data storage and query.
[0049] According to the typical embodiments of the present application, it can be known that the technical solution of the present application has many advantages, including but not limited to the following aspects:
[0050] The present application captures the result data set to be stored in the data storage system, calculates the average storage size of a single partition, and determines whether to perform a file merging operation according to the preset size threshold, thereby optimizing the storage structure of the data storage system. When there are too many small files, the small files in the result data set are merged to reduce the number of small files.
[0051] Through the merging operation of small files, this application significantly reduces the I / O overhead during file reading, avoids frequent file system operations, thereby greatly improving the data retrieval and loading efficiency, and making the overall data processing process smoother and faster. Secondly, by reducing the number of files, the connection pressure on the database is effectively reduced, avoiding problems such as the exhaustion of the connection pool of the data storage system or service crashes caused by high-concurrency writes, and ensuring the stability of the data storage system and the stable operation of the business system. Especially during peak hours on e-commerce platforms, this application can significantly reduce the burden on the data storage system and improve the overall performance of the system. In addition, eliminating the small file problem can also optimize the utilization efficiency of storage resources, reduce the storage cost on the cloud, and further improve the resource utilization rate.
[0052] In a further embodiment, please refer to Figure 2 , when the average storage size is less than a preset size threshold, perform a file merging operation on the result data set to adjust the number of partitions and the storage size of each partition, so that the result data set contains files corresponding to multiple partitions, and there is no more than one file with a storage size less than the size threshold in each partition. Before storing the merged result data set into the data storage system, the following steps are included:
[0053] Step S6100: Traverse the existing files in the data storage system to obtain the number and total size of the existing files;
[0054] Access the existing files or data records in the data storage system. For a distributed file system (such as HDFS), use the API of the file system to traverse all files in the specified directory to obtain the metadata information of each file, including the file size and the number of files. For a database, retrieve the data records in the specified table through an SQL query statement and calculate the size or occupied space of each record. Then traverse all existing files or data records, count the number of files, and calculate the size of each file. For a distributed file system, directly obtain the file size information; for a database, calculate by querying the size or occupied space of the data records. Finally, accumulate the sizes of all files to obtain the total size of the existing files. Through this step, the number and total size of the existing files in the data storage system can be accurately obtained, providing basic data support for adjusting the preset size threshold in the following.
[0055] Step S6200: Adjust the preset size threshold according to the number and total size of the existing files for comparison with the average size.
[0056] Based on the quantity and total size of existing files, the current storage structure of the current data storage system can be quantified. When the number of existing files is excessive and the total size is large, it indicates that there may be a large number of small files stored in the data storage system, resulting in low storage efficiency and degraded query performance. At this time, the preset size threshold can be appropriately increased to reduce the number of small files and optimize the storage structure. For example, assume that the average size of each existing file in the data storage system is 64 MB, and the preset size threshold is 128 MB. Then, the threshold can be appropriately increased to 256 MB to reduce the number of small files and improve storage efficiency. On the other hand, if the number of existing files is small and the total size is small, it indicates that the file distribution in the storage system is relatively reasonable. The preset size threshold can be maintained or appropriately decreased to avoid excessive file merging resulting in overly large individual files, which may affect data processing and query performance. For example, if the average size of each existing file is 256 MB and the preset size threshold is 128 MB, the threshold can be appropriately decreased to 64 MB to ensure a more uniform file size distribution. By dynamically adjusting the preset size threshold, the file merging strategy can be optimized according to the actual situation of the data storage system, improving storage efficiency and query performance.
[0057] In one embodiment, the small file density can be calculated to quantify the current storage structure of the data storage system, and the preset size threshold can be adjusted according to the small file density. In addition, multi-dimensional metrics such as file access frequency, storage system load, file type, and historical data can be combined, which do not affect the manifestation of the creative spirit of this application. For example, for files with a high access frequency, even if their size is smaller than the preset threshold, they can be selected not to be merged to avoid frequent merging operations affecting system performance; for situations where the storage system load is high, the preset size threshold can be appropriately increased to reduce the frequency of file merging operations to reduce the pressure on the data storage system; for different types of data files (such as log files, index files, calculation result files, etc.), different preset size thresholds can be set to achieve differential processing; historical data analysis can be used to analyze the change trend of file size and quantity, predict the future file distribution, and dynamically adjust the preset size threshold based on the prediction results. By comprehensively considering multi-dimensional metrics, the file merging strategy can be more finely controlled, thereby improving the performance and resource utilization rate of the data storage system.
[0058] In this embodiment, it can dynamically adapt to the actual state of the data storage system. By analyzing the number and total size of existing files in real time and combining multi-dimensional metrics (such as small file density, file access frequency, storage system load, file type, and historical data, etc.), it can intelligently adjust the preset size threshold, thereby achieving fine-grained control of the file merging strategy. This dynamic adjustment mechanism can not only effectively reduce the number of small files, optimize the storage structure, improve storage efficiency and query performance, but also avoid the negative impact on the performance of the data storage system caused by overly large individual files or frequent merging operations due to excessive merging. In addition, by comprehensively considering multi-dimensional metrics, this embodiment can more accurately balance the load and performance requirements of the data storage system, adapt to the storage optimization requirements in different business scenarios, significantly improve the resource utilization rate and overall performance of the data storage system, and has higher flexibility and adaptability.
[0059] In a further embodiment, please refer to Figure 3 , and adjust the preset size threshold according to the number and total size of the existing files for comparison with the average storage size, including the following steps:
[0060] Step S6210: Divide the total size of the existing files by the number to obtain the average size of the existing files, calculate the ratio of the average size of the existing files to the preset size threshold, and determine it as the small file density;
[0061] Traverse all existing files in the data storage system to determine the total size and number of all files, and then divide the total size by the number of files to calculate the average size of each file. Further, compare the calculated average size with the preset size threshold, and this ratio is used to quantify the small file density. Specifically, whether dividing the average size by the preset size threshold or dividing the preset size threshold by the average size, the following embodiments are adaptively corresponding and do not affect the manifestation of the creative spirit of this application. The small file density reflects the distribution of small files in the current data storage system. The lower the density, the more small files there are, and the lower the storage efficiency may be; the higher the density, the closer the file size distribution is to the preset threshold, and the relatively reasonable the storage structure is. By calculating the small file density, it is possible to quantitatively evaluate the file distribution state of the current storage system, provide data support for dynamically adjusting the preset size threshold, thereby optimizing the file merging strategy and improving storage efficiency and query performance.
[0062] Step S6220: If the small file density is less than the preset density threshold, determine an adjustment coefficient based on the small file density and the preset first mapping relationship, and increase the size threshold according to the adjustment coefficient;
[0063] First, compare the calculated small file density with the preset density threshold, which is set by those skilled in the art as needed. If the calculated small file density is less than the density threshold, it indicates that there are many small files in the current data storage system and the storage efficiency is low. In this case, the preset size threshold needs to be adjusted. At this time, determine the adjustment coefficient based on the small file density and the preset first mapping relationship. The first mapping relationship can be a linear or non-linear functional relationship, such as a linear function, an exponential function, or other mathematical relationships, without affecting the manifestation of the inventive concept of this application. For example, the lower the small file density, the larger the adjustment coefficient; the higher the small file density, the smaller the adjustment coefficient. For example, if the small file density is 0.5 and the density threshold is 0.8, the adjustment coefficient is determined to be 1.5 according to the first mapping relationship. Then, multiply the adjustment coefficient by the currently preset size threshold to obtain the adjusted size threshold. For example, if the currently preset size threshold is 128 MB and the adjustment coefficient is 1.5, the adjusted size threshold is 192 MB. The first mapping relationship can be represented in the form of a table or a linear function. In this way, the size threshold can be dynamically adjusted according to the small file density, so as to optimize the file merging strategy, reduce the number of small files, and improve the storage efficiency and query performance.
[0064] Step S6230: If the small file density is greater than the preset density threshold, determine the adjustment coefficient based on the small file density and the preset second mapping relationship, and perform a reduction adjustment on the size threshold according to the adjustment coefficient.
[0065] Compare the calculated small file density with the preset density threshold. If the small file density is greater than the density threshold, it indicates that the file distribution in the current data storage system is relatively reasonable, the number of small files is small, and the storage efficiency is high. At this time, the preset size threshold can be appropriately reduced to avoid excessive file merging resulting in an overly large single file and affecting data processing and query performance. Determine the adjustment coefficient based on the small file density and the preset second mapping relationship. The second mapping relationship can be a linear or non-linear functional relationship. For example, the higher the small file density, the smaller the adjustment coefficient; the lower the small file density, the larger the adjustment coefficient. For example, if the small file density is 0.9 and the density threshold is 0.8, the adjustment coefficient is determined to be 0.8 according to the second mapping relationship. Then, multiply the adjustment coefficient by the currently preset size threshold to obtain the adjusted size threshold. In this way, the size threshold can be dynamically adjusted according to the small file density, so as to optimize the file merging strategy, ensure a more uniform file size distribution, and improve the storage efficiency and query performance.
[0066] In this embodiment, it is possible to adjust the preset size threshold by dynamically calculating the small file density and combining the preset mapping relationship, so as to achieve refined control of the file merging strategy. By analyzing the ratio of the average size of existing files to the preset threshold in real time, the distribution state of small files in the current storage system is quantitatively evaluated, and the size threshold is flexibly adjusted according to the comparison result between the small file density and the density threshold. When the small file density is low, the threshold is increased to reduce the number of small files and optimize the storage structure; when the small file density is high, the threshold is decreased to avoid excessive merging resulting in an overly large single file and ensure a more uniform distribution of file sizes. This dynamic adjustment mechanism can not only effectively improve storage efficiency and query performance, but also avoid waste of storage resources or degradation of system performance caused by unreasonable setting of the size threshold, significantly enhancing the adaptability and flexibility of the file merging strategy and meeting the storage optimization requirements in different business scenarios.
[0067] In a further embodiment, please refer to Figure 4 , when the average storage size is less than the preset size threshold, perform a file merging operation on the result data set to adjust the number of partitions and the storage size of each partition, so that the result data set contains files corresponding to multiple partitions, and there is no more than one file with a storage size less than the size threshold in each partition. Store the merged result data set in the data storage system, including the following steps:
[0068] Step S5310: When the average storage size is less than the preset size threshold, start the read-write data process, continuously read the data in the result data set, accumulate the size of the data read each time, and obtain the cumulative read data volume;
[0069] If the average storage size is less than the preset size threshold, it indicates that there are a large number of small files in the result data set and a file merging operation is required. Specifically, when the average storage size is less than the preset size threshold, start the data reading process and read the data in the partition files in the result data set one by one. Each time data is read, record the amount of data read and add it to the cumulative read data volume, so as to monitor the cumulative read data volume in real time, provide data support for subsequent file merging operations, ensure that the size of the merged file is equal to the preset size threshold, thereby optimizing the storage structure, improving storage efficiency and query performance.
[0070] In one embodiment, start the read-write data process, read the data in the partition files in the result data set one by one through the I / O stream. Assume that the size of each data block read is 1024 bytes (1KB), and record the amount of data read each time. During the process of reading data, accumulate the amount of data read each time in real time to obtain the cumulative read data volume.
[0071] Step S5320: When the cumulative read data volume reaches the size threshold, merge and write the respective read data into the same file, and reset the cumulative read data volume to zero;
[0072] During the data reading process, the data volume read each time is accumulated in real time. When the cumulative read data volume reaches a preset size threshold (such as 128MB), multiple currently read data blocks are merged and written into a temporary file. For example, if the cumulative read data volume reaches 128MB, then this 128MB of data is written into a temporary file through an I / O stream to ensure that the file size is consistent with the preset threshold. After writing is completed, the cumulative read data volume is reset to zero for subsequent continuous data reading. In this way, the size of each merged file can be precisely controlled to ensure that the size of each file is equal to the preset size threshold.
[0073] Step S5330: Iterate the above data reading and writing process until the result data set is empty and the corresponding cumulative read data volume is less than the size threshold, then terminate the data reading and writing process;
[0074] In the previous step, after completing a file merging operation, the cumulative read data volume is reset to zero. In this step, the remaining data in the result data set is continuously read. Each time data is read, the read data volume is accumulated until the cumulative read data volume reaches the preset size threshold again to generate the next merged file. Repeat the above data reading and writing process until all the data in the result data set is read. If the remaining cumulative read data volume is less than the preset size threshold at the end, then this part of the data is written into the last merged file. For example, if the remaining cumulative read data volume is 50MB and the preset size threshold is 128MB, then this 50MB of data is written into the last file. In this way, it can be ensured that all data is merged into files with sizes close to the preset threshold.
[0075] It should be noted that the size of the last file may exactly equal the preset size threshold. For example, if the remaining cumulative read data volume is exactly 128MB and the preset size threshold is 128MB, then this 128MB of data is written into the last file to make its size exactly the same as the preset threshold. In this way, it can be ensured that all data is merged into files with sizes close to or equal to the preset threshold, thereby optimizing the storage structure and improving the storage efficiency and query performance.
[0076] Step S5340: Write the merged file into a temporary directory, rename the temporary directory, and store it in the distributed file system.
[0077] In one embodiment, when writing the result data set to the distributed file system through the offline computing framework, the overall data is not readable until all the result data set, that is, all files are written. However, the file merging method of the present application is independent of the offline computing framework and is carried out through file I / O. In this case, if the conventional writing method is used, when reading data during the writing process, the result path or the parent directory may be read, resulting in the inability to read the complete result data set. Therefore, it is necessary to introduce a temporary directory to store the result data set. After all the result data set is written, the temporary directory is renamed and then stored in the distributed file system, which can solve the problem of not being able to read the complete result data set mentioned above.
[0078] Write the merged result data set to a temporary directory, which can be a temporary path in the distributed file system (such as HDFS). During the writing process, ensure that all the merged files are correctly stored according to the target size and the number of files, and retain the metadata information of the files (such as partition information, file size, etc.). After the writing is completed, perform a rename operation on the temporary directory and move it to the target storage path in the distributed file system. The rename operation is atomic, which can ensure the consistency and integrity of the data and avoid data loss or damage caused by system failures or interruptions during the writing process. In this way, the merged result data set can be safely and efficiently stored in the distributed file system, ensuring the optimization of the data storage structure and the improvement of the query performance.
[0079] In this embodiment, by monitoring the cumulative read data volume in real time and dynamically adjusting the file merging operation, the size of the merged file can be accurately controlled to ensure that the size of each file is equal to the preset size threshold, thus effectively avoiding the problem of too many small files in the storage system. At the same time, by introducing a temporary directory and an atomic rename operation, the consistency and integrity of the data writing process are ensured, and data loss or damage caused by system failures or interruptions is avoided. In addition, this method is independent of the offline computing framework and is implemented through file I / O, which can maintain the readability and availability of the data during the writing process, significantly improving the storage efficiency and query performance.
[0080] In a further embodiment, please refer to Figure 5 , when the average storage size is less than the preset size threshold, perform a file merging operation on the result data set to adjust the number of partitions and the storage size of each partition, so that the result data set contains files corresponding to multiple partitions, and the number of files with a storage size less than the size threshold in each partition does not exceed one. Store the merged result data set in the data storage system, including the following steps:
[0081] Step S5340: When the average storage size is less than the preset size threshold, determine the target number of partitions according to the ratio of the total storage size to the size threshold using the ceiling operation.
[0082] Divide the total storage size of the result dataset by the preset size threshold to obtain a preliminary ratio of the number of partitions. For example, if the total storage size is 80 MB and the preset size threshold is 64 MB, the ratio is 80 MB / 64 MB ≈ 1.25. Then, process this ratio using the ceiling operation to ensure that the number of partitions is an integer and can cover all the data. For example, performing the ceiling operation on 1.25 gives a target number of partitions of 2. In this way, it can be ensured that the storage size of each partition is as close as possible to the preset size threshold, while avoiding generating too many small files, thereby optimizing the storage structure and improving the performance and query efficiency of the data storage system.
[0083] Step S5350: Re-partition the result dataset based on the target number of partitions to obtain the re-partitioned result dataset.
[0084] According to the determined target number of partitions, perform a re-partitioning operation on the result dataset. In one embodiment, by invoking the partitioning tool provided by an offline computing framework (such as Apache Spark or Apache Hive), the original partitioned files are re-divided according to the target number of partitions to ensure that the storage size of each new partition is as close as possible to the preset size threshold. During the re-partitioning process, optimize according to the distribution characteristics of the data (such as data skew degree, partitioning keys, etc.) to ensure that the data is evenly distributed among the partitions and avoid problems such as data skew or partitions being too large or too small. For example, if the target number of partitions is 782, the original partitioned files are re-divided into 782 new partitions, and the storage size of each partition is close to 128 MB. In this way, the number of small files can be effectively reduced, ensuring that the re-partitioned result dataset meets the preset storage requirements, thereby improving the performance and query efficiency of the data storage system.
[0085] Step S5360: Store the re-partitioned result dataset into the specified table of the corresponding database through a database connection.
[0086] First, establish a connection to the target database. Using the API or data writing interface provided by the database (such as JDBC, ODBC, etc.), write the result data set after re-partitioning into the specified table of the database partition by partition. During the writing process, ensure that the data in each partition is correctly mapped according to the field structure and data type of the database table, and retain the necessary metadata information (such as partition keys, file sizes, etc.). For example, if using JDBC to connect to a MySQL database, the data in each partition can be efficiently written into the target table through batch insertion. After writing is completed, close the database connection and verify the integrity and consistency of the data to ensure that all data is successfully stored without loss or damage. In this way, the result data set after re-partitioning can be efficiently and securely stored in the database, ensuring the optimization of the data storage structure and the improvement of query performance.
[0087] In a further embodiment, refer to Figure 6 , before capturing the result data set composed of files in multiple partitions to be stored in the data storage system and obtaining the number of partitions and the total storage size corresponding to the result data set, the following steps are included:
[0088] Step S7100: Respond to the product query request and recall the result data set that matches the query conditions from the product database;
[0089] In an offline computing framework (such as Apache Spark or Apache Hive), when a product query request is received, first parse the query conditions (such as product category, price range, keywords, etc.), and then generate corresponding query statements or call the database query interface according to the parsed conditions. By executing the query statements or calling the interface, retrieve the product data that matches the query conditions from the product database to form a result data set. The recalled result data set contains multiple product records, and each record contains information such as product ID, name, price, description, etc. In this way, the result data set that matches the query conditions can be efficiently extracted from the product database.
[0090] Step S7200: According to the preset sharding rules and the parallelism of the computing task, divide the result data set into multiple partitions and determine the number of partitions;
[0091] In the offline computing framework, first, the result data set is logically partitioned according to a preset partitioning rule (such as by product category, price range, or hash value, etc.). The selection of the partitioning rule is usually based on the distribution characteristics of the data and the requirements of the computing task to ensure that the data is evenly distributed among the partitions and avoid data skew. Then, according to the parallelism of the computing task (i.e., the number of computing nodes or the allocation of computing resources), the final number of partitions is determined. For example, if the parallelism of the computing task is 10, the result data set is divided into 10 partitions, and each partition contains part of the product data. In this way, the result data set can be divided into multiple partitions to ensure that the data volume of each partition is appropriate, which is convenient for parallel processing on multiple computing nodes, thereby improving the execution efficiency of the computing task and the data processing performance.
[0092] Step S7300: Accumulate the storage sizes of each partition to determine the total storage size of the result data set;
[0093] In one embodiment, the storage size information of each partition file can be obtained through the API of the distributed file system or the data operation tool provided by the computing framework. For example, for the partition files in HDFS, the metadata information of each file, including the file size, can be obtained through the API of the file system; for the partition data in the database, the partition size can be calculated by querying the size or occupied space of each record. Then, traverse all the partition files or data records, and accumulate the storage sizes of each partition in turn to obtain the total storage size of the result data set. For example, if the result data set contains 10 partitions, and the sizes of each partition are 1GB, 1.2GB, 0.9GB, etc., then the total storage size is the sum of these partition sizes. In this way, the total storage size of the result data set can be accurately calculated, providing key data support for subsequent file merging and storage optimization.
[0094] Step S7400: Store the number of partitions and the total storage size as metadata information in association with the corresponding result data set.
[0095] Pack the calculated number of partitions and total storage size into metadata information, including the number of partitions, total storage size, partition file list and its size, etc. Then, store this metadata information in association with the corresponding result data set. For example, the metadata information can be stored in a specific directory of the distributed file system (such as HDFS) as an attached file to the result data set; or the metadata information can be written into the metadata table of the database and associated with the storage path or table name of the result data set. In this way, the number of partitions and total storage size can be bound and stored as key metadata with the result data set, which is convenient for quickly obtaining and using this information in subsequent file merging, storage optimization, and query operations, thereby improving data management efficiency and system performance.
[0096] In this embodiment, through the optimization of the entire process from the commodity query request to data partitioning, storage size calculation, and metadata management, the efficient processing and storage management of the result data set are realized. The number of partitions and the total storage size are stored as metadata associated with the data set, facilitating direct acquisition when the number of partitions and the total storage size are required in this application.
[0097] In a further embodiment, please refer to Figure 7 , before storing the merged result data set into the data storage system, it includes:
[0098] Step S8100: Compare the check value of the result data set with the check value of the result data set after the file merging operation. If the check value of the merged result data set is consistent with the check value of the result data set before merging, store the merged result data set into the data storage system;
[0099] Before the file merging operation, calculate the check value of the result data set before merging. For example, calculate the hash value of the result data set using a hash algorithm as the check value and store it as a reference value. Then, after the file merging operation is completed, calculate the check value of the merged result data set and compare it with the check value of the result data set before merging. If the two are consistent, it indicates that the file merging operation has not caused any modification or damage to the data content. At this time, store the merged result data set into the target storage medium through the interface of the data storage system (such as the file writing interface of a distributed file system or the data insertion interface of a database). Through this step, the integrity and consistency of the data during the merging process can be ensured, avoiding data errors or losses caused by the merging operation, thereby improving the reliability and security of the data storage system.
[0100] Step S8200: If the check value of the merged result data set is inconsistent with the check value of the result data set before merging, terminate the file merging operation and generate an error log.
[0101] If the checksum of the merged result dataset is inconsistent with the checksum of the result dataset before merging, terminate the file merging operation and generate an error log. Specifically, after the file merging operation is completed, calculate the checksum of the merged result dataset and compare it with the checksum of the result dataset before merging. If the two are inconsistent, it indicates that the file merging operation may have modified or damaged the data content. At this time, immediately terminate the file merging operation to avoid further data errors or losses. Then, generate a detailed error log, record the specific information of the inconsistent checksum (such as the original checksum, the checksum after merging, the file path with the inconsistency, etc.), and store the error log in the specified log directory for subsequent analysis and troubleshooting. This enables the timely discovery and handling of abnormal situations in the file merging operation, ensuring the reliability and security of the data storage system.
[0102] In this embodiment, through the checksum comparison mechanism, the integrity and consistency of the data during the file merging operation are ensured. Calculate and store the checksum of the original dataset before file merging, calculate the checksum again after merging and compare them. If they are consistent, it is confirmed that the data is not damaged and is allowed to be stored in the data storage system; if they are inconsistent, immediately terminate the operation and generate an error log to prevent data errors or losses. This mechanism can not only effectively avoid data damage or loss caused by the merging operation, but also record the problem details through the error log, facilitating subsequent troubleshooting and repair, significantly improving the reliability, security and maintainability of the data storage system, and is applicable to business scenarios with high requirements for data integrity.
[0103] Please refer to Figure 8, A file merging device provided to meet one of the purposes of this application is a functional embodiment of the file merging method of this application. On the other hand, a file merging device provided to meet one of the purposes of this application includes a data set capture module 5100, an average storage calculation module 5200, a file merging module 5300, and a direct storage module 5400. Among them, the data set capture module 5100 is used to capture a result data set composed of files in multiple partitions to be stored in the data storage system, and obtain the number of partitions and the total storage size corresponding to the result data set; the average storage calculation module 5200 is used to determine the average storage size of a single partition according to the number of partitions and the total storage size; the file merging module 5300 is used to, when the average storage size is less than a preset size threshold, perform a file merging operation on the result data set to adjust the number of partitions and the storage size of each partition, so that the result data set contains files corresponding to multiple partitions, and the number of files with a storage size less than the size threshold in each partition does not exceed one, and store the merged result data set in the data storage system; the direct storage module 5400 is used to, when the average size is greater than the preset size threshold, store the result data set in the data storage system.
[0104] In a further embodiment, before the file merging module 5300, it includes: an existing file traversal sub-module, which is used to traverse the existing files in the data storage system and obtain the number and total size of the existing files; a size threshold adjustment sub-module, which is used to adjust the preset size threshold according to the number and total size of the existing files for comparison with the average storage size.
[0105] In a further embodiment, the size threshold adjustment sub-module includes: a density determination sub-module, which is used to divide the total size of the existing files by the number to obtain the average size of the existing files, calculate the ratio of the average size of the existing files to the preset size threshold, and determine it as the small file density; an increase adjustment sub-module, which is used to, when the small file density is less than the preset density threshold, determine an adjustment coefficient based on the small file density and the preset first mapping relationship, and perform an increase adjustment on the size threshold according to the adjustment coefficient; a decrease adjustment sub-module, which is used to, when the small file density is greater than the preset density threshold, determine an adjustment coefficient based on the small file density and the preset second mapping relationship, and perform a decrease adjustment on the size threshold according to the adjustment coefficient.
[0106] In a further embodiment, the file merging module 5300 includes: a merged data determination sub-module, configured to determine the number of files to be merged and the target size of each file after merging according to the size threshold and the total storage size when the average storage size is less than a preset size threshold; a merged data adjustment sub-module, configured to perform a file merging operation on the result data set based on the determined number of files and the target size, so as to adjust the number of partitions and the storage size of each partition to meet the number of files and the target size; a temporary directory writing sub-module, configured to write the result data set after merging the files into a temporary directory, rename the temporary directory, and store it in a distributed file system.
[0108] When the cumulative read data volume reaches the size threshold, merge the respective read data into the same file, and reset the cumulative read data volume to zero.
[0109] Iterate the above read and write data process until the result data set is empty and the corresponding cumulative read data volume is less than the size threshold, and then terminate the read and write data process.
[0110] Write the merged file into a temporary directory, rename the temporary directory, and store it in a distributed file system.
[0111] In a further embodiment, the file merging module 5300 includes: a target partition number determination sub-module, configured to determine the target number of partitions by using ceiling operation according to the ratio of the total storage size to the size threshold when the average storage size is less than a preset size threshold; a re-partitioning sub-module, configured to perform a re-partitioning operation on the result data set based on the target number of partitions to obtain a re-partitioned result data set; a database writing sub-module, configured to store the re-partitioned result data set into a specified table of a corresponding database through a database connection.
[0112] In a further embodiment, after the alternative text generation module 5500, it includes: a data set recall sub-module, configured to recall a result data set matching the query condition from a commodity database in response to a commodity query request; a partition division sub-module, configured to divide the result data set into multiple partitions according to a preset sharding rule and the parallelism of a computing task, and determine the number of partitions; a storage size determination sub-module, configured to accumulate the storage size of each partition to determine the total storage size of the result data set; an associated storage sub-module, configured to associate and store the number of partitions and the total storage size as metadata information with the corresponding result data set.
[0113] In a further embodiment, the dataset capture module 5100 includes: a check value comparison sub-module, configured to compare the check value of the result dataset with the check value of the result dataset after the file merging operation. If the check value of the merged result dataset is the same as that of the result dataset before merging, the merged result dataset is stored in the data storage system; a merge operation termination sub-module, configured to terminate the file merging operation and generate an error log if the check value of the merged result dataset is different from that of the result dataset before merging.
[0114] To solve the above technical problems, an embodiment of the present application further provides a computer device. As Figure 9 shown, it is a schematic internal structure diagram of the computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected through a system bus. Among them, the computer-readable storage medium of the computer device stores an operating system, a database, and computer-readable instructions. The control information sequence can be stored in the database. When the computer-readable instructions are executed by the processor, the processor can implement a file merging method. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The memory of the computer device can store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can execute the file merging method of the present application. The network interface of the computer device is used to communicate with the terminal. Those skilled in the art can understand that Figure 9 the structure shown in
[0115] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements. Figure 8 In this embodiment, the processor is used to execute the specific functions of each module and its sub-modules in
[0116] The memory stores the program codes and various types of data required to execute the above modules or sub-modules. The network interface is used for data transmission between the user terminal or the server. The memory in this embodiment stores the program codes and data required to execute all modules / sub-modules in the file merging device of the present application. The server can call the program codes and data of the server to execute the functions of all sub-modules.
[0117] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments of the present application can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a computer-readable storage medium such as a magnetic disk, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc.
[0118] In summary, the present application can provide users with timely and accurate responses.
[0119] Those skilled in the art of this technology can understand that the steps, measures, and solutions in the various operations, methods, and processes discussed in the present application can be alternated, changed, combined, or deleted. Further, the other steps, measures, and solutions in the various operations, methods, and processes discussed in the present application can also be alternated, changed, rearranged, decomposed, combined, or deleted. Further, the steps, measures, and solutions in the prior art that are the same as those in the various operations, methods, and processes open-sourced in the present application can also be alternated, changed, rearranged, decomposed, combined, or deleted.
[0120] The above are only some embodiments of the present application. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A file merging method, characterized in that: The steps include: Capturing a result data set composed of files in multiple partitions to be stored in a data storage system, and obtaining the number of partitions and the total storage size corresponding to the result data set; Determine an average storage size of a single partition according to the number of partitions and the total storage size; When the average storage size is smaller than a preset size threshold, performing a file merging operation on the result data set to adjust the number of partitions and the storage size of each partition, so that the result data set includes files corresponding to multiple partitions, and the storage size of each partition does not exceed one file smaller than the size threshold, and storing the merged result data set in the data storage system; When the average storage size is greater than a preset size threshold, the result data set is stored in the data storage system.
2. The file merging method according to claim 1, characterized in that: When the average storage size is less than a preset size threshold, performing a file merging operation on the result data set to adjust the number of partitions and the storage size of each partition, so that the result data set includes files corresponding to multiple partitions, and the storage size of each partition does not exceed one file that is less than the size threshold, and before storing the merged result data set in the data storage system, including: Traversing existing files in the data storage system to obtain the number and total size of the existing files; According to the number and total size of the existing files, a preset size threshold is adjusted for comparison with the average storage size.
3. The file merging method according to claim 2, characterized in that: According to the number and total size of the existing files, adjusting a preset size threshold for comparison with the average storage size includes: The total size of the existing files is divided by the number to obtain an average size of the existing files, and the ratio of the average size of the existing files to the preset size threshold is calculated to determine the density of small files; If the small file density is less than a preset density threshold, determining an adjustment coefficient based on the small file density and a preset first mapping relationship, and increasing and adjusting the size threshold according to the adjustment coefficient; If the small file density is greater than a preset density threshold, an adjustment coefficient is determined based on the small file density and a preset second mapping relationship, and the size threshold is reduced and adjusted according to the adjustment coefficient.
4. The file merging method according to claim 1, characterized in that: When the average storage size is less than a preset size threshold, performing a file merging operation on the result data set to adjust the number of partitions and the storage size of each partition, so that the result data set includes files corresponding to multiple partitions, and the storage size of each partition does not exceed one file that is less than the size threshold, and storing the merged result data set in the data storage system, including: When the average storage size is less than a preset size threshold, start the data reading and writing process, continue to read the data in the result data set, accumulate the size of the data read each time, and obtain the accumulated read data amount; When the accumulated read data volume reaches the size threshold, the corresponding read data are merged and written into the same file, and the accumulated read data volume is reset to zero; The above-mentioned data reading and writing process is iterated until the result data set is empty and the corresponding accumulated read data amount is less than the size threshold, and then the data reading and writing process is terminated. The merged file is written into a temporary directory, the temporary directory is renamed, and stored in a distributed file system.
5. The file merging method according to claim 1, characterized in that: When the average storage size is less than a preset size threshold, performing a file merging operation on the result data set to adjust the number of partitions and the storage size of each partition, so that the result data set includes files corresponding to multiple partitions, and the storage size of each partition does not exceed one file that is less than the size threshold, and storing the merged result data set in the data storage system, including: When the average storage size is less than a preset size threshold, determining the target partition quantity by rounding up according to the ratio of the total storage size to the size threshold; Repartitioning the result data set based on the target number of partitions to obtain a repartitioned result data set; The repartitioned result data set is stored in the specified table of the corresponding database through the database connection.
6. The file merging method according to any one of claims 1 to 5, characterized in that: Before capturing a result data set composed of files in multiple partitions to be stored in a data storage system and obtaining the number of partitions and the total storage size corresponding to the result data set, the method includes: In response to a product query request, a result data set matching the query condition is retrieved from the product database; According to the preset sharding rules and the parallelism of the computing task, the result data set is divided into multiple partitions and the number of partitions is determined; The storage size of each partition is accumulated to determine the total storage size of the result data set; The partition quantity and the total storage size are used as metadata information and stored in association with the corresponding result data set.
7. The file merging method according to any one of claims 1 to 5, characterized in that: Before storing the merged result data set in the data storage system, the method includes: comparing the check value of the result data set with the check value of the result data set after the file merging operation, and if the check value of the result data set after merging is consistent with the check value of the result data set before merging, storing the merged result data set in the data storage system; If the checksum of the result data set after merging is inconsistent with the checksum of the result data set before merging, the file merging operation is terminated and an error log is generated.
8. A file merging device, characterized in that: include: A data set capture module, used to capture a result data set consisting of files in multiple partitions to be stored in a data storage system, and obtain the number of partitions and the total storage size corresponding to the result data set; An average storage calculation module, used to determine the average storage size of a single partition according to the number of partitions and the total storage size; A file merging module, configured to, when the average storage size is smaller than a preset size threshold, perform a file merging operation on the result data set to adjust the number of partitions and the storage size of each partition, so that the result data set includes files corresponding to multiple partitions, and the storage size of each partition does not exceed one file smaller than the size threshold, and store the merged result data set in the data storage system; The direct storage module is used to store the result data set into the data storage system when the average storage size is greater than a preset size threshold.
9. A computer device comprising a central processing unit and a memory, characterized in that: The central processing unit is used to call and run the computer program stored in the memory to execute the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: It stores a computer program implemented according to the method described in any one of claims 1 to 7 in the form of computer-readable instructions, and when the computer program is called and executed by a computer, the steps included in the corresponding method are executed.