Data reading method, device, storage medium, and program product
Patent Information
- Application Number
- PCT/CN2026/079093
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-28
- Filing Date
- 2026-02-12
- Publication Date
- 2026-10-01
Smart Images

Figure CN2026079093_01102026_PF_FP_ABST
Abstract
Description
Data reading methods, devices, storage media and software products
[0001] This disclosure claims priority to Chinese Patent Application No. 202510381288.X, filed with the China Patent Office on March 28, 2025, entitled “Data Reading Method, Apparatus, Storage Medium and Program Product”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This disclosure relates to the field of data storage technology, and in particular to a data reading method, device, storage medium, and program product. Background Technology
[0003] In the fields of Artificial Intelligence (AI) and Machine Learning (ML), the quality, suitability, and validity of data are key factors determining the performance of AI models. To ensure these characteristics, a process of scanning, analyzing, and understanding the data is typically required. Data collection is a crucial step in data scanning, especially when dealing with large-scale datasets. Training datasets can often reach hundreds of terabytes (TB) or even petabytes (PB), containing hundreds of millions of files.
[0004] For training purposes, AI training systems need to read these data files from storage systems. However, existing data reading methods, especially when dealing with large-scale datasets, suffer from high latency, making them a significant bottleneck in terms of data reading efficiency. Summary of the Invention
[0005] This disclosure provides a data reading method, device, storage medium, and program product to reduce data reading latency and improve data reading efficiency.
[0006] This disclosure provides a data reading method, including:
[0007] In response to a data read request, the storage address information and data length of multiple objects to be read are obtained; the storage address information of the multiple objects includes: the storage address information of multiple data fragments corresponding to each of the multiple objects; the multiple data fragments corresponding to each object are stored on multiple magnetic storage modules;
[0008] Based on the storage address information of the multiple data fragments corresponding to each of the multiple objects, the storage address information of the data fragments stored by each of the multiple magnetic storage modules is determined;
[0009] With the goal of maximizing sequential access to the magnetic storage modules, the multiple objects are sorted according to the storage address information of the data fragments stored in each of the multiple magnetic storage modules to obtain the target reading order.
[0010] Based on the target reading order, the storage address information of the multiple data fragments corresponding to each of the multiple objects, and the data length, the multiple objects are read from the multiple magnetic storage modules.
[0011] This disclosure also provides an electronic device, including: a memory and a processor; wherein the memory is used to store a computer program;
[0012] The processor is coupled to the memory and is used to execute the computer program to perform the steps in the data reading method described above.
[0013] This disclosure also provides a computer-readable storage medium storing computer instructions, which, when executed by one or more processors, cause the one or more processors to perform the steps in the data reading method described above.
[0014] This disclosure also provides a computer program product, including a computer program that, when executed by one or more processors, causes the one or more processors to perform the steps in the data reading method described above.
[0015] In this embodiment of the disclosure, for a distributed storage scenario where multiple objects to be read include multiple data fragments and the multiple data fragments are stored on multiple magnetic storage modules, with the goal of maximizing sequential access to the magnetic storage modules, the reading order of the multiple objects is sorted according to the storage address of the data fragments stored in each of the multiple magnetic storage modules to obtain the target reading order; and multiple objects are read from the multiple magnetic storage modules according to the target reading order, thereby maximizing sequential access to the magnetic storage modules, reducing the randomness of random access to the magnetic storage modules, and thus helping to reduce the access latency of data reading and improve data reading efficiency. Attached Figure Description
[0016] The accompanying drawings, which are included to provide a further understanding of this disclosure and form part of this disclosure, illustrate exemplary embodiments of the present disclosure and are used to explain the disclosure, but do not constitute an undue limitation of the disclosure. In the drawings:
[0017] Figure 1 is a flowchart illustrating the data reading method provided in an embodiment of this disclosure;
[0018] Figures 2 and 3 are schematic diagrams of the data SCAN optimization process provided in the embodiments of this disclosure;
[0019] Figure 4 is a schematic diagram of the structure of the electronic device provided in the embodiment of this disclosure. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this disclosure clearer, the technical solutions of this disclosure will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.
[0021] It should be noted that, in the cases involving user information in the embodiments of this disclosure, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0022] In AI training or big data analytics scenarios, data scanning typically refers to the process of scanning, analyzing, and understanding data to ensure its quality, applicability, and validity. This process is crucial for any successful machine learning project because the model's performance is highly dependent on the data used. In AI training scenarios, key steps in data scanning include:
[0023] 1. Data Collection: During data collection, it is necessary to determine which data is relevant. Then, this relevant data is collected from various sources. The collected relevant data may include structured data (such as tables in a database), semi-structured data (such as JSON files), or unstructured data (such as text, images, or videos). Generally, data collection can be implemented by reading relevant data from data sources.
[0024] 2. Data Cleaning: Once the data is collected, the next step is to clean it. This step aims to remove inaccurate, incomplete, or irrelevant data, while also handling missing values, duplicate records, and outliers.
[0025] 3. Data Analysis: In this stage, the data undergoes in-depth analysis to understand its characteristics, distribution, and potential patterns. This helps identify trends, correlations, and factors that may affect model performance.
[0026] 4. Data Labeling: In AI training scenarios, especially in supervised learning scenarios, data labeling is necessary (e.g., adding labels to images in image classification tasks). High-quality labeling is crucial for training accurate models.
[0027] 5. Data transformation and preprocessing: In order to make the data suitable for a specific algorithm, it may be necessary to convert it into a specific format or apply certain preprocessing techniques (such as normalization, dimensionality reduction, etc.).
[0028] 6. Validation and test set preparation: Divide the data into training, validation and test sets to evaluate the model's performance and prevent overfitting.
[0029] The above data SCAN has the following advantages: (1) By carefully scanning the data, the accuracy of the final AI model can be significantly improved; (2) Careful examination of the data during the data SCAN process can help identify and reduce bias in the dataset, thereby making the model more fair and just; (3) Optimize resource use: Effective data SCAN can help identify the most relevant and valuable data, thereby making more efficient use of computing resources.
[0030] In summary, data scanning is an indispensable step in AI training, directly impacting the success of the training process. Proper data processing not only improves model performance but also accelerates development cycles and reduces maintenance costs.
[0031] Data collection is a crucial step in data scanning. It typically involves reading trillions of bytes (TB) of data, or even more. Therefore, improving data reading efficiency can significantly shorten the data scanning time. A common method for data reading is to iterate through (list) the prefix (or directory) of the "training dataset" to obtain the identifiers or filenames of the objects to be accessed. In AI training or big data analysis scenarios, the number of identifiers or filenames of the objects to be accessed can reach hundreds of millions, or even more. The identifiers or filenames of the objects to be accessed are shown below:
[0032] Then, the files are read in the order described above to complete data collection. For data storage systems where the data source is a magnetic storage device such as a hard disk drive (HDD), the above scheme results in high data read latency. For example, File 1 is stored in the middle of HDD1, File 2 is stored at the beginning of HDD2, and File 3 is stored at the beginning of HDD1. If data is read in the order of File1->File2->File3, it leads to significant random access, which has a large impact on backend storage and results in longer data scan times. This is mainly because when reading data from an HDD, the read / write head needs to move to the correct track and wait for the disk to rotate to the correct position (i.e., seek time and rotational latency). Random data reading requires frequent head rotation, resulting in high read latency. For example, for File 1 and File 3, the read / write head of HDD1 needs to first rotate from its current position to the middle of HDD1 to read the data corresponding to File 1, and then return from the middle of HDD1 to the beginning of HDD1, resulting in long seek time and high rotational latency, thus causing high read latency.
[0033] In some embodiments of this disclosure, in order to reduce data reading latency, for distributed storage scenarios where multiple objects to be read include multiple data fragments and the multiple data fragments are stored on multiple magnetic storage modules, with the goal of maximizing sequential access to the magnetic storage modules, the reading order of multiple objects is sorted according to the storage addresses of the data fragments stored in each of the multiple magnetic storage modules to obtain a target reading order; and multiple objects are read from the multiple magnetic storage modules according to the target reading order, thereby maximizing sequential access to the magnetic storage modules, reducing the randomness of random access to the magnetic storage modules, and thus helping to reduce data reading latency and improve data reading efficiency.
[0034] The technical solutions provided by the embodiments of this disclosure are described in detail below with reference to the accompanying drawings.
[0035] It should be noted that the same reference numerals in the following figures and embodiments denote the same object or the same step. Therefore, once an object or step is defined in one figure or embodiment, it does not need to be discussed further in subsequent figures and embodiments.
[0036] Figure 1 is a flowchart illustrating the data reading method provided in this embodiment. As shown in Figure 1, the data reading method mainly includes the following steps:
[0037] 101. In response to a data read request, obtain the storage address information and data length of the multiple objects to be read; the storage address information of the multiple objects includes: the storage address information of the multiple data fragments corresponding to each of the multiple objects; the multiple data fragments corresponding to each object are stored on multiple magnetic storage modules.
[0038] 102. Based on the storage address information of the multiple data fragments corresponding to each of the multiple objects, determine the storage address information of the data fragments stored by each of the multiple magnetic storage modules.
[0039] 103. With the goal of maximizing sequential access to magnetic storage modules, sort multiple objects according to the storage address information of the data fragments stored in each magnetic storage module to obtain the target reading order.
[0040] 104. Based on the target reading order, the storage address information and data length of the multiple data fragments corresponding to each of the multiple objects, read multiple objects from multiple magnetic storage modules.
[0041] In this embodiment of the disclosure, a magnetic storage module is a hardware device that uses magnetic materials to record and store data. These devices represent binary data (0 and 1) by changing the direction of the magnetic field on the magnetic medium, thereby enabling data storage and retrieval. The magnetic storage module can be a mechanical hard disk, also known as a hard disk drive (HDD) and / or a magnetic disk, etc.
[0042] An object refers to a data object stored in a data source (such as a database or data storage system), which may include: structured data (such as tables in a database) and / or semi-structured data (such as JSON files) and / or unstructured data (such as text, images, or videos). In embodiments of this disclosure, the object may be a file, a data table, or other forms.
[0043] To improve data reliability, objects are typically divided into multiple data shards, which are distributed and stored across multiple magnetic storage modules (two or more). Objects can include metadata and application data. Metadata and application data are stored independently, separating control flow from data flow to achieve higher policy scalability and input / output (IO) concurrency. Metadata refers to the metadata of application data, which may include the storage address information and data length of the application data. In an embodiment where an object is divided into multiple data shards, the metadata may include: the storage address information of the multiple data shards corresponding to the object and the data length of the corresponding data shards. For example, as shown in Figure 2, the metadata of object 1 is stored on HDD1, data shard 1 of object 1 is stored on HDD2, data shard 2 of object 1 is stored on HDD4, and data shard 3 of object 1 is stored on HDD5. Similarly, the metadata of object 2 is stored on HDD1, data shard 1 of object 2 is stored on HDD4, data shard 2 of object 2 is stored on HDD7, and data shard 3 of object 2 is stored on HDD100. The metadata of object 3 is stored on HDD4, data fragment 1 of object 3 is stored on HDD2, data fragment 2 of object 3 is stored on HDD6, and data fragment 3 of object 3 is stored on HDD100. Here, HDDn refers to the HDD numbered n, where n = 1, 2, 3, ...
[0044] In this embodiment, a magnetic storage module can store data fragments of different objects, and the data fragments of different objects are stored in different locations within the magnetic storage module. For example, HDD2 stores data fragment 1 of object 1 and data fragment 1 of object 3, and the storage location of data fragment 1 of object 1 is located before the storage location of data fragment 1 of object 3.
[0045] Magnetic storage modules typically consist of one or more rotating disks, each with two surfaces for storing data. A read / write head, positioned above each disk surface, is responsible for reading and writing data. Seek time is the time required for the head to move from the current track to the target track. This is one of the most time-consuming parts of read / write operations in a magnetic storage module. Random access often requires frequent head movements, resulting in high seek times. Rotational latency is the time required to wait for the target sector to rotate under the head. Since disks rotate at a constant speed, rotational latency depends on the disk's rotational speed and the location of the target sector. Transfer time is the time actually spent reading or writing data. Once the head is positioned correctly, data can be transferred quickly.
[0046] In random access mode, the read / write head needs to move frequently between different tracks, and each movement requires a certain amount of seek time. On the other hand, since data blocks may be scattered across different locations on the disk, after the head reaches the target track, it needs to wait for the target sector to rotate under the head, which adds additional rotational latency. Each read or write of a small amount of data involves the aforementioned seek time and rotational latency, resulting in low overall transfer efficiency.
[0047] In sequential access mode, data is typically stored contiguously on the same or adjacent tracks. This reduces the need for frequent head movements, significantly decreasing seek time. Even when switching to an adjacent track is necessary, this movement is much faster than moving across the entire disk. If data is stored sequentially, the head can read multiple consecutive blocks of data in a single rotational cycle, greatly reducing the impact of rotational latency. Ideally, the head only needs to wait for one rotational latency to read a large amount of data. Therefore, sequential access magnetic storage modules, compared to random access magnetic storage modules, reduce the latency of reading data from the magnetic storage module and improve data reading efficiency.
[0048] The random access mode and sequential access mode can be compared to elevator scheduling. The random access mode is analogous to random floor scheduling in an elevator. For example, suppose user 1 is on the 30th floor, user 2 is on the 2nd floor, and user 3 is on the 27th floor. Random floor scheduling would follow the order in which users pressed the elevator buttons. If the order is: User 1, User 2, and User 3, then the elevator would first ascend to the 30th floor, then descend to the 2nd floor, then ascend again to the 27th floor, and finally descend from the 27th floor to the 1st floor, resulting in a long travel time. The sequential access mode, analogous to sequential floor scheduling, follows the order of the elevator's floors. In this case, the elevator would first ascend to the 30th floor, then descend to the 27th floor, then descend to the 2nd floor, and finally descend from the 2nd floor to the 1st floor. This sequential scheduling method undoubtedly saves travel time compared to random floor scheduling.
[0049] Based on the above analysis, in order to improve the data reading efficiency in the magnetic storage module, a sequential access mode can be used to read data from the magnetic storage module in this embodiment. In order to reduce memory space usage during data reading, the data of one object is read first, and then another object is read. This is mainly because: if multiple object data fragments are directly read from the magnetic storage module according to their storage address information in the same magnetic storage module, the read data fragments of multiple objects need to be stored in memory. After all data fragments of multiple objects are read from other magnetic storage modules, the multiple objects are then transmitted to the downstream process for data processing. This undoubtedly leads to a large number of data fragments remaining in the data reading device, occupying the memory of the data reading device. By reading all data fragments of one object first, and then reading another object, the object can be directly transmitted to the downstream process for data processing after all data fragments of one object have been read, reducing the memory usage of data fragments on the data reading device. Therefore, the data reading method provided in this embodiment reads data on an object-by-object basis. Therefore, in order to maximize sequential access to the magnetic storage module, the data reading method provided in this disclosure requires sorting the reading order of multiple objects to be read.
[0050] Therefore, in this embodiment, to maximize sequential access to the magnetic storage module, in step 101, in response to a data read request, the storage address information and data length of the multiple objects to be read can be obtained. Since the objects are divided into multiple data fragments and distributed across multiple magnetic storage modules, the storage address information of the multiple objects may include: the storage address information of the multiple data fragments corresponding to each of the multiple objects. The data length of the multiple objects may include: the data length of the multiple data fragments corresponding to each of the multiple objects.
[0051] Specifically, in response to a data read request, the identifiers of multiple objects to be read can be obtained from the data read request; then, based on the identifiers of the multiple objects, the storage address information and corresponding data length of the data fragments corresponding to each of the multiple objects can be obtained from the metadata storage nodes in the distributed storage system.
[0052] The storage address information of a data fragment may include: the identifier of the magnetic storage module where the data fragment resides and the offset address of the data fragment within the magnetic storage module. The identifier of the magnetic storage module determines which magnetic storage module the data fragment is located in. The offset address of the data fragment within the magnetic storage module determines its relative address within the magnetic storage module, i.e., its specific location within the magnetic storage module.
[0053] In this disclosure, the source of the data read request is not limited. In some embodiments, the data read request may be provided by the user processing the data. The user may send a data read request to the distributed storage system when data collection is required. Accordingly, the storage service node in the distributed storage system may receive the data read request. In other embodiments, the storage service node may also automatically initiate a data read request upon being triggered by a specific event, etc. The data read request may include identifiers of multiple objects to be read.
[0054] Furthermore, based on the identifiers of multiple objects, the storage address information and corresponding data length of each data fragment corresponding to the object can be obtained from the metadata storage node in the distributed storage system.
[0055] Since the storage address information of each object's data fragment reflects the relative position of the object's data fragment within the magnetic storage module, and sequential access refers to rotating the magnetic storage module's head in the same seek direction, and since the storage address information of the data fragments stored in each magnetic storage module reflects the relative position of the data fragments within that magnetic storage module, in step 102, the storage address information of the data fragments stored in each of the multiple magnetic storage modules can be determined based on the storage address information of the multiple data fragments corresponding to multiple objects.
[0056] As shown in Table 1, the storage address information of the data fragments may include the identifier of the magnetic storage module where the data fragment is located. For example, Tables 1 and 2 show that the identifiers of the magnetic storage modules where the data fragment of object 1 is located are HDD1, HDD2, HDD4, and HDD5, respectively. As shown in Table 2, the identifiers of the magnetic storage modules where the data fragment of object 2 is located are HDD1, HDD4, HDD7, and HDD100, respectively; and the identifiers of the magnetic storage modules where the data fragment of object 3 is located are HDD5, HDD2, HDD6, and HDD100, respectively. In Tables 1 and 2, the identifier of the data block (Chunk) represents the identifier of the data fragment. Accordingly, the storage address information of the data fragments stored in each magnetic storage module can be determined based on the identifiers of the magnetic storage modules where the multiple data fragments corresponding to multiple objects are located.
[0057] Table 1 shows the storage address information and data length of multiple data fragments corresponding to a single object.
[0058] Table 2 shows the storage address information and data length of multiple data fragments corresponding to multiple objects.
[0059] For the same magnetic storage module, data can be read in the order of the data fragments in the corresponding column of the magnetic storage module to achieve sequential access to the magnetic storage module. Since the data fragments of the same object are distributed and stored on multiple magnetic storage modules, in order to achieve the goal of sequential access, the frequency of the read / write head rotating back and forth on the same magnetic storage module can be minimized to ensure sequential access within the same magnetic storage module. This requires sorting the reading order of multiple objects. Based on this, in step 103, with the goal of maximizing sequential access to the magnetic storage module, multiple objects can be sorted according to the storage address information of the data fragments stored in each of the multiple magnetic storage modules to obtain the target reading order. Specifically, sorting multiple objects according to the storage address information of the data fragments stored in each of the multiple magnetic storage modules with the goal of maximizing sequential access to the magnetic storage module can make the data fragments in each magnetic storage module as much as possible be sorted in the same direction. This can reduce the randomness of reading multiple objects according to the target reading order in the future, achieve or approximate sequential access, and thus reduce data reading latency.
[0060] In this embodiment, the specific implementation of sorting multiple objects based on the storage address information of the data fragments stored in each of the multiple magnetic storage modules is not limited. In some embodiments, for any magnetic storage module A among the multiple magnetic storage modules, the data fragments stored in magnetic storage module A are sorted according to the size relationship of the offset addresses of the data fragments stored in magnetic storage module A, so as to obtain the arrangement order of the data fragments in magnetic storage module A.
[0061] In some embodiments, the data fragments stored in magnetic storage module A can be sorted according to their offset addresses in ascending order to obtain the arrangement order of the data fragments stored in storage device A. In other embodiments, the data fragments stored in magnetic storage module A can be sorted according to their offset addresses in descending order to obtain the arrangement order of the data fragments in magnetic storage module A. Using the same method, the arrangement order of the data fragments stored in each of the magnetic storage modules can be obtained.
[0062] Furthermore, with the goal of maximizing sequential access to the magnetic storage module, multiple objects can be sorted according to the arrangement order of the data fragments stored in each magnetic storage module to obtain the target read order. Specifically, sorting the data fragments stored in magnetic storage module A according to the offset address relationship ensures that the data fragments in magnetic storage module A are arranged in the same direction, providing a basis for subsequent sequential access to the magnetic storage module. Thus, by prioritizing sequential access to the magnetic storage module and sorting multiple objects according to the arrangement order of the data fragments stored in each magnetic storage module, the target read order obtained can reduce the randomness of accessing the magnetic storage module, which is beneficial for achieving sequential access to the magnetic storage module.
[0063] In some embodiments, for any magnetic storage module A, the offset addresses of the data fragments stored in magnetic storage module A can be written into the corresponding column of the address mapping table according to the arrangement order of the data fragments stored in magnetic storage module A, so as to obtain an address mapping table for multiple objects. Using the same method, the offset addresses of the data fragments of all objects stored in all magnetic storage modules can be written into the corresponding columns of the respective magnetic storage modules to obtain an address mapping table for multiple objects. That is, as shown in Figure 2, the data mapping module can map the storage address information of the data fragments of the multiple objects to be read shown in Table 2 to the address mapping table shown in Table 3 below using the above method. Figure 2 only illustrates Table 3 by using the relative positions of the data fragments in the magnetic storage module (such as HDD). The address mapping table is shown in Table 3 below. The rows in the address mapping table represent multiple magnetic storage modules, such as HDD1, HDD, HDD4, etc. shown in Table 3; the columns in the address mapping table represent the offset addresses of the data fragments in the corresponding magnetic storage module within that magnetic storage module.
[0064] Table 3 Address Mapping Table
[0065] As shown in Table 3, the order of the offset addresses of data fragments in each column of the address mapping table indicates the sorting distribution of data fragments within the magnetic storage module. Therefore, to achieve sequential access to the magnetic storage module, data fragments within the same magnetic storage module should be read as much as possible according to the order of their corresponding columns in the address mapping table. Based on this, to achieve the goal of sequential access to the magnetic storage module, multiple objects can be sorted according to the distribution of their data fragments in the address mapping table, with the goal of maximizing sequential access, to obtain the target reading order. In this embodiment, the order of data fragments in each column of the address mapping table, representing the sorting distribution of data fragments within the magnetic storage module, provides a clear view showing the position of data fragments on each magnetic storage module. This facilitates the management and tracking of the data distribution of each object within the magnetic storage module, thereby enabling the sorting of multiple objects with the goal of maximizing sequential access to the magnetic storage module.
[0066] As shown in Table 3, the distribution of data fragments for multiple objects in the address mapping table includes the number of data fragments for each object in each row. To achieve sequential access to the magnetic storage module, the number of data fragments for each object in each row is compared sequentially from smallest to largest row number (i.e., from front to back). The reading order of multiple objects is determined based on the number of data fragments in the row that appears first. For objects whose reading order cannot be determined based on the number of data fragments in the row that appears first, the number of data fragments in the next row is used to determine the reading order of objects in the preceding rows whose reading order is still uncertain. The earlier a row contains more data fragments for a particular object, the greater the likelihood that reading that object first will achieve sequential access to the magnetic storage module.
[0067] Based on this, in some embodiments, the number of data fragments of multiple objects contained in the first row of the address mapping table can be compared first. The number of data fragments for each object in each row is determined by the actual situation and can be 0. Furthermore, multiple objects can be sorted according to the descending order of the number of data fragments contained in the first row. For example, as shown in Table 3, the data fragments in Table 3 refer to the data fragments of the application data of the objects. The first row of the address mapping table includes: 3 data fragments of object 1 (i.e., the metadata of object 1, data fragment 1, and data fragment 3), 2 data fragments of object 2 (i.e., data fragment 1 and data fragment 2 of object 2), and 2 data fragments of object 3 (i.e., data fragment 2 and data fragment 3 of object 3). Therefore, based on the number of data fragments of each object contained in the first row of the address mapping table shown in Table 3, the reading order of object 1 can be determined to be first.
[0068] Since the first row of the address mapping table contains the same number of data fragments for objects 2 and 3, the reading order of objects 2 and 3 cannot be determined based solely on the number of data fragments for each object in the first row. Therefore, after traversing the first row, if there are multiple target objects whose reading order is not determined (such as objects 2 and 3 in Table 3), the number of data fragments for the multiple target objects in the second row of the address mapping table can be compared, and the multiple target objects can be sorted based on the number of data fragments for the multiple target objects in the second row. If the second row cannot determine the reading order of all target objects, the number of data fragments for the target objects in the third row of the address mapping table can be compared, and so on, until the reading order of the multiple objects to be read is determined or all rows of the address mapping table have been traversed.
[0069] The process of sorting multiple objects to be read can be summarized as follows: When multiple target objects exist whose reading order is not yet determined, the following steps are executed repeatedly until a set loop stopping condition is met. The target reading order is then determined based on the reading order of the multiple objects at the loop stopping point. The repeatedly executed steps include: determining the target row with the smallest row number from the rows not yet traversed in the address mapping table; and sorting the multiple target objects according to the number of data fragments of the multiple target objects in the target row from largest to smallest, so as to access the magnetic storage module sequentially. The set loop stopping condition may include: the reading order of all multiple objects to be read is determined, and / or all rows of the address mapping table have been traversed.
[0070] Since the order of rows in the address mapping table reflects the order of addresses of the magnetic storage modules, the reading order is determined by comparing the number of fragments of each object in the address mapping table row by row. Prioritizing the reading of objects with more data fragments in each row can maximize the chance of sequential reading and minimize the randomness of magnetic storage module access.
[0071] When the aforementioned loop terminates, the reading order of all objects to be read may have been determined. Accordingly, if the loop stops when the reading order of the multiple objects has been determined, the reading order of the multiple objects determined when the loop stops is determined as the target reading order.
[0072] In some embodiments, when the loop terminates, such as when all rows of the address mapping table have been traversed, there may still be multiple objects whose reading order has not been determined. These objects then need to be sorted. For ease of description, objects whose reading order remains undetermined when the loop stops are defined as the first object. For example, as shown in Table 3 above, using the method of comparing the size of the data fragments of each object row by row to sort the reading order of multiple objects, if the reading order of object 2 and object 3 still cannot be determined after all rows of the address mapping table have been traversed, then the first object includes object 2 and object 3.
[0073] In some embodiments, the reading order of the first objects can be determined randomly, or the first objects can be sorted according to the original reading order of these objects in the data reading request, etc. This method of determining the reading order of multiple first objects is simple and can improve the speed of determining the reading order of objects. However, this method of determining the reading order of multiple first objects still makes the reading of multiple first objects random.
[0074] To maximize sequential access to the magnetic storage module, multiple first objects can be sorted with the goal of maximizing sequential access. In some embodiments, all permutations of the multiple first objects can be performed to obtain multiple read orders for the multiple first objects. Here, "all permutations" refers to a mathematical concept encompassing all possible arrangements of the multiple objects. For example, assuming there are n first objects, there are n! (n factorial) possible read orders for the n first objects. As shown in Table 3, if there are two first objects, including object 2 and object 3, then the read orders for the two objects are as follows. Since sequential access to the magnetic storage module refers to the read / write head rotating in the same direction, maximizing sequential access should reduce the frequency of head rotations in the magnetic storage module. Therefore, to maximize sequential access, the required head rotation frequency for each of the multiple read orders for the multiple first objects can be determined.
[0075] Specifically, the target columns where the data fragments of multiple first objects reside can be determined from the address mapping table. As shown in Table 3, the target columns where the data fragment of object 2 resides are: the column corresponding to HDD1, the column corresponding to HDD4, the column corresponding to HDD7, and the column corresponding to HDD100; the target columns where the data fragment of object 3 resides are: the column corresponding to HDD2, the column corresponding to HDD5, the column corresponding to HDD6, and the column corresponding to HDD100.
[0076] Furthermore, the reading order of the objects whose reading order was determined at the aforementioned loop cutoff can also be obtained. For ease of description and distinction, the objects whose reading order was determined at the loop cutoff are defined as the second objects. There may be one or more second objects. "Multiple" refers to two or more (including two). In Table 3 above, the second object at the loop cutoff is object 1. Then, based on the distribution of the objects to which the data slices stored in the multiple target columns belong and the reading order of the second objects, the required head rotation frequency for reading each of the multiple first objects according to multiple reading orders can be determined.
[0077] For example, regarding Table 3 above, if the reading order is to read object 1 first, then object 2, and finally object 3, for HDD4, reading object 1 first and then object 2 will cause one head rotation; for HDD100, reading object 2 first and then object 3 will also cause one head rotation. Therefore, the head rotation frequency required for the reading order is 2 times. If the reading order is to read object 1 first, then object 3, and finally object 2, for HDD4, reading object 1 first and then object 2 will cause one head rotation, while other target columns are read sequentially. Therefore, the head rotation frequency required for the reading order is 1 time. Therefore, to maximize sequential access, the reading order for objects 2 and 3 can be determined as reading object 3 first and then object 2. Based on this, the reading order with the lowest head rotation frequency can be determined as the reading order for multiple first objects. Accordingly, the target reading order can be determined based on the reading order with the lowest head rotation frequency and the reading order of the second object. Specifically, the reading order with the lowest head rotation frequency can be arranged after the reading order of the second object to obtain the target reading order. For example, as shown in Figure 3, for the address mapping table shown in Table 3 above, the target reading order of the multiple objects to be read (i.e., object 1, object 2, and object 3) can be obtained as: object 1 -> object 3 -> object 2.
[0078] In this embodiment, the sorting method that minimizes the frequency of magnetic head rotation supplements the aforementioned sorting method that compares the size of data fragments of each object row by row. This helps to maximize sequential access and reduce the probability of random access.
[0079] In some embodiments of this disclosure, the reading order of multiple objects can be sorted using head rotation parameters. These head rotation parameters reflect the rotation of the magnetic storage module's head when reading data in a predetermined order, and the head rotation can, to some extent, reflect the sequential nature of the head's rotation. The head rotation parameters may include head rotation frequency and / or head rotation distance. Head rotation frequency refers to the number of times the head first rotates to a large offset address to read data, and then returns to a small offset address to read data again. Head rotation distance refers to the distance the head rotates during data reading.
[0080] Based on this, we can perform a full permutation of the multiple objects to be read, resulting in multiple candidate reading orders for the objects. Assuming there are n objects to be read, there are n! (n factorial) possible reading orders. Assuming there are 3 objects to be read, including object 1, object 2, and object 3, there are 6 possible reading orders for the 3 objects, including: object 1 → object 2 → object 3, object 1 → object 3 → object 2, object 2 → object 1 → object 3, object 2 → object 3 → object 1, object 3 → object 1 → object 2, and object 3 → object 2 → object 1.
[0081] Furthermore, based on the distribution of data fragments of multiple objects in the address mapping table, the required head rotation parameters for reading each of the multiple objects according to multiple candidate read orders can be determined. Accordingly, with the goal of maximizing access to the magnetic storage module, a target read order can be determined from multiple candidate read orders based on the head rotation parameters. Sorting multiple objects according to the head rotation parameters can minimize the randomness of reading multiple objects, thereby maximizing sequential access to the magnetic storage module.
[0082] In some embodiments, the head rotation parameters include the head rotation frequency. The multiple objects to be read can be sorted by minimizing the head rotation frequency. Specifically, the multiple objects to be read can be fully permuted to obtain multiple candidate reading orders. Further, the required head rotation frequency for each of the multiple objects can be determined based on the distribution of data fragments of the multiple objects in the address mapping table, according to the multiple candidate reading orders. For example, according to the address mapping table shown in Table 3 above, it can be determined that the required head rotation frequency for reading object 1 -> object 2 -> object 3 is 2 times. The required head rotation frequency for reading object 1 -> object 3 -> object 2 is 1 time. The required head rotation frequency for reading object 2 -> object 1 -> object 3 is 2 times, including: 1 head rotation caused by reading object 2 first and then object 1 in HDD1, and 1 head rotation caused by reading object 2 first and then object 3 in HDD100. The required head rotation frequency for reading from object 2 to object 3 to object 1 is 4 times, including: 1 head rotation caused by reading object 2 first and then object 1 in HDD1, 1 head rotation caused by reading object 3 first and then object 1 in HDD2, 1 head rotation caused by reading object 3 first and then object 1 in HDD5, and 1 head rotation caused by reading object 2 first and then object 3 in HDD100. The required head rotation frequency for reading from object 3 to object 1 to object 2 is 3 times, including: 1 head rotation caused by reading object 3 first and then object 1 in HDD2, 1 head rotation caused by reading object 1 first and then object 2 in HDD4, and 1 head rotation caused by reading object 3 first and then object 1 in HDD5. The required head rotation frequency for reading from object 3 to object 2 to object 1 is 3 times, including: 1 head rotation caused by reading object 2 first and then object 1 in HDD1, 1 head rotation caused by reading object 3 first and then object 1 in HDD2, and 1 head rotation caused by reading object 3 first and then object 1 in HDD5.
[0083] Therefore, to maximize sequential access to the magnetic storage module, the candidate read order with the lowest head rotation frequency can be determined as the target read order, thereby maximizing sequential access to the magnetic storage module. In this embodiment, by minimizing the head rotation frequency to determine the target read order for multiple objects, the randomness of reading multiple objects can be minimized, thus maximizing sequential access to the magnetic storage module.
[0084] In other embodiments, the head rotation parameter may include the head rotation distance. Based on the working principle of the magnetic storage module, the magnetic head rotates in the same direction to read data. Compared to random data reading, this shortens the head rotation distance, thereby reducing data read latency. Therefore, by minimizing the head movement distance, the goal of maximizing sequential access to the magnetic storage module can be achieved. Based on this, multiple objects are permuted to obtain multiple candidate read orders. Then, based on the distribution of data fragments of the multiple objects in the address mapping table, the required head movement distance for reading each of the multiple objects according to the multiple candidate read orders can be determined.
[0085] Specifically, for any magnetic storage module C, the head movement distance required to read the data fragments in that magnetic storage module C according to each candidate read order can be determined based on the offset addresses of the data fragments stored in that magnetic storage module. For example, for HDD1 shown in Table 3, if the read order of object 1 and object 2 in the candidate read order is: object 1 first, then object 2, then the head will first rotate to offset address 10MB, and then rotate from offset address 10MB to offset address 20MB, with a head rotation distance of 20MB. If the read order of object 1 and object 2 in the candidate read order is: object 2 first, then object 1, then the head will first rotate to offset address 20MB, and then return from offset address 20MB to offset address 10MB, with a head rotation distance of 30MB. Using the same method, the head rotation distance required to read the data fragments in each magnetic storage module according to each candidate read order can be determined. Subsequently, for any candidate read order, the sum of the head rotation distances required to read data fragments from multiple magnetic storage modules in that candidate read order can be calculated, thus obtaining the head rotation distance required to read multiple objects in that candidate read order. Using the same method, the head rotation distance required to read multiple objects in each candidate read order can be obtained.
[0086] Subsequently, the target read order corresponding to the shortest head movement distance can be determined from multiple candidate read orders to maximize sequential access to the magnetic storage module. In this embodiment, the aforementioned process of generating an address mapping table for multiple objects can be omitted. The target read order of multiple objects can be determined directly by minimizing the head rotation distance, which can speed up the sorting process and minimize the randomness of reading multiple objects, thereby maximizing sequential access to the magnetic storage module.
[0087] For example, regarding the offset addresses of the data fragments of Object 1, Object 2, and Object 3 in the corresponding magnetic storage modules shown in Table 3 above, we can obtain the following: According to the candidate read order of Object 1 -> Object 2 -> Object 3, the head of HDD1 first moves to offset address 10MB, and then moves from offset address 10MB to position 20MB, that is, the head moves a distance of 20MB of data; the head of HDD2 first moves to offset address 100MB, and then moves from offset address 100MB to position 300MB, that is, the head moves a distance of 300MB of data; the head of HDD4 first moves to offset address 1000MB, and then returns from offset address 1000MB to position 200MB, that is, the head moves a distance of 1800MB of data. Using the same method, we can determine the data length required for head movement when reading data from HDD5 in the candidate read order of object 1 -> object 2 -> object 3: 1000 MB; from HDD6: 300 MB; from HDD7: 2000 MB; and from HDD100: 3700 MB. Therefore, if we read multiple objects in the candidate read order of object 1 -> object 2 -> object 3, the total required head movement distance is 9120 MB.
[0088] Using the same method, it can be determined that the required head movement distance for reading multiple objects in the candidate read order of object 1 -> object 3 -> object 2 is 7420 MB. The required head movement distance for reading multiple objects in the candidate read order of object 2 -> object 1 -> object 3 is 8330 MB. The required head movement distance for reading multiple objects in the candidate read order of object 2 -> object 3 -> object 1 is 9500 MB. The required head movement distance for reading multiple objects in the candidate read order of object 3 -> object 1 -> object 2 is 8590 MB. The required head movement distance for reading multiple objects in the candidate read order of object 3 -> object 2 -> object 1 is 7800 MB. Therefore, the shortest head movement distance is the required head movement distance (i.e., 7420 MB) for reading multiple objects in the candidate read order of object 1 -> object 3 -> object 2. Therefore, the candidate reading order of object 1 -> object 3 -> object 2 can be determined, which is the target reading order.
[0089] In some embodiments of this disclosure, it is also possible to determine the head rotation parameters required for reading multiple objects according to multiple candidate read orders, without relying on the address mapping table of data fragments of multiple objects, and directly based on the arrangement order of the data fragments stored in each of the multiple magnetic storage modules. With the goal of maximizing sequential access to the magnetic storage modules, the target read order is determined from the multiple candidate read orders based on the head rotation parameters required for reading multiple objects. For a detailed implementation of determining the target read order from multiple candidate read orders based on the head rotation parameters, please refer to the relevant content of the foregoing embodiments, which will not be repeated here. This implementation eliminates the aforementioned process of generating an address mapping table for multiple objects, directly determining the target read order of multiple objects through the head rotation parameters, which can speed up the sorting process and minimize the randomness of reading multiple objects, thereby maximizing sequential access to the magnetic storage modules.
[0090] In other embodiments, based on the working principle of the magnetic storage module, the data read latency of sequential access to the magnetic storage module is lower than that of random access. Therefore, the goal of maximizing sequential access to the magnetic storage module can be achieved by minimizing the data read time. Specifically, multiple objects can be fully permuted to obtain multiple candidate read orders for the multiple objects. Then, based on the offset addresses of the data fragments stored in each of the multiple magnetic storage modules, the read time required to read each of the multiple objects according to the multiple candidate read orders can be predicted.
[0091] In this disclosure, the specific implementation of predicting the read time required to read multiple objects in multiple candidate read orders based on the offset addresses of the data fragments stored in each of the multiple magnetic storage modules is not limited. In some embodiments, the time required for the magnetic storage module's head to rotate per unit data length can be pre-tested and preset. Then, the head rotation distance required to read multiple objects in multiple candidate read orders can be determined based on the offset addresses of the data fragments stored in each of the multiple magnetic storage modules. For details on determining the head rotation distance required to read multiple objects in multiple candidate read orders, please refer to the relevant content in the foregoing embodiments. Then, the product of the head rotation distance required to read multiple objects in multiple candidate read orders and the time required for the head to rotate per unit data length can be determined as the read time required to read multiple objects in multiple candidate read orders.
[0092] The rotation speed of the read / write head of a magnetic storage module is related to the operating status of the electronic device in which it resides. Generally, the head rotation speed is faster when the electronic device is operating well. Therefore, it is possible to obtain operating status information for multiple electronic devices housing magnetic storage modules. This operating status information may include processor resource status information and memory status information, etc. For the same magnetic storage module, the larger the memory space and the more processor cores in the electronic device, the faster the magnetic storage module can read data.
[0093] Based on this, the data read paths corresponding to multiple candidate read sequences can be determined according to the storage address information of the data fragments stored in each of the multiple magnetic storage modules. The data read path corresponding to the candidate read sequence refers to the path to access the magnetic storage module when reading multiple objects according to the candidate read sequence. For example, for objects 1-3 in Table 3, if the candidate read sequence is to read object 1 first, then object 2, and finally object 3, then the data read path corresponding to this candidate read sequence is: offset 10MB from HDD1 → offset 10MB from HDD2 → offset 1000MB from HDD4 → offset 30MB from HDD5 → offset 20MB from HDD1 → offset 200MB from HDD4 → offset 2000MB from HDD7 → offset 2000MB from HDD100 → offset 300MB from HDD2 → offset 1000MB from HDD5 → offset 300MB from HDD6 → offset 300MB from HDD4.
[0094] Subsequently, the data reading paths corresponding to each of the multiple candidate reading sequences and the operating status information of the electronic devices containing the multiple magnetic storage modules can be input into a pre-built reading time prediction model. The reading time prediction model can be a pre-trained neural network model or a pre-built function module that uses the data reading path and the operating status information of the electronic device as independent variables and the reading time as the dependent variable.
[0095] Furthermore, a pre-built read time prediction model can be used to predict the read time corresponding to each of the multiple candidate read sequences based on the data read path and operating status information. In determining the read time corresponding to each of the multiple candidate read sequences, the operating status of the electronic device containing the magnetic storage module is taken into account, which helps to improve the accuracy of the predicted read time corresponding to each of the multiple candidate read sequences.
[0096] Furthermore, a target read order corresponding to the minimum read time can be determined from multiple candidate read orders to maximize sequential access to the magnetic storage module. This implementation achieves the goal of maximizing sequential access to the magnetic storage module by minimizing data read time. Therefore, when reading multiple objects according to the target read order determined in this way, data time can also be minimized, thereby reducing data read latency.
[0097] The implementation methods for maximizing sequential access to the magnetic storage module shown in the foregoing embodiments are merely illustrative and do not constitute a limitation. Referring to the data SCAN optimization process shown in Figures 2 and 3, the aforementioned step 103 and its specific implementation can be completed by the "optimizer" in Figures 2 and 3. Figures 2 and 3 illustrate only an implementation method where the optimizer, aiming to maximize sequential access to the magnetic storage module, sorts multiple objects according to the distribution of data fragments in the address mapping table to obtain the target read order, but this is not a limitation. That is, the optimizer can also use other methods shown in the foregoing embodiments to sort multiple objects. As shown in Figures 2 and 3, the input to the optimizer can be the address mapping table shown in Table 3 of the foregoing embodiments, and the output result is the target read order after sorting multiple objects, such as "object 1 -> object 3 -> object 2," that is, first read object 1, then read object 3, and finally read object 2.
[0098] After sorting the reading order of multiple objects by maximizing sequential access to the magnetic storage module to obtain the target reading order, in step 104, multiple objects can be read from multiple magnetic storage modules according to the target reading order, the storage address information and data length of the multiple data fragments corresponding to each object, thereby achieving maximized sequential access to the magnetic storage module. Specifically, multiple objects can be read from multiple magnetic storage modules according to the target reading order, based on the storage address information and data length of the multiple data fragments corresponding to each object. For any object i, data of the corresponding data length can be read starting from the storage address information of the multiple data fragments of object i corresponding to the multiple magnetic storage modules to obtain multiple data fragments of object i; then, the multiple data fragments of object i can be concatenated according to the ascending order of the identifiers of the multiple data fragments of object i to obtain object i. Using the same method, all objects can be read from the magnetic storage module according to the target reading order, achieving maximized sequential access to the magnetic storage module.
[0099] In some embodiments of this disclosure, for multiple objects to be read, the reading order of the multiple objects is sorted by maximizing sequential access to the magnetic storage module to obtain a target reading order; and multiple objects are read from multiple magnetic storage modules according to the target reading order. Maximizing sequential access to the magnetic storage module can reduce the randomness of random access to the magnetic storage module, thereby helping to reduce the access latency of data reading and improve data reading efficiency.
[0100] In this disclosure, the application scenarios of the data reading method provided in the foregoing embodiments are not limited. In some embodiments, the data reading method provided in this disclosure is applied to AI training scenarios. Accordingly, in response to a data reading request, multiple objects required for training a neural network model can be identified as multiple objects to be read. Further, the storage address information and data length of the multiple objects can be obtained from the metadata storage node of the distributed storage system. Further, the multiple objects can be read according to the data reading method provided in the foregoing embodiments; subsequently, the multiple objects can be used to train the neural network model.
[0101] Specifically, the dataset can be divided into training, test, and validation sets. The training set is the subset of data used to train the model. The model learns from samples in the training set, adjusting its internal parameters to minimize the loss function, thus improving its fit to the training data. The validation set is another subset separated from the training set. It is used to evaluate the model's performance on unseen data and helps tune hyperparameters. The test set is a third subset of data, independent of the training and validation sets. It is used to ultimately evaluate the model's generalization ability, i.e., its performance on completely unseen data.
[0102] Next, the neural network model can be trained using the training set to obtain the first neural network model; then, the hyperparameters of the first neural network model can be optimized using the validation set to obtain the target neural network model. Finally, the performance of the target neural network model can be evaluated using the test set to complete the model training.
[0103] In other embodiments, the data reading method provided in this disclosure is applied to big data analysis scenarios. Accordingly, in response to a data reading request, multiple objects required for data analysis can be identified as objects to be read. Further, the storage address information and data length of the multiple objects can be obtained from the metadata storage nodes of the distributed storage system. Further, the multiple objects can be read according to the data reading method provided in the foregoing embodiments; subsequently, the multiple objects can be used for big data analysis, etc. For example, the multiple objects can be used to analyze the development trend of a target or predict the development trend of a target. Alternatively, the multiple objects can be used to trace back or restore a target, etc.
[0104] The inventors of this disclosure have found through testing that using the data reading mechanism provided in the embodiments of this disclosure to collect data in data SCAN can improve the efficiency of AI scanning data. For embodiments using a Graphics Processing Unit (GPU) for AI training, it can save GPU time and reduce GPU resource usage costs.
[0105] It should be noted that the execution subject of each step of the method provided in the above embodiments can be the same device, or the method can be executed by different devices. For example, the execution subject of steps 101 and 102 can be device A; or the execution subject of step 101 can be device A, and the execution subject of step 102 can be device B; and so on.
[0106] Furthermore, some processes described in the above embodiments and accompanying drawings include multiple operations that appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or they may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel.
[0107] Accordingly, this disclosure also provides a computer-readable storage medium storing computer instructions, which, when executed by one or more processors, cause one or more processors to perform the steps in the data reading methods provided in the foregoing embodiments.
[0108] Computer-readable storage media include volatile or non-volatile or a combination thereof, and may be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital video disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transfer medium.
[0109] This disclosure also includes a computer program product comprising a computer program that, when executed by one or more processors, causes the one or more processors to perform the steps in the data reading methods provided in the foregoing embodiments.
[0110] In this disclosure, the specific implementation form of the computer program product is not limited. In some embodiments, the computer program product may be implemented as an application (APP), a mini-program, a computer-side client, a program module, a plug-in, an installation package, a software development kit (SDK), an image file of an optical disc (such as an ISO file), a plug-in, or software in the form of Software as a Service (SaaS), etc., but is not limited thereto.
[0111] The computer program product should understand that each or a combination of the above-described method flow can be implemented by a computer program or instructions. Furthermore, these computer programs or instructions can be applied to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device, enabling the processor of the general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to function as an apparatus for implementing the corresponding functions in the above-described method embodiments.
[0112] Figure 4 is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. As shown in Figure 4, the electronic device includes a memory 40a and a processor 40b. The memory 40a is used to store computer programs and can be configured to store various other data to support operation on a computing platform. Examples of this data include instructions for any application or method operating on the electronic device, data structures, contact data, phonebook data, messages, pictures, videos, etc.
[0113] The processor 40b is coupled to the memory 40a and is used to execute a computer program to perform the steps in the data reading methods provided in the foregoing embodiments. Specific implementation details of each step can be found in the relevant descriptions of the foregoing embodiments, and will not be repeated here.
[0114] In some alternative embodiments, as shown in FIG4, the electronic device may further include optional components such as a communication component 40c, a power supply component 40d, a display component 40e, and an audio component 40f. FIG4 only schematically shows some components and does not mean that the electronic device must include all the components shown in FIG4, nor does it mean that the electronic device can only include the components shown in FIG4.
[0115] Furthermore, the components within the dashed boxes in Figure 4 are optional, not mandatory, and their specific requirements depend on the product form of the electronic device. The electronic device in this embodiment can be a desktop computer, laptop computer, mobile phone, or IoT device; it can also be a traditional server, cloud server, or server cluster, or other server equipment.
[0116] In embodiments of this disclosure, the memory is used to store computer programs and can be configured to store various other data to support operation on its host device. The processor can execute the computer programs stored in the memory to implement corresponding control logic. The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), electrically erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0117] In this embodiment of the disclosure, the processor can be any hardware processing device capable of executing the above-described method logic. Optionally, the processor can be a central processing unit (CPU), a graphics processing unit (GPU), or a microcontroller unit (MCU); it can also be a programmable device such as a field-programmable gate array (FPGA), a programmable array logic (PAL), a general array logic (GAL), or a complex programmable logic device (CPLD); or it can be an advanced RISC machine (ARM) or a system on chip (SoC), etc., but is not limited thereto.
[0118] In this embodiment of the disclosure, the communication component is configured to facilitate wired or wireless communication between its host device and other devices. The device hosting the communication component can access wireless networks based on communication standards, such as 2G or 3G, 4G, 5G, or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel.
[0119] In embodiments of this disclosure, the display component may include a liquid crystal display (LCD) and a touch panel (TP). If the display component includes a touch panel, the display component may be implemented as a touchscreen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation.
[0120] In embodiments of this disclosure, a power supply component is configured to provide power to various components of the device in which it resides. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply component resides.
[0121] In embodiments of this disclosure, the audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals can be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals. For example, in devices with voice interaction capabilities, voice interaction with a user can be achieved through the audio component.
[0122] It should be noted that the terms "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a chronological order, nor do they limit "first" and "second" to different types.
[0123] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the aforementioned element.
[0124] The above description is merely an embodiment of this disclosure and is not intended to limit the scope of this disclosure. Various modifications and variations can be made to this disclosure by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of the claims of this disclosure.
Claims
1. A data reading method, wherein, include: In response to a data read request, obtain the storage address information and data length of the multiple objects to be read; The storage address information of the multiple objects includes: the storage address information of multiple data fragments corresponding to each of the multiple objects; the multiple data fragments corresponding to each object are stored on multiple magnetic storage modules; Based on the storage address information of the multiple data fragments corresponding to each of the multiple objects, the storage address information of the data fragments stored by each of the multiple magnetic storage modules is determined; With the goal of maximizing sequential access to the magnetic storage modules, the multiple objects are sorted according to the storage address information of the data fragments stored in each of the multiple magnetic storage modules to obtain the target reading order; Based on the target reading order, the storage address information of the multiple data fragments corresponding to each of the multiple objects, and the data length, the multiple objects are read from the multiple magnetic storage modules.
2. The method of claim 1, wherein, The storage address information of the data fragment includes: the offset address of the data fragment in the magnetic storage module; the step of sorting the multiple objects according to the storage address information of the data fragments stored in each of the multiple magnetic storage modules to obtain the target reading order, with the goal of maximizing sequential access to the magnetic storage modules, includes: For any one of the plurality of magnetic storage modules, the data fragments stored in any one magnetic storage module are sorted according to the size relationship of the offset addresses of the data fragments stored in any one magnetic storage module, so as to obtain the arrangement order of the data fragments stored in any one magnetic storage module; With the goal of maximizing sequential access to the magnetic storage modules, the multiple objects are sorted according to the arrangement order of the data fragments stored in each of the multiple magnetic storage modules to obtain the target reading order.
3. The method of claim 2, wherein, The goal of maximizing sequential access to the magnetic storage modules involves sorting the multiple objects according to the arrangement order of the data fragments stored in each of the multiple magnetic storage modules to obtain the target reading order, including: For any magnetic storage module, according to the arrangement order of the data fragments stored in the magnetic storage module, the offset address of the data fragments stored in the magnetic storage module is written into the column corresponding to the magnetic storage module in the address mapping table to obtain the address mapping table of the multiple objects; With the goal of maximizing sequential access to the magnetic storage module, the multiple objects are sorted according to the distribution of their data fragments in the address mapping table to obtain the target reading order.
4. The method of claim 3, wherein, The distribution of the data fragments of the multiple objects in the address mapping table includes: the number of data fragments of the multiple objects in each row of the address mapping table; the sorting of the multiple objects according to the distribution of the data fragments of the multiple objects in the address mapping table with the goal of maximizing sequential access to the magnetic storage module includes: If multiple target objects exist among the multiple objects for which the reading order is not determined, the following steps are executed repeatedly until a set loop stop condition is met, and the target reading order is determined based on the reading order of the multiple objects determined at the time of loop stop; wherein the repeatedly executed steps include: From the rows that have not yet been traversed in the address mapping table, determine the target row with the smallest row number; The multiple target objects are sorted according to the number of data fragments of the multiple target objects in the target row from largest to smallest, so as to maximize sequential access to the magnetic storage module.
5. The method of claim 4, wherein, Determining the target reading order based on the reading order of the multiple objects determined when the loop stops includes: If the loop stops when the reading order of the multiple objects is determined, then the reading order of the multiple objects determined when the loop stops is determined as the target reading order; or, If the loop stops after all rows of the address mapping table have been traversed, and if there are multiple first objects among the multiple objects whose reading order has not been determined when the loop stops, then the multiple first objects are permuted to obtain multiple reading orders for the multiple first objects; multiple target columns where the data fragments of the multiple first objects are located are determined from the address mapping table; based on the distribution of the objects to which the data fragments stored in the multiple target columns belong and the reading order of the second objects, the required head rotation frequency for reading each of the multiple first objects according to the multiple reading orders is determined; and the target reading order is determined based on the reading order with the lowest head rotation frequency and the reading order of the second objects. The second object is the object whose reading order is determined among the plurality of objects when the loop stops.
6. The method according to any one of claims 3-5, wherein, The method further includes: Perform a full permutation of the multiple objects to obtain multiple candidate reading orders for the multiple objects; The goal of maximizing sequential access to the magnetic storage module involves sorting the multiple objects according to their data fragments in the address mapping table to obtain the target read order, including: Based on the distribution of the data fragments of the multiple objects in the address mapping table, determine the head rotation parameters required for each of the multiple objects to be read in the order of the multiple candidate reads; With the goal of maximizing sequential access to the magnetic storage module, the target read order is determined from the plurality of candidate read orders based on the magnetic head rotation parameters.
7. The method of claim 6, wherein, The head rotation parameters include: head rotation frequency and / or head rotation distance; the step of determining the target read order from the plurality of candidate read orders based on the head rotation parameters, with the goal of maximizing sequential access to the magnetic storage module, includes: From the multiple candidate read orders, determine the target read order corresponding to the minimum head rotation frequency to maximize sequential access to the magnetic storage module; or, From the multiple candidate read orders, the target read order corresponding to the shortest head movement distance is determined to maximize sequential access to the magnetic storage module.
8. The method of claim 2, wherein, The method further includes: Perform a full permutation of the multiple objects to obtain multiple candidate reading orders for the multiple objects; The goal of maximizing sequential access to the magnetic storage modules involves sorting the multiple objects according to the arrangement order of the data fragments stored in each of the multiple magnetic storage modules to obtain the target reading order, including: Based on the arrangement order of the data fragments stored in each of the multiple magnetic storage modules, determine the required magnetic head rotation parameters for reading each of the multiple objects according to the multiple candidate reading order; With the goal of maximizing sequential access to the magnetic storage module, the target read order is determined from the plurality of candidate read orders based on the magnetic head rotation parameters.
9. The method of claim 1, wherein, The method further includes: Perform a full permutation of the multiple objects to obtain multiple candidate reading orders for the multiple objects; The goal of maximizing sequential access to the magnetic storage modules involves sorting the multiple objects according to the storage address information of the data fragments stored in each of the multiple magnetic storage modules to obtain the target reading order, including: Based on the storage address information of the data fragments stored in each of the multiple magnetic storage modules, predict the reading time required to read each of the multiple objects in the order of the multiple candidate reads; From the multiple candidate read orders, a target read order corresponding to the minimum read time is determined to maximize sequential access to the magnetic storage module.
10. The method of claim 9, wherein, Also includes: Obtain the operating status information of the electronic device where the plurality of magnetic storage modules are located; The step of predicting the required read time for each of the multiple objects to be read in the order of the multiple candidate reads, based on the storage address information of the data fragments stored in each of the multiple magnetic storage modules, includes: Based on the storage address information of the data fragments stored in each of the multiple magnetic storage modules, the data reading path corresponding to each of the multiple candidate reading orders is determined; Input the data reading path and the running status information into the pre-built reading time prediction model; Using a pre-built read time prediction model, the read time corresponding to each of the multiple candidate read sequences is predicted based on the data read path and the running status information.
11. The method of any one of claims 1-10, wherein, The step of responding to a data read request by obtaining the storage address information and data length of multiple objects to be read includes: In response to the data reading request, the objects required for training the neural network model are determined to be the plurality of objects; The storage address information and data length of the multiple objects are obtained from the metadata storage nodes of the distributed storage system. as well as, The method further includes: The neural network model is trained using the multiple objects.
12. The method of claim 11, wherein, The step of training the neural network model using the multiple objects includes: The multiple objects are divided into a training set, a validation set, and a test set; The neural network model is trained using the training set to obtain a first neural network model; The first neural network model is optimized using the validation set to obtain the target neural network model; The target neural network model is evaluated using the test set to complete model training.
13. The method of any one of claims 1-12, wherein, The object includes structured data, semi-structured data, and / or unstructured data; the object is divided into multiple data fragments.
14. The method of any one of claims 1-13, wherein, The object also includes metadata and application data. The metadata and application data are stored independently to achieve separation of control flow and data flow. The metadata includes the storage address information and data length of the application data.
15. The method of any one of claims 1-14, wherein, The storage address information of the data fragment also includes the identifier of the magnetic storage module where the data fragment is located; the identifier of the magnetic storage module where the data fragment is located is used to determine the magnetic storage module to which the data fragment belongs, and the offset address is used to determine the specific location of the data fragment in the corresponding magnetic storage module.
16. The method of any one of claims 1-15, wherein, The step of responding to a data read request by obtaining the storage address information and data length of multiple objects to be read includes: In response to the data read request, the identifiers of the plurality of objects are obtained from the data read request; Based on the identifiers of the multiple objects, the storage address information and corresponding data length of the data shards corresponding to each of the multiple objects are obtained from the metadata storage nodes of the distributed storage system.
17. The method of any one of claims 1-16, wherein, The step of reading the multiple objects from the multiple magnetic storage modules according to the target reading order, the storage address information of the multiple data fragments corresponding to each of the multiple objects, and the data length includes: Following the target reading order, read multiple data fragments corresponding to each object sequentially; For each object, the multiple data fragments are concatenated according to the ascending order of their identifiers to obtain the object.
18. An electronic device, comprising: include: A memory and a processor; wherein the memory is used to store computer programs; The processor is coupled to the memory for executing the computer program to perform the steps of the method according to any one of claims 1-17.
19. A computer readable storage medium having stored thereon computer instructions, wherein, When the computer instructions are executed by one or more processors, the one or more processors are caused to perform the steps of the method according to any one of claims 1-17.
20. A computer program product, wherein, Includes a computer program that, when executed by one or more processors, causes the one or more processors to perform the steps of the method according to any one of claims 1-17.