Method, electronic device, and computer program product for storage management
By simultaneously determining the verification information and identification information of the storage page in a data loading cycle in storage management, the problem of inefficiency in traditional methods is solved, and the improvement of CPU efficiency and IO performance is achieved.
Patent Information
- Application Number
- CN202110014398.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-06
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-01-06
AI Technical Summary
In the traditional dirty data refresh preprocessing process, cyclic redundancy checksum non-cryptographic hash calculations are performed independently, resulting in inefficient data loading and traversing processes, excessive CPU load, and poor IO performance.
In a data loading cycle, the verification information and identification information of the storage page are determined simultaneously, avoiding two cycles of traversals, reducing data loading, improving CPU efficiency and improving IO performance.
By halving the data load, reduce CPU load, save CPU cycles, improve CPU efficiency and improve IO performance.
Smart Images

Figure CN114721586B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure generally relate to storage management, and more particularly, to methods, electronic devices, and computer program products for storage management. Background Art
[0002] To improve data processing efficiency, data in a persistent storage device such as a RAID (Redundant Arrays of Independent Disks) is loaded into a memory for use during task execution. During task execution, the data in the memory may be temporarily updated while the corresponding data in the persistent storage device has not been updated yet. In this case, the temporarily updated data can be referred to as dirty data. The dirty data will be flushed to the persistent storage device to update the corresponding data in the persistent storage device. Before the dirty data is flushed, preprocessing needs to be performed. However, the traditional preprocessing process is inefficient. Summary of the Invention
[0003] Embodiments of the present disclosure provide methods, electronic devices, and computer program products for storage management.
[0004] In a first aspect of the present disclosure, a method for storage management is provided. The method includes: obtaining target data in a target storage page in a memory; determining, based on the target data, check information and identification information associated with the target data, the check information being used to verify whether the target data is correct, and the identification information being used to identify the target data; and determining, based on the identification information, storage information associated with the target data and the check information, the storage information indicating whether to store the target data and the check information into a persistent storage device.
[0005] In a second aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit and at least one memory. The at least one memory is coupled to the at least one processing unit and stores instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the device to perform actions, the actions including: obtaining target data in a target storage page in a memory; determining, based on the target data, check information and identification information associated with the target data, the check information being used to verify whether the target data is correct, and the identification information being used to identify the target data; and determining, based on the identification information, storage information associated with the target data and the check information, the storage information indicating whether to store the target data and the check information into a persistent storage device.
[0006] In a third aspect of the present disclosure, a computer program product is provided. The computer program product is tangibly stored on a non-transitory computer-readable medium and includes machine-executable instructions that, when executed, cause the machine to implement any of the steps of the method described in the first aspect of the present disclosure.
[0007] The Summary of the Invention section is provided to introduce a selection of concepts in a simplified form, which will be further described in the Detailed Description below. The Summary of the Invention section is not intended to identify the key features or essential features of the present disclosure, nor is it intended to limit the scope of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] By describing the exemplary embodiments of the present disclosure in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present disclosure will become more apparent, wherein, in the exemplary embodiments of the present disclosure, the same reference numerals generally represent the same components.
[0009] Figure 1 A schematic diagram showing an example of a storage management environment in which some embodiments of the present disclosure can be implemented;
[0010] Figure 2 A flowchart showing an example of a method for storage management according to some embodiments of the present disclosure;
[0011] Figure 3 A flowchart showing an example of a method for determining storage information according to some embodiments of the present disclosure; and
[0012] Figure 4 A schematic block diagram showing an example device that can be used to implement the embodiments of the present disclosure.
[0013] In the respective drawings, the same or corresponding reference numerals denote the same or corresponding parts. DETAILED DESCRIPTION
[0014] The preferred embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Instead, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0015] As used herein, the term "comprising" and its variations mean open-ended inclusion, i.e., "including but not limited to". Unless specifically stated otherwise, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "an example embodiment" and "an embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter.
[0016] As described above, preprocessing needs to be performed before dirty data is flushed. In some cases, the data in the memory can be organized in units of storage pages (pages). Each storage page has a predetermined size, such as 4 kilobytes (4KB). In this case, the dirty data can be referred to as dirty storage page data. During the preprocessing of the dirty storage page data, a cyclic redundancy check value needs to be calculated for verifying data consistency, and then a non-cryptographic hash value needs to be calculated to find data in the persistent storage device that duplicates the dirty storage page data.
[0017] Traditionally, the calculation of the cyclic redundancy check value and the non-cryptographic hash value is independent. In the cyclic redundancy check, 8 bytes (8B) of dirty storage page data need to be loaded from the memory to the register each time, and the dirty storage page data is traversed in a loop for cyclic redundancy check calculation. In the non-cryptographic hash calculation, the same process needs to be gone through again. That is, in the non-cryptographic hash calculation, 8B of dirty storage page data also need to be loaded from the memory to the register each time, and the dirty storage page data is traversed in a loop for non-cryptographic hash calculation. For a 4KB storage page, a total of 4096 / 8 * 2 = 1024 times of loading are required.
[0018] According to an example embodiment of the present disclosure, an improved solution for storage management is proposed. In this solution, the target data in the target storage page in the memory can be obtained. The check information and identification information associated with the target data can be determined based on the target data. The check information is used to verify whether the target data is correct. The identification information is used to identify the target data. Thus, based on the identification information, the storage information associated with the target data and the check information can be determined. The storage information indicates whether the target data and the check information are to be stored in the persistent storage device.
[0019] In this way, in the preprocessing process of the dirty storage page data, the proposed solution can determine the verification information and the identification information simultaneously in one data loading loop, so as to avoid traversing the dirty storage page data twice for the verification information and the identification information. Since a single loop is used instead of two loops for storage page refreshing, the amount of data to be loaded can be halved, thereby reducing the CPU (Center Processing Unit) load, saving CPU cycles, improving CPU efficiency, and improving IO (Input / Output) performance. The embodiments of the present disclosure will be specifically described below with reference to the accompanying drawings.
[0020] Figure 1 FIG. shows a schematic diagram of an example of a storage management environment 100 in which some embodiments of the present disclosure can be implemented. The storage management environment 100 includes a processor 110, a memory 120, and a persistent storage device 130. As an example, the processor 110 can be any device with computing capabilities. For example, the processor 110 can be a processor of a personal computer, a tablet computer, a wearable device, a cloud server, a mainframe, a distributed computing system, etc.
[0021] The memory 120 can be any device with storage capabilities. For example, the memory can be a volatile memory such as SDRAM (Synchronous Dynamic random access memory) and DDR SDRAM (Double Data Rate Synchronous Dynamic random access memory). Similarly, the persistent storage device 130 can be any device with storage capabilities. Different from the memory 120, the persistent storage device 130 can be a non-volatile memory such as a disk, a solid-state drive, and a RAID.
[0022] In the storage management environment 100, the processor 110 is configured to perform storage management. As described above, the data in the memory 120 can be organized in units of storage pages (pages). Each storage page has a predetermined size, such as 4 kilobytes (4KB). During the execution of a task, the data in some storage pages in the memory 120 may be temporarily updated, while the corresponding data in the persistent storage device 130 has not been updated. In this case, the temporarily updated data can be referred to as dirty storage page data. Such dirty storage page data can be used as the target data 150, and the target data 150 will be refreshed to the persistent storage device 130 to update the corresponding data in the persistent storage device 130.
[0023] Before refreshing, the processor 110 may perform preprocessing. Specifically, during the preprocessing, the processor 110 may simultaneously determine the verification information 160 and the identification information 170 of the target data 150 in one data loading loop. The verification information 160 is used to verify whether the target data 150 is correct. For example, in the case where the target data is stored in the persistent storage device 130, the verification information 160 may be used to verify the consistency of the target data 150 when reading the target data 150 from the persistent storage device 130 in the future. The identification information 170 is used to identify the target data. For example, the identification information 170 may be used to find out whether the target data 150 has been stored in the persistent storage device 130. Thus, since the verification information 160 and the identification information 170 are simultaneously determined in one data loading loop, it is possible to avoid traversing the target data 150 twice for the verification information 160 and the identification information 170.
[0024] In some embodiments, the processor 110 includes a register 140. In this case, the processor 110 may load the target data 150 from the memory 120 into the register 140, and determine the verification information 160 and the identification information 170 based on the target data 150. Further, the processor 110 may determine the storage information 180 based on the identification information 160. The storage information 180 indicates whether to store the target data 150 and the verification information 160 in the persistent storage device 130, that is, whether to refresh the target data 150 and the verification information 160 to the persistent storage device 130. Hereinafter, reference will be made to Figure 2 and Figure 3 to describe in detail the storage management operations performed by the processor 110.
[0025] Figure 2 The flowchart of a method 200 for storage management according to some embodiments of the present disclosure is shown. The method 200 may be implemented by a processor 110 as shown in Figure 1 Alternatively, the method 200 may also be implemented by other entities other than the processor 110. It should be understood that the method 200 may further include additional steps not shown and / or steps shown may be omitted, and the scope of the present disclosure is not limited in this regard.
[0026] At 210, the processor 110 acquires the target data 150 in the target storage page in the memory 120. In some embodiments, refreshing is only started when the number of storage pages with dirty data exceeds a threshold number (for example, 2048 storage pages). In view of this, the processor 110 may determine whether the number of at least one candidate storage page in the memory 120 exceeds the threshold number. For example, these candidate storage pages may be storage pages with dirty data. If the number exceeds the threshold number, the processor 120 may determine the target storage page from these candidate storage pages.
[0027] In addition, as described above, in some embodiments, the processor 110 includes a register 140. In this case, the processor 110 may load the target data 150 from the memory 120 into the register 140 to obtain the target data 150. In some embodiments, since the storage space size of the register 140 is generally smaller than the size of the target data 150, multiple parts of the target data 150 may be separately loaded from the memory 120 into the register 140. For example, the size of the target data 150 may be 4KB. In this case, a part of the target data 150 (e.g., 8B) may be loaded from the memory 120 into the register 140 in each loop until the target data 150 is traversed.
[0028] At 220, the processor 110 determines the check information 160 and the identification information 170 associated with the target data 150 based on the target data 150. The check information 160 is used to verify whether the target data 150 is correct. In some embodiments, the processor 110 may perform a check operation on the target data 150 to determine the check information 160. In this case, the check information 160 may be a check value obtained after performing a check operation on the target data 150. For example, the check operation may include parity check operation, CRC (Cyclic Redundancy Check) operation, and BCC (Block Check Character) exclusive OR check operation, etc. In some embodiments, the processor 110 may also allocate a buffer for the check information 160 to store the check information 160.
[0029] The identification information 170 is used to identify the target data 150. In some embodiments, the processor 110 may perform a hash operation on the target data 150 to determine the identification information 170. In this case, the identification information 170 may be a hash value obtained after performing a hash operation on the target data 150. For example, the hash operation may include non - cryptographic hash operations such as MurmurHash, xxHash, SipHash, etc. Non - cryptographic hash operations are suitable for hash - based lookups. Different from cryptographic hash operations, non - cryptographic hash operations are not specifically designed to be difficult to reverse, making them not suitable for encryption purposes.
[0030] In some embodiments, when multiple parts of the target data 150 are separately loaded from the memory 120 into the register 140, the processor 110 may separately perform a verification operation on the multiple parts of the target data 150 to generate a first set of intermediate values, and determine the verification information 160 based on the first set of intermediate values. When a buffer is allocated, the first set of intermediate values may also be stored in the buffer. Similarly, the processor 110 may separately perform a hashing operation on the multiple parts of the target data 150 to generate a second set of intermediate values, and determine the identification information 170 based on the second set of intermediate values.
[0031] For example, when a part of the target data 150 is loaded from the memory 120 into the register 140, the processor 110 may perform a verification operation on this part to generate an intermediate value associated with the verification information 160, and may also perform a hashing operation on this part to generate another intermediate value associated with the identification information 170. In addition, in some embodiments, the intermediate value associated with the verification information 160 may also be stored in the allocated buffer.
[0032] At 230, the processor 110 determines the storage information 180 associated with the target data 150 and the verification information 160 based on the identification information 170. The storage information 180 indicates whether to store the target data 150 and the verification information 160 into the persistent storage device 130. Hereinafter, reference will be made to Figure 3 a detailed description of the operation of the processor 110 for determining the storage information. Figure 3 A flowchart showing an example of a method 300 for determining storage information according to some embodiments of the present disclosure is shown.
[0033] In some embodiments, the processor 110 may perform deduplication before flushing. Thus, at 310, the processor 110 may obtain a set of reference identification information items associated with corresponding reference data in a set of reference storage pages that have been stored in the persistent storage device 130. At 320, the processor 110 may separately compare the identification information 170 with this set of reference identification information items.
[0034] If the identification information 170 does not match each item in this set of reference identification information items, this means that the target data 150 has not been stored in the persistent storage device 130. In this case, at 330, the processor 110 may determine the storage information 180 as indicating to store the target data 150 and the verification information 160 into the persistent storage device 130.
[0035] If the identification information 170 matches one of the set of reference identification information items, this means that the target data 150 has already been stored in the persistent storage device 130. In this case, the processor 110 may determine the storage information 180 as indicating not to store the target data 150 and the check information 160 in the persistent storage device 130. In some embodiments, the processor 110 may mark the target storage page or the reference storage page corresponding to the matching reference identification information item as a duplicate storage page. Additionally, in some embodiments, the target data 150 stored in the persistent storage device 130 may have a reference count value. In the case where it is determined that the target data 150 has already been stored in the persistent storage device 130, the processor 110 may increment the reference count value of the target data 150 to record the number of times the target data 150 has been referenced. Thereby, deduplication can be achieved.
[0036] In some embodiments, to ensure a correct determination of whether the target data 150 has already been stored in the persistent storage device 130, in the case where the identification information 170 matches one of the set of reference identification information items, at 340, the processor 110 may also compare the target data 150 with the reference data corresponding to the matching reference identification information item. If the target data 150 does not match the reference data, this means that the target data 150 has not yet been stored in the persistent storage device 130. In this case, at 350, the processor 110 may determine the storage information 180 as indicating to store the target data 150 and the check information 160 in the persistent storage device 130. If the target data 150 matches the reference data, this means that the target data 150 has already been stored in the persistent storage device 130. In this case, at 360, the processor 110 may determine the storage information 180 as indicating not to store the target data 150 and the check information 160 in the persistent storage device 130.
[0037] In some embodiments, in the case where the storage information 180 indicates to store the target data 150 and the check information 160 in the persistent storage device 130, the processor 110 may append the check information 160 to the target data 150 to obtain the data to be stored. Further, the processor 110 may store the data to be stored in the persistent storage device 130. Thereby, when reading the target data 150 from the persistent storage device 130 in the future, the check information 160 can be used to perform a consistency check on the read target data 150.
[0038] In this way, in the preprocessing process of the dirty storage page data, the proposed solution can determine the verification information and the identification information simultaneously in one data loading loop, so as to avoid traversing the dirty storage page data twice for the verification information and the identification information. Since a single loop is used instead of two loops for storage page refreshing, the amount of data to be loaded can be halved, thereby reducing the CPU load, saving CPU cycles, improving CPU efficiency, and improving the IO performance. The embodiments of the present disclosure will be specifically described below with reference to the accompanying drawings.
[0039] Figure 4 FIG. shows a schematic block diagram of an exemplary device 400 that can be used to implement embodiments of the present disclosure. For example, the processor 110 as shown can be implemented by the device 400. As shown in the figure, the device 400 includes a central processing unit (CPU) 410, which can execute various appropriate actions and processes according to computer program instructions stored in the read-only memory (ROM) 420 or computer program instructions loaded from the storage unit 480 into the random access memory (RAM) 430. In the RAM 430, various programs and data required for the operation of the device 400 can also be stored. The CPU 410, the ROM 420, and the RAM 430 are connected to each other through a bus 440. The input / output (I / O) interface 450 is also connected to the bus 440. Figure 1 As shown, the processor 110 can be implemented by the device 400. As shown in the figure, the device 400 includes a central processing unit (CPU) 410, which can execute various appropriate actions and processes according to computer program instructions stored in the read-only memory (ROM) 420 or computer program instructions loaded from the storage unit 480 into the random access memory (RAM) 430. In the RAM 430, various programs and data required for the operation of the device 400 can also be stored. The CPU 410, the ROM 420, and the RAM 430 are connected to each other through a bus 440. The input / output (I / O) interface 450 is also connected to the bus 440.
[0040] A plurality of components in the device 400 are connected to the I / O interface 450, including: an input unit 460, such as a keyboard, a mouse, etc.; an output unit 470, such as various types of displays, speakers, etc.; a storage unit 480, such as a disk, an optical disc, etc.; and a communication unit 490, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 490 allows the device 400 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0041] The various processes and processes described above, such as processes 200 and 300, can be executed by the processing unit 410. For example, in some embodiments, processes 200 and 300 can be implemented as computer software programs, which are tangibly included in a machine-readable medium, such as the storage unit 480. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 400 via the ROM 420 and / or the communication unit 490. When the computer program is loaded into the RAM 430 and executed by the CPU 410, one or more actions of the processes 200 and 300 described above can be executed.
[0042] The present disclosure may be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for performing various aspects of the present disclosure.
[0043] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in a groove having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not to be construed as a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0044] The computer-readable program instructions described herein may be downloaded to each computing / processing device from a computer-readable storage medium or may be downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0045] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Java, Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.
[0046] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.
[0047] These computer - readable program instructions can be provided to a processing unit of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that, when the instructions are executed by the processing unit of the computer or other programmable data - processing apparatus, a device is created that implements the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture that includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0048] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0049] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.
[0050] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the technical improvement of technologies in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A method for storage management, comprising: Obtaining target data in a target storage page within a memory; Based on the target data, determining check information and identification information associated with the target data, the check information being used to verify whether the target data is correct, and the identification information being used to identify the target data; And Based on the identification information, determining storage information associated with the target data and the check information, the storage information indicating whether to store the target data and the check information into a persistent storage device, wherein determining the check information and the identification information includes: Circularly loading each part of a plurality of parts of the target data from the memory into a register; For each part of the target data when loaded into the register, performing a check operation and a hash operation, the check operation generating a corresponding first intermediate value and the hash operation generating a corresponding second intermediate value, the first intermediate values being accumulated into a first set of intermediate values, and the second intermediate values being accumulated into a second set of intermediate values; and Subsequently determining the check information based on the first set of intermediate values, and determining the identification information based on the second set of intermediate values.
2. The method according to claim 1, wherein determining the storage information includes: Obtaining a set of reference identification information items associated with corresponding reference data in a set of reference storage pages that have been stored in the persistent storage device; Comparing the identification information with each of the set of reference identification information items respectively; And If the identification information matches one of the set of reference identification information items, determining the storage information as indicating not to store the target data and the check information into the persistent storage device.
3. The method according to claim 1, wherein determining the storage information includes: Obtaining a set of reference identification information items associated with corresponding reference data in a set of reference storage pages that have been stored in the persistent storage device; Comparing the identification information with each of the set of reference identification information items respectively; And If the identification information matches one of the set of reference identification information items, comparing the target data with the reference data corresponding to the matching reference identification information item; And If the target data matches the reference data, determining the storage information as indicating not to store the target data and the check information into the persistent storage device.
4. The method according to claim 1, wherein the storage information indicates to store the target data and the check information into the persistent storage device, and the method further includes: Attaching the check information to the target data to obtain data to be stored; And Storing the data to be stored into the persistent storage device.
5. The method according to claim 1, further comprising: Determining whether the number of at least one candidate storage page in the memory exceeds a threshold number; And If the number exceeds the threshold number, determining the target storage page from the at least one candidate storage page.
6. An electronic device, comprising: at least one processing unit; at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which when executed by the at least one processing unit cause the electronic device to perform actions, the actions including: obtaining target data in a target storage page within the memory; based on the target data, determining verification information and identification information associated with the target data, the verification information being used to verify whether the target data is correct, and the identification information being used to identify the target data; and based on the identification information, determining storage information associated with the target data and the verification information, the storage information indicating whether to store the target data and the verification information in a persistent storage device, wherein determining the verification information and the identification information includes: cyclically loading each part of a plurality of parts of the target data from the memory into a register; for each part of the target data when loaded into the register, performing a verification operation and a hash operation, the verification operation generating a corresponding first intermediate value and the hash operation generating a corresponding second intermediate value, the first intermediate values being accumulated as a first set of intermediate values and the second intermediate values being accumulated as a second set of intermediate values; and subsequently determining the verification information based on the first set of intermediate values and determining the identification information based on the second set of intermediate values.
7. The device according to claim 6, wherein determining the storage information includes: obtaining a set of reference identification information items associated with corresponding reference data in a set of reference storage pages that have been stored in the persistent storage device; comparing the identification information with each of the set of reference identification information items respectively; and if the identification information matches one of the set of reference identification information items, determining the storage information as indicating not to store the target data and the verification information in the persistent storage device.
8. The device according to claim 6, wherein determining the storage information includes: obtaining a set of reference identification information items associated with corresponding reference data in a set of reference storage pages that have been stored in the persistent storage device; comparing the identification information with each of the set of reference identification information items respectively; and if the identification information matches one of the set of reference identification information items, comparing the target data with the reference data corresponding to the matching reference identification information item; and if the target data matches the reference data, determining the storage information as indicating not to store the target data and the verification information in the persistent storage device.
9. The device according to claim 6, wherein the storage information indicates to store the target data and the verification information in the persistent storage device, and the actions further include: appending the verification information to the target data to obtain data to be stored; and storing the data to be stored in the persistent storage device.
10. The apparatus according to claim 6, wherein the action further comprises: determining whether the number of at least one candidate storage page in the memory exceeds a threshold number; and if the number exceeds the threshold number, determining the target storage page from the at least one candidate storage page.
11. A computer program product tangibly stored on a non-transitory computer-readable medium and comprising machine-executable instructions that, when executed, cause the machine to perform the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Data sharing and recovery within a network of untrusted storage devices using data object fingerprinting
US20110099200A1