Method, system and non-transitory readable storage medium for application patching
By dividing application data and update data into variable-sized blocks and hash value generation, compression and encryption, the problem of frequent data erasing and writing in SSD storage is solved, efficient and secure application and update delivery is achieved, and the life of the SSD is extended.
Patent Information
- Application Number
- CN202010715549.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-07-30
- Filing Date
- 2020-07-23
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-05-13
AI Technical Summary
When using solid-state drive (SSD) storage application updates, prior art is difficult to effectively reduce data erasing and writing operations, thereby shortening the lifetime of SSDs.
Efficient packaging and delivery of data is achieved by dividing application data and update data into variable-sized data blocks, and using sliding windows to generate hash values, compress and encrypt these blocks.
This method can ensure the secure delivery of applications and updates while minimizing data erasing and writing as much as possible, and extend the service life of the SSD.
Smart Images

Figure CN112306395B_ABST
Abstract
Description
Technical Field
[0001] Aspects of the present disclosure relate to encryption and compression, and in particular, aspects of the present disclosure relate to encryption and compression of data blocks for transmission over a network in a software patching system. Background Art
[0002] Virtual file delivery is becoming the standard for how users receive their applications. Companies are expected to support their applications by pushing updates to user systems. Efficient virtual file and update delivery is often critical to the security and availability of the application.
[0003] The rise of virtual markets for apps such as the Apple App Store, Google Play Store, Steam Store, Nintendo Store, Microsoft Store, and PlayStation Store means that app developers must push updates and provide app downloads through third parties. In addition, updates and / or apps are often compressed for easier delivery and storage for greater efficiency. Apps and / or updates pushed by app developers can also be encrypted to protect privacy and avoid unnecessary file checks.
[0004] Virtual marketplaces are often more than just platforms that allow downloading of applications and updates. Many marketplaces are designed to be integrated with the file system of the user's device. Therefore, third-party platforms are interested in ensuring that client devices have longevity and efficient file structures.
[0005] The introduction of solid-state drives (SSDs) has made file access very fast compared to hard disk drives (HDDs). Data blocks on SSDs are always accessible, while the accessibility of data blocks on HDDs is determined by the position of the read head and the speed of the platters. Despite the greatly improved performance, SSDs still have several disadvantages. First, the life of an SSD is determined by the number of write and erase cycles the SSD can withstand before it no longer retains data after power is removed. Each bit on an SSD can only be written and erased so many times, after which the bit cannot retain data in the event of a power outage, essentially losing data when the device is powered off. Second, writing and erasing bits on an SSD occurs at different levels. SSD memory is organized into pages and blocks. A page contains a certain number of bits, and each block has a certain number of pages. Data on an SSD cannot be overwritten and must first be deleted before a write can occur. Reading and writing on an SSD occurs at the page level, while deletion only occurs at the block level. This means that if a block contains both data marked for deletion and valid data, the valid data must be written to another location on the SSD when the block is deleted. Disadvantageously, each deletion may result in additional writing to the SSD, thus reducing the life of the SSD.
[0006] Therefore, third-party market operators who push application updates stored on SSDs are interested in reducing the amount of writing and erasing required to store and update applications. This interest conflicts with the interest of application developers who want their applications and updates to be delivered efficiently and securely. Therefore, there is a need for a system or method that can efficiently and securely deliver applications and updates to users in a manner that requires as little data erasure and writing as possible.
[0007] It is against this background that aspects of the present disclosure arise. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Various aspects of the present disclosure may be readily understood by considering the following detailed description taken in conjunction with the accompanying drawings, in which:
[0009] Figure 1 is a flow chart depicting a method for packing data for updating using variable-sized blocks in accordance with aspects of the present disclosure.
[0010] Figure 2 is a flow chart depicting a method for detecting changes in update data having blocks of variable size in accordance with aspects of the present disclosure.
[0011] Figure 3 is a diagram depicting the packaging of variable-sized blocks for delivery to a user device in accordance with aspects of the present disclosure.
[0012] Figure 4 is a diagram illustrating file system updates and merges in variable-sized blocks in accordance with aspects of the present disclosure.
[0013] Figure 5 is a diagram illustrating fused application and patch data according to aspects of the present disclosure.
[0014] Figure 6 is a block diagram illustrating a method for determining block referentiality according to aspects of the present disclosure.
[0015] Fig. 7A is a block diagram depicting a method for packaging data using feedback regarding the probability of the data changing, in accordance with aspects of the present disclosure.
[0016] Figure 7B is a block diagram illustrating a method for updating and merging data using feedback regarding the probability of data changes in accordance with aspects of the present disclosure.
[0017] Figure 8 is a block diagram illustrating a method for variable-sized block deduplication according to aspects of the present disclosure.
[0018] Fig. 9 is a diagram illustrating variable-sized block deduplication in accordance with aspects of the present disclosure.
[0019] Fig.10 is a block diagram illustrating a system for packing data using variable-sized blocks in accordance with aspects of the present disclosure. DETAILED DESCRIPTION
[0020] Although the following specific embodiments include many specific details for illustrative purposes, it will be appreciated by any person skilled in the art that many variations and modifications to the following details are within the scope of the present invention. Therefore, the exemplary embodiments of the present invention described below are set forth without loss of generality and without implying limitations to the claimed invention.
[0021] According to aspects of the present disclosure, the conflict of interest between third-party market operators and software developers in updating applications in SSDs can be resolved by implementing variable-sized data blocks for application data and update data. Variable-sized data blocks allow updates to be applied more efficiently in terms of bytes written compared to solutions that implement only fixed-sized data blocks.
[0022] Apply the patch
[0023] Figure 1 A method for packaging data using variable-sized blocks according to aspects of the present disclosure is depicted. The method can be performed by a tool used by an application developer or a third-party virtual marketplace server, or both. First, as indicated at 101, the tool or server can receive application data. The received data can be organized as an uncompressed data set, a compressed data set, or a mixture of the two, which are concatenated together to create a continuous data set, as indicated at 102. It can be reasonably expected that the application data will subsequently need to be patched, and therefore preparation for the patching process begins during packaging of the application data. First, application data block boundaries can be described; thereafter, the application data will be divided along these boundaries.
[0024] As indicated at 103, a hash may be generated for the sliding window. According to aspects of the present disclosure, in some implementations, the sliding window may run the length of the continuous data set, thereby shifting its length each time the sliding window moves. In other embodiments, the sliding window may move less than the length of the sliding window, for example and in the case of no restrictions, the sliding window may shift half its length. In yet other embodiments, the sliding window may shift more than its length each time it moves, for example and in the case of no restrictions, the sliding window may shift one and a half window lengths each time it moves along the continuous data set. In yet other embodiments, the sliding window may shift the distance of the application block from the previous window, thereby effectively obtaining the hash of the window at the starting point of each application block. In this way, a continuous hash value (also referred to as a rolling hash value) of the data window is generated for the continuous data set. In some implementations, the size of the sliding window also corresponds to a fixed encryption block size. The size of the sliding window may be, for example, but not limited to, 64 kilobytes (KiB) or less. According to aspects of the present disclosure, a rolling hash may be created from the window at the starting point (or bottom) of each application block.
[0025] After generating the rolling hash for the continuous data set, a weak hash 104 may be generated for each depicted application block in the continuous application data set. The weak hash algorithm may be, for example, but not limited to, a checksum. The weak hash may be stored separately or may be part of the metadata of the application.
[0026] As indicated at 105, a strong hash may be generated for each depicted variable-sized chunk. The strong hash for each depicted variable-sized chunk of data may be stored separately or as metadata. The strong hash may be any cryptographic hash function known in the art, such as SHA-3-256, SHA-2-256, SHA-512, MD-5, etc. The strong hash value for the data set may be stored in memory separately from the continuous data set or as part of the metadata for the application data or as part of the application data.
[0027] As indicated at 106, once a hash value is generated for the depicted continuous data set, the data may be divided into variable-sized data chunks. Typically, the uncompressed continuous data will be segmented into smaller variable-sized data chunks, which may vary in size between 2 megabytes (MiB) and 100 kilobytes (KiB). As will be discussed below, developer feedback may be used to guide the segmentation of the continuous data set. As indicated at 107, after the depicted continuous data set is segmented into variable-sized chunks, each variable-sized chunk is compressed. The variable-sized data chunks may be compressed to reduce the data size by any compression algorithm known in the art, such as, but not limited to, the Lempel-Ziv-Welsh compression algorithm (LZW), ZLIB, DEFLATE, and the like.
[0028] After compression, the compressed variable-sized data blocks are merged 108 and then divided into fixed-sized blocks as indicated at 109. Fixed-sized blocks are necessary for encryption and may be smaller than variable-sized data blocks. For example, but not limited to, fixed-sized blocks may be less than or equal to 64KiB. As indicated at 110, after dividing the variable-sized data blocks into fixed-sized data blocks, the data blocks are encrypted. The encryption method may be any strong symmetric encryption method known in the art for protecting communications, such as, but not limited to, DES, 3DES, CAST5, etc. The encrypted data blocks may then be stored or sent to a marketplace server or client device. The encrypted data blocks may be application data or patch data.
[0029] To better visualize the packaging of compressed or uncompressed files for easier delivery, Figure 3 A method for packaging application data according to aspects of the present disclosure is depicted. As shown, compressed or uncompressed files and other data types 301 can be received by a marketplace server over a network, or by a packaging tool application. The server or tool can concatenate the compressed or uncompressed files and other data types 301 into a single continuous data set 302. The continuous data set can then be divided into variable-sized chunks, and each variable-sized chunk of data can then be compressed and merged together to generate compressed variable-sized chunks of application data 303. The compressed variable-sized chunks of application data are further divided into fixed-sized chunks and encrypted; generating encrypted fixed-sized chunks that contain portions of the compressed variable-sized chunk application data 304. The encrypted fixed-sized chunks can then be sent to a client device.
[0030] Figure 4 and Figure 5 An example of using variable-sized blocks for patching and data fusion according to aspects of the present disclosure is shown. Figure 4As shown in , first the client device may have application data 401 stored on the device (eg, stored in any suitable memory or data storage device). As discussed above, the application data may be encrypted in fixed-size blocks and compressed in variable-size blocks.
[0031] During the patching process, the client may receive updates. These updates may consist only of non-referenceable patch data 404. To minimize erases and writes, only the application data replaced by the patch data is invalidated 409 in the metadata, and the location of the patch data is referenced in the metadata. The location of the patch data may be different from the location of the application data because other data blocks 405 may have been written to memory in the time between writing the application data to memory and writing the patch data to memory. In addition, the patch data need not be a simple 1:1 replacement of the patch block for the application block. For example, in Figure 4 In the prediction example shown in , the application block B has been replaced by two patch blocks B1 and B2.
[0032] Application developers can push additional patches. Figure 4 In the example shown in , the location of the additional patch data is next to the previously received patch X.1 402. If X.1 402 is received simultaneously with X.2 403 or no other writes to the memory occur between when X.1 402 is received and when X.2 403 is received, there is no gap in the memory block between X.1 403 and X.2 403. As previously described, the application data block C has expired in the metadata, and the patch data block C1 has been referenced in the metadata. The result of this patching scheme is that as more and more patches are added to the application data, the patch data is scattered throughout the data storage device. Data scattered throughout the memory space is called file or data fragmentation. Although modern SSDs are not severely affected by data fragmentation in SSDs as HDDsif the data is very severely fragmented, this will affect the read bandwidth of the drive because multiple different blocks must be accessed. Therefore, the patching process can implement remote triggering of data fusion to reduce data fragmentation.
[0033] Remote triggering of data fusion
[0034] After updating or before sending to the client device, a fragmentation metric for the application and patch data is obtained. By way of example and not limitation, the fragmentation metric may be, for example, the read bandwidth of the application, and when the read bandwidth drops below a threshold, the patching process will issue a merge command 407. Another example of a fragmentation metric may be, but is not limited to, wasted space due to stale data on storage space. In some embodiments, the tool or server may model the data stored on the client device and use the model to calculate the fragmentation metric before updating.
[0035] Figure 5 Operations generated by remote triggering of a fusion command are shown. When a fusion command is issued, instructions are sent to read application data 501 in a linear fashion. The application data is written 505 to a new location in memory 506 and erased 509. When a reference to patch data 503 is encountered, the invalid application data 502 is erased 509, and the patch data is read 504, written to a new location in memory 508 in place of the invalid application data and erased 509. In some implementations, the patch data is read and then decrypted and re-encrypted before writing the patch data.
[0036] After the patch data referenced in the metadata has been written and erased, the system reads the next application data in the sequence 510, writes the application data to the new location and erases the read application data. This process continues until all application data and patch data have been read, written to the new location and erased at their old location 509. The result is that the application and patch data spanning multiple locations in the memory are now compressed into a single new location 507, 408. In some implementations, erasure of application and patch data does not occur until all application and patch data have been written to the new location. In other embodiments, the application and patch data are erased shortly after being written to the new location. In yet other embodiments, erasure occurs before a block of application data or patch data is written to the new location. A marker can be placed in the metadata that directs the system to the new location of the application and patch data.
[0037] In some cases, the variable-sized blocks do not divide equally into the existing data blocks during the merge operation. In such cases, the variable-sized blocks of application data and patch data can be merged into a few fixed-sized blocks. Figure 5During the fusion operation example shown in , blocks on the storage device associated with variable-sized data blocks A, D and patch data B1, C1 can be read and erased and rewritten without data from the variable-sized blocks. In other embodiments, only the blocks associated with the variable-sized data on the storage device remain unerased, and references are made in the metadata to inform the system of the new location of the application and patch data. This is because the starting points of the variable-sized blocks and the fixed-sized blocks may not be completely aligned during writing. Incomplete alignment means that the fixed-sized blocks may have information from other applications that have no intention of modifying it. The locations of the start and end points of the variable-sized blocks within the fixed-sized area can be listed in the metadata.
[0038] In an embodiment of the present disclosure, the data block 511 may be encrypted as a fixed-size block. In such cases, additional steps must be performed to fuse the data. After performing the read operation, the encrypted data block 511 must be decrypted. The decrypted data is then stored in the working memory 505 and combined with other decrypted variable-size block data 506. The decrypted data is combined according to the size of the fixed-size block. This combination is performed to ensure that when the variable-size block is written to a new location, the fixed-size block is aligned with the memory block of the client device. In addition, during this operation, an offset may be applied to the first fixed-size block of the variable-size block to take into account that the memory block in the client device has some existing data at the location where the fused application and patch data will be written. After the data is combined in a fixed-size block, the block is encrypted and written to the new location 506.
[0039] The above data fusion process may be triggered by the client device detecting fragmentation and performed by the client device. Alternatively, the metric may be sent to the market server, and the market server may issue a fusion command. Alternatively, the fragmentation metric may be calculated by the market server using a model of the application stored in the client device, and the market server may issue a fusion command based on the fragmentation metric generated from the model.
[0040] Data change detection
[0041] Figure 2A method for detecting changes in patch data according to various aspects of the present disclosure is depicted. As indicated at 202, a third-party market server or packaging tool may receive patch data. As indicated at 201, the tool or server may have previously received application data, or the application data may be received simultaneously with the patch data, or after receiving the patch data. According to various aspects of the present disclosure, patch data may be received by the tool or server together with the application data. In order to patch the application data, the data added or changed by the patch must be extracted from the previous version of the application data (hereinafter referred to as the application data) to deliver efficient updates to the client. To assist in this operation, metadata about the application may optionally be received together with the application data, as indicated at 203, or may be stored in a memory from a previous packaging operation. The received application data, patch data, and metadata may be encrypted to protect privacy, and as indicated at 204, the tool or server may decrypt the received encrypted data.
[0042] The patch data and application data may have been divided into compressed variable-sized chunks, so the patch data may be decompressed and, after decompression, each variable-sized chunk of the application data may have a hierarchy of hash values generated for it, as indicated at 205. The hierarchy of hashes may be, for example, but not limited to, strong hashes and weak hashes calculated for the entire chunk or a subset of the data within the chunk. Additionally, the hierarchy may include a rolling hash value for a first window of the chunk data to detect byte-level changes. In some implementations, the application data and / or previous patch data may also be decompressed and have a hierarchy of hashes generated for them. In other embodiments, metadata or some other memory location may have a hierarchy of hashes associated with each chunk in the application data.
[0043] A Bloom filter 206 is created from a rolling hash for the application data. According to some aspects of the present disclosure, a rolling hash value may be generated for a first rolling window of each variable-sized block. According to other aspects of the present disclosure, a rolling hash value may be generated for each window over each variable-sized block. Additionally, a Bloom filter may contain hash values for the entire application's blocks or a specific set of blocks, such as, but not limited to, a hash value for a previous version of a file.
[0044] The strong hash value of the variable-sized chunk of patch data is compared to the hash value of each chunk of application data, as indicated at 207. A hash table may be generated for each strong hash of the application data. Based on the comparison of the strong hashes, a decision is made at 208. When the strong hash of the patch data matches the hash of the chunk of application data, as indicated at 209, the next variable-sized chunk in the patch data is selected for comparison at 207 because the patch data is a duplicate of the existing application data.
[0045] When an equivalent strong hash is not found, a rolling window hash of a first window of variable-sized patch data chunks is created from the bottom of the failed variable-sized patch data chunks, as indicated at 210. The size of the rolling window may be, but is not limited to, less than 64KiB.
[0046] As indicated at 211, a window of a rolling hash of a variable-sized chunk of patch data is searched in a Bloom filter. In the event that no match is found in the Bloom filter, the bytes at the bottom of the current window are determined to be unreferenceable data for patching purposes, as indicated at 216. The window is then moved to the next position, as indicated at 218, and the Bloom filter is applied, and the process is restarted, as indicated at 219, restarting at 211 for the new window position. Alternatively, the process reaches the end of the variable-sized chunk of patch data and the process jumps over 219 to start at 207, or the process reaches the end of the patch data and ends 219.
[0047] If a match is found for the window of the rolling hash of the variable-sized patch data block in the Bloom filter, a weak hash value of the matching variable-sized patch data block is generated 212. At 213, the weak hash starting from the bottom of the matching window of the variable-sized patch data block is compared to the weak hash of the matching application data block found in the Bloom filter. If there is no match, the bytes at the bottom of the current window are determined to be non-referenceable data for patching purposes at 216. The window is then moved to the next position at 218 and the Bloom filter is queried, restarting the process by returning at 219 to search the window of the rolling hash of the patch data in the Bloom filter at 211 for a new window position. Alternatively, the process reaches the end of the variable-sized patch data block and the process restarts by returning at 219 to the hash comparison at 207, or the process reaches the end of the patch data and ends at 219.
[0048] If an exact match of the hash of the patch data block is found in the hash of the application data, then at 214, a strong hash of the patch data block and the application data block is generated. The strong hashes of the matching windows of the application data and patch data are compared at 215. If the strong hashes of the matching windows are not identical, then the bytes at the bottom of the window of patch data are determined to be unreferenceable data for patching purposes 216. The window is then moved to the next position at 218, and the Bloom filter is queried, restarting the process by returning at 219 to search the window of rolling hash of patch data in the Bloom filter at 211 for a new window position. Alternatively, the process reaches the end of the variable-sized patch data block and the process restarts by returning at 219 to the hash comparison at 207, or the process reaches the end of the patch data and ends at 219.
[0049] If the strong hash of the patch data block is equivalent to the strong hash of the application data block, then the patch data is determined to be referenceable at 217. If the location of the referenceable window of the patch data is different from the window location in the application data, a reference may be made in the metadata to reference the location in the application data for the referenceable patch data. Otherwise, the referenceable block of patch data is considered redundant and may not need to be sent with the patch during the patching process. If the location of the referenceable block of patch data is different from the block location in the application data, a reference may be made in the metadata to reference the location in the application data for the referenceable patch data. Otherwise, the reference may not need to be sent with the patch during the patching process. After determining that the patch data block is referenceable, the window is then moved to the next location at 218 and a Bloom filter is applied, and the process is restarted by returning to 211 at 219 to perform a Bloom filter search to find a new window location. Alternatively, the process reaches the end of the variable-sized chunk of patch data and the process restarts at 219 by returning to the hash comparison at 207 , or the process reaches the end of the patch data and ends at 219 .
[0050] According to aspects of the method for detecting patch changes, data may also be applied to the patched application. In the case where the application has already pushed a patch to the client, the application data is considered to contain the patch data that has been provided to the client in the previous patch. Therefore, in these embodiments, all comparisons to the application data also include the patch data previously sent to the client device.
[0051] Unreferenceable data is new and unique data that has not yet been provided to the client device or that the application has not previously obtained. Ideally, during the patching process, unreferenceable data will constitute the majority of the patch data sent to the client device. Referenceable data, on the other hand, may be data found in other parts of the application. This data may be found within a window of the data chunk, and thus while the chunk as a whole may be unreferenceable, portions of the chunk may be identical to application data due to some unreferenceable windows of data within the chunk. In some embodiments where referenceable and unreferenceable data are found within the same chunk, two separate chunks may be created. See Figure 6 for more information about splitting and merging referenceable and nonreferenceable chunks.
[0052] Figure 6 Methods for merging and splitting variable-sized chunks according to aspects of the present disclosure are depicted. According to aspects of the present disclosure, the merging and splitting operations can be performed by a tool used by an application developer or at a marketplace server. The merging and splitting operations can occur after determining portions of a variable-sized chunk of data as unreferenceable (hereinafter referred to as unreferenceable data) at 601. Figure 2 The method for detecting unreferenceable data described in can be used to detect unreferenceable data. Once detected, the unreferenceable data is compared to a first threshold at 602. The first threshold can be, for example, but not limited to, a merge threshold of 128KiB. If the size of the unreferenceable data is not less than the threshold, at 603, a new variable-sized data block is created from the unreferenceable data. If the unreferenceable data is less than the first threshold, it is merged with an adjacent referenceable application (if available) at 604. According to various aspects of the present disclosure, the first threshold can be derived empirically from different factors that affect storage operations. For example, but not limited to, the compression rate of smaller blocks is relatively reduced, the area of minimum possible change is encapsulated in a variable-sized block, the decompression setup cost based on hardware capabilities, etc.
[0053] Next, at 605, the new data chunk or the merged data chunk is compared to a second threshold. The second threshold may be, but is not limited to, a block length threshold of 240 KiB. It should be noted that in some implementations, the second threshold is always greater than the first threshold. If the size of the merged data chunk or the new data chunk is less than the second threshold, the size of the new data chunk or the merged data chunk is retained at 606, and no further segmentation or merging operations are performed on those data chunks. On the other hand, if the size of the new data chunk or the merged data chunk is greater than the second threshold, another comparison is applied to the data. The threshold comparison is whether two variable-sized chunks can be created at 607, each of which is greater than the first threshold. If two chunks greater than or equal to the first threshold cannot be created from the merged or new data chunk, the size of the data chunk is retained at 608, and no further comparison is performed. If two chunks greater than the first threshold can be created, a final comparison is performed.
[0054] The final comparison is whether the new data block or the merged data block is divisible by the second threshold with a remainder greater than or equal to the first threshold at 609. If the remainder of the division is greater than or equal to the first threshold, the merged data block or the new data block is split at the size of the second threshold at 611. If the size of the remainder is less than the first threshold, the block is split at the size of the first threshold at 610. After the split operation has been performed, the system returns to normal operation at 612.
[0055] Use feedback to update optimizations
[0056] Fig. 7A An example of a method for using feedback about probability of change in data regions for application data packet optimization is depicted. As part of a patching process, at 701, an application developer may provide information about an application to a tool or server. The application metadata may include information about the probability of change in data regions. Alternatively, the application metadata may include labels for different regions of application data. Based on the application metadata, regions with high probability of change are determined at 702.
[0057] For example, but not limited to, application metadata may simply indicate that certain areas of data contain data that may change in later patches. In other embodiments, metadata may mark certain areas of the data, such as headers and directories. Data in directories may be marked as having a high likelihood of change because this information typically changes each time a new patch is added.
[0058] Once the application data areas with high probability of change are determined, variable-sized data blocks can be created based on the application data areas with high probability of change at 703. Alternatively, the boundaries of the variable-sized blocks of application data can be adjusted to fit the areas with high probability of change. In other implementations according to various aspects of the present disclosure, the boundaries of the variable-sized blocks of application data can fit the areas provided by the user, and the probability of change in those areas is relatively low. The adjusted generation of the variable-sized data blocks can be performed, for example but not limited to, by creating variable-sized data blocks, which contain data with high probability of change and end at data with low probability of change. Alternatively, the boundaries of existing variable-sized blocks can be adjusted so that blocks containing a lot of data with high probability of change have end boundaries at data with low probability of change. The generation efficiency is improved because during the erase-write cycle, data with low probability of change is not deleted and must be rewritten for each change. In addition, as the boundaries between blocks are better defined and less overlapping data is transmitted during patching, the download size of the patch is reduced with area.
[0059] Figure 7B A method for performing patch data merging using feedback about the probability of change of data regions is shown. An application developer may provide metadata about the likelihood that regions of patch data and application data will change at 704. At 705, the metadata may be used to determine which regions of the patch data are likely to change in future patches. Using this information, the creation of variable-sized patch data blocks may be guided. For example, but not limited to, adjacent patch data that is determined to have a high probability of change may be included in a variable-sized patch data block, while other patch data that is not determined to have a high probability of change may be included in a different patch data block. In addition, in the case of performing a merge operation, the patch metadata and application metadata may guide the merge operation. For example, in the case of determining a non-referenceable region of patch data and a merge is appropriate, if the non-referenceable patch data is determined to have a high probability of change, a merge operation may be performed at 706 on adjacent referenceable data that has also been determined to have a high probability of change.
[0060] According to aspects of the present disclosure, during packaging, application data and patch data may be deduplicated to reduce the amount of writes performed on the client device. Figure 8An exemplary method for deduplication according to various aspects of the present disclosure is depicted. During the deduplication process, at 801, a strong hash value for each variable-sized block of an application is compared to a strong hash value for each other variable-sized block of the application or a subset of variable-sized blocks. Through the comparison, at 802, it is determined whether any variable-sized block in the application has the same data as another block of the application. If it is determined that two variable-sized blocks have the same data, only the first identical variable-sized block is written to a storage device. At 803, a reference is made in metadata to the location of each block in the storage device that is associated with the first identical variable-sized block for the second variable-sized block, and at 804, writing of blocks associated with duplicate compressed variable-sized blocks to the storage device is omitted.
[0061] Fig. 9 Memory blocks 904 and metadata 905 after deduplication are shown. As shown, variable-sized chunks 0, 1, 2, 4, 5 are written across multiple fixed-sized memory blocks. Metadata 905 contains information for variable-sized chunks 0, 1, 2, 3, 4, 5. In the illustrative case, variable-sized chunks 0 and 3 are identical, and therefore, when writing data chunks to memory, the deduplication process omits chunk 3. In the metadata, variable-sized chunk 0 901 references variable-sized chunk 0 in memory 906. After deduplication, the metadata for variable-sized chunk 3 902 references 907 variable-sized chunk 0 902 in memory.
[0062] Application patching system
[0063] Fig.10 A system 1000 configured to package application data and patch data or store application and patch data according to aspects of the present disclosure is depicted. The system 1000 may include one or more processor units 1003, which may be configured according to a well-known architecture such as, for example, a single core, dual core, quad core, multi-core, processor-coprocessor, cell processor, etc. The system may also include one or more memory units 1004 (e.g., random access memory (RAM), dynamic random access memory (DRAM), read-only memory (ROM), etc.).
[0064] Processor unit 1003 may execute one or more programs, portions of which may be stored in memory 1004, and processor 1003 may be operatively coupled to memory (e.g., accessing memory via data bus 1005). The program may be configured to download application information 1021 including application data, and according to the above description of Figure 1 and / or Figure 3 and / or Figure 4The application packaging method described in the foregoing description may package application information. In addition, the memory 1004 may contain information about the connection 1008 between the system and one or more application developer servers or networks and client devices. Such connection information may include, for example, Internet Protocol (IP) addresses, Network Address Translator (NAT) traversal information, and connection performance information, such as bandwidth and latency. Patch data 1010 may also be stored in the memory 1004. The program may also be configured to download the patch data 1010 and perform the patching according to the patching method by Figure 2 and / or Figure 5 As part of packaging and patching, the program may generate hash values for the application and / or patch data, which hash values 1009 may be stored in memory before being added to the application information as metadata. In addition, memory 1004 may contain compressed data 1022 and / or one or more encrypted data 1023 for use during packaging of the application and / or patch data. The program may also be configured to Fig. 7A and Figure 7B The method described in 1004 modifies variable-sized application blocks based on application metadata. In addition, the memory 1004 may contain application information 1021, such as metadata about the size of the application blocks, the number of application blocks that make up the application, the order of the application blocks, and the possibility of changes associated with the application blocks. Application information, patch information, hash values, compressed data, encrypted data, and connection information may also be stored as data in the mass storage device 1015.
[0065] The system 1000 may also include well-known support circuits, such as input / output (I / O) 1007, circuits, power supply (P / S) 1011, clock (CLK) 1012, and cache 1013, which may communicate with other components of the system, for example, via bus 1005. The system may include a network interface 1014. The processor unit 1003 and the network interface 1014 may be configured to implement a local area network (LAN) or a personal area network (PAN), via a suitable network protocol for the PAN (e.g., Bluetooth). The system may optionally include a mass storage device 1015 (such as a disk drive, CD-ROM drive, tape drive, flash memory, etc.), and the mass storage device may store programs and / or data. The market server may also include a user interface 1016 for facilitating interaction between the system and the user. The user interface may include a monitor, a television screen, a speaker, a headset, or other means of conveying information to the user. A user input device 1002 (such as a mouse, keyboard, game controller, joystick, etc.) may communicate with the I / O interface and provide system control to the user.
[0066] Although the above is a complete description of the preferred embodiment of the present invention, it is possible to use various alternatives, modifications and equivalents. Therefore, the scope of the present invention should not be determined with reference to the above description, but should instead be determined with reference to the appended claims and their full range of equivalents. Any feature described herein (whether preferred or not) may be combined with any other feature described herein (whether preferred or not). In the appended claims, unless explicitly stated otherwise, Indefinite article "A / An" Refers to the quantity of one or more of the item following the article. The appended claims should not be interpreted as including means-plus-function limitations unless such a limitation is explicitly recited in a given claim using the phrase "means for..."
Claims
1. A method for applying a patch, comprising: a) concatenate uncompressed data into continuous data sets; b) dividing the continuous data set into data blocks of variable sizes; c) compressing each of said variable-sized data blocks; d) dividing each of the compressed variable-sized data blocks into fixed-sized data blocks; e) encrypting the fixed-size data block to generate an encrypted fixed-size data block; f) sending the encrypted fixed-size data block over a network; g) dividing the patch data into patch data blocks of variable sizes; h) determining a relationship between a patch data block of variable size and the data block of variable size; i) generating patch metadata for said relationship between said variable-sized patch data block and said variable-sized data block; j) compressing the variable-sized patch data block to generate a compressed variable-sized patch data block; k) dividing the compressed variable-size patch data block into fixed-size patch data blocks; l) encrypting the fixed-size patch data block to generate an encrypted fixed-size patch data block; m) sending the encrypted fixed-size patch data block and metadata over the network.
2. The method of claim 1 further comprising generating metadata for the location of the variable-sized data chunks within the encrypted fixed-sized chunks.
3. The method of claim 1, wherein each of the variable-sized data chunks is greater than 64 kilobytes in size.
4. The method of claim 1, further comprising deduplicating the compressed variable-sized data blocks after dividing the compressed variable-sized data blocks into fixed-sized data blocks.
5. The method of claim 4, further comprising generating metadata for the variable-sized data chunks, wherein the metadata includes references to repeated chunks.
6. The method of claim 4, wherein deduplicating the compressed variable-sized data blocks comprises: Determining a hash value for each variable-sized block, comparing the hash value of a first variable-sized block of data to a hash value of a second variable-sized block of data, wherein the first variable-sized block of data and the second variable-sized block of data have matching hash values, deleting the second variable-sized block of data, and creating a reference to the first variable-sized block of data in a memory.
7. A system for application patching, comprising: processor; a memory coupled to the processor; non-transitory instructions embedded in the memory, which when executed cause the processor to perform a method comprising: a) concatenate uncompressed data into continuous data sets; b) dividing the continuous data set into data blocks of variable sizes; c) compressing each of said variable-sized data blocks; d) dividing each of the compressed variable-sized data blocks into fixed-sized data blocks; e) encrypting the fixed-size data block to generate an encrypted fixed-size data block; f) sending the encrypted fixed-size data block over a network; g) dividing the patch data into patch data blocks of variable sizes; h) determining a relationship between a variable-sized patch data block and a variable-sized data block; i) generating patch metadata for the relationship between the variable-sized patch data blocks and the variable-sized data blocks; j) compressing the variable-sized patch data block to generate a compressed variable-sized patch data block; k) dividing the compressed variable-size patch data block into fixed-size patch data blocks; l) encrypting the fixed-size patch data block to generate an encrypted fixed-size patch data block; m) sending the encrypted fixed-size patch data block and metadata over the network.
8. The system of claim 7, wherein the method further comprises generating metadata for the location of the variable-sized data chunks within the encrypted fixed-sized chunks.
9. The system of claim 7, wherein each of the variable-sized data chunks is greater than 64 kilobytes in size.
10. The system of claim 7, wherein the method further comprises deduplicating the compressed variable-sized data blocks after dividing the compressed variable-sized data blocks into fixed-sized data blocks.
11. The system of claim 10, wherein the method further comprises generating metadata for the variable-sized data chunks, wherein the metadata includes references to repeated chunks.
12. The system of claim 10, wherein deduplicating the compressed variable-sized data blocks comprises: Determining a hash value for each variable-sized block, comparing the hash value of a first variable-sized block of data to a hash value of a second variable-sized block of data, wherein the first variable-sized block of data and the second variable-sized block of data have matching hash values, deleting the second variable-sized block of data, and creating a reference to the first variable-sized block of data in a memory.
13. A non-transitory computer-readable storage medium having instructions embedded thereon, the instructions, when executed, causing a computer to perform a method for applying a patch, the method comprising: a) concatenate uncompressed data into continuous data sets; b) dividing the continuous data set into data blocks of variable sizes; c) compressing each of said variable-sized data blocks; d) dividing each of the compressed variable-sized data blocks into fixed-sized data blocks; e) encrypting the fixed-size data block to generate an encrypted fixed-size data block; f) sending the encrypted fixed-size data block through a network; g) dividing the patch data into patch data blocks of variable sizes; h) determining a relationship between a patch data block of variable size and the data block of variable size; i) generating patch metadata for the relationship between the variable-sized patch data blocks and the variable-sized data blocks; j) compressing the variable-sized patch data block to generate a compressed variable-sized patch data block; k) dividing the compressed variable-size patch data block into fixed-size patch data blocks; l) encrypting the fixed-size patch data block to generate an encrypted fixed-size patch data block; m) sending the encrypted fixed-size patch data block and metadata over the network.
14. The non-transitory computer-readable storage medium of claim 13, wherein the method further comprises generating metadata for the location of the variable-sized data chunks within the encrypted fixed-sized chunks.
15. The non-transitory computer-readable storage medium of claim 13, wherein each of the variable-sized data chunks is greater than 64 kilobytes in size.
16. The non-transitory computer-readable storage medium of claim 13, wherein the method further comprises deduplicating the compressed variable-sized data blocks after dividing the compressed variable-sized data blocks into fixed-sized data blocks.
17. The non-transitory computer-readable storage medium of claim 16, wherein deduplicating the compressed variable-sized data chunks comprises: Determining a hash value for each variable-sized block, comparing the hash value of a first variable-sized block of data to a hash value of a second variable-sized block of data, wherein the first variable-sized block of data and the second variable-sized block of data have matching hash values, deleting the second variable-sized block of data, and creating a reference to the first variable-sized block of data in a memory.
Citation Information
Patent Citations
Secure Relational File System With Version Control, Deduplication, And Error Correction
US20150293817A1
Data write to subset of memory devices
US20170220488A1