Ship increment updating method based on improved Bsdiff algorithm

By improving the Bsdiff algorithm and using the EM-SA-DS algorithm for suffix sorting, the problems of large memory usage and general performance during remote updates of ships are solved, and more efficient incremental updates are achieved, reducing the amount of data transmitted and improving the upgrade success rate.

CN120010896APending Publication Date: 2025-05-16RES INST 708 OF CHINA STATE SHIPBUILDING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510163889.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing Bsdiff algorithm requires a large memory space to generate incremental files when ships are updated remotely, and its performance is average, resulting in the server being unable to allocate enough memory when large applications are updated, affecting the generation of differential packages.

Method used

Using the improved Bsdiff differential algorithm, the EM-SA-DS algorithm is used instead of qSufSort for suffix sorting, generating a suffix array, reducing memory usage and improving efficiency.

Benefits of technology

It effectively reduces the size of incremental update files, reduces the amount of data transmitted, improves the upgrade success rate in a weak network environment, and improves the efficiency of the update process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010896A_ABST
    Figure CN120010896A_ABST
Patent Text Reader

Abstract

The invention relates to a ship increment updating method based on an improved Bsdiff algorithm, and the method is characterized in that the method comprises the following steps: 1, generating a suffix array based on an old file through employing an EM-SA-DS algorithm; 2, initializing a difference file; 3, comparing a new file by using the generated suffix array, and generating an approximately matched region; 4, determining the boundary of the difference file; 5, writing in a control file, a difference file and a newly added file; an incremental update file is generated by the old file and the new file at the shore end and transmitted to the ship end, and the new file is generated by the old file and the incremental update file at the ship end; the problem that the memory space required for generating the increment file is large when the Bsdiff algorithm is used for remote updating of the ship-end equipment is solved, the size of an updating package can be effectively reduced, the transmitted data volume is reduced, rollback caused by upgrading failure is reduced, memory use is reduced, and efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a ship software remote upgrade technology, and in particular to a ship incremental update method based on an improved Bsdiff algorithm. Background Art

[0002] As large ships accumulate data, iterate training, and change ship conditions during long-distance voyages, the corresponding application algorithms or model data packages attached to the applications need to be updated. In the past, personnel went on board for maintenance and updates after arriving at the shore, which was very time-effective. The future trend is to upgrade through remote offshore. Since the network has low bandwidth, high latency, and high packet loss when the ship is sailing offshore, incremental upgrades can be used to transmit only the updated part rather than the entire software. Incremental update refers to the calculation of the differentiated sequence of two versions based on the differentiation algorithm. The client only needs to update and download the sequence, which can effectively reduce the traffic of network transmission, improve security, and reduce large-scale rollback events caused by the interruption of full updates.

[0003] Currently, the industry mainly uses the Bsdiff algorithm for incremental upgrades. This algorithm generates smaller differential packages and can ensure that the upgraded files are consistent with the original files. The Bsdiff algorithm uses a differential algorithm to perform differential calculations on executable files before and after modification. A slight change in the source file of the executable file may cause a large number of changes to the executable file. There are two types of changes: one is the part directly affected by the source code modification. The new and old files of this part are inconsistent on a large scale. The other is the change in the pointer address or register address of the unmodified code part caused by the insertion or modification of the code. The change in the binary content of this part is extremely sparse, usually not more than one or two bytes. In the two old and new code blocks affected by the second type, most of the content byte difference is 0, and a small part of the address displacement data that needs to be updated has a large number of identical displacement values. This part of the difference data can be efficiently compressed. Therefore, the core idea of ​​the Bsdiff algorithm is to find these two parts to form new files and difference files. Such as Figure 1 shown.

[0004] The bottleneck of the algorithm is the construction of the suffix array, which takes up a lot of memory and has average performance. bsdiff uses qSufSort to sort the suffixes of old files. The code is as follows:

[0005]

[0006] free(V);

[0007] Bsdiff uses the qSufSort algorithm to sort suffixes. This algorithm requires two arrays, I and V. Array I stores the generated suffix array, and V is an auxiliary array that can be released after sorting. If the size of the old file is n bytes, off_t is an integer type used to represent the file offset. It generally occupies 4 bytes in a 32-bit environment. The two arrays I and V will occupy a total of 8n bytes of memory space, plus a previously loaded old file, a total of 9n bytes of memory are required; off_t generally occupies 8 bytes in a 64-bit environment. The two arrays I and V will occupy 16n bytes of memory space, plus a previously loaded old file, a total of 17n bytes of memory are required. As the size of the application continues to increase, if the server cannot allocate such a large memory space, it will not be able to generate a difference package. In addition, the performance of the qSufSort algorithm is relatively general. Summary of the invention

[0008] Aiming at the problem that the memory space required to generate incremental files when remotely updating ship-side equipment using the Bsdiff algorithm is large, a ship incremental update method based on the improved Bsdiff algorithm is proposed. The improved Bsdiff difference algorithm is used to generate incremental update files.

[0009] The technical solution of the present invention is:

[0010] A ship incremental update method based on an improved Bsdiff algorithm comprises the following steps:

[0011] The first step is to use the EM-SA-DS algorithm to generate a suffix array based on the old file;

[0012] The second step is to initialize the difference file;

[0013] The third step is to compare the new file with the generated suffix array to generate an approximate matching area;

[0014] Step 4: Determine the boundary of the difference file;

[0015] Step 5: Write control files, difference files, and new files;

[0016] At the shore end, an incremental update file is generated from the old file and the new file through the above steps, and the incremental update file is transmitted to the ship end, where a new file is generated from the old file and the incremental update file.

[0017] Further, the specific steps are as follows:

[0018] The first step is to generate a suffix array based on the old file;

[0019] Build an array based on the old file, use the EM-SA-DS algorithm to sort the suffixes of the array built from the old file, and generate a suffix array;

[0020] The second step is to initialize the difference file;

[0021] The header file is 32 bytes in total and is divided into four parts. The first part is the file format signature, which is an 8-byte constant "BSDIFF40". The second part is the length of the control file compression stream, which is an 8-byte integer. The third part is the length of the difference file compression stream, which is an 8-byte integer. The fourth part is the length of the new file, which is an 8-byte integer.

[0022] After the header file are compressed streams of the control file, difference file, and newly added files;

[0023] The third step is to compare the new file with the generated suffix array to generate an approximate matching area;

[0024] Initialize an empty approximate matching area in the old file starting from the end of the previous segment; start from the end of the previous segment of the new file, start to compare with the suffix array at each byte, and find a section in the old file that is exactly the same as the new file, with a length of len; compare the fields of length len at the beginning of the approximate matching area of ​​the new file and the old file. If they are exactly the same, the next time the new file is compared directly from len bytes; if there are more than 8 bytes that are different, or the end of the new file has been reached, then this round ends and the approximate matching area of ​​the old file is output; otherwise, the comparison continues from the next byte of the new file;

[0025] Step 4: Determine the boundary of the difference file;

[0026] For the approximate matching area generated in the third step, compare the beginning of the area with the beginning of the new file to obtain a segment lenf with a 50% byte identity rate; compare the end of the area with the end of the new file to obtain a segment lenb with a 50% byte identity rate; if the two segments overlap, find a position in the overlapping part to maximize the number of bytes in the approximate matching area as the boundary;

[0027] Step 5: Write control files, difference files, and new files;

[0028] The control file is an array, each element of which is a triple (x, y, z) consisting of three 8-byte integers; x is the size of each difference file; y is the size of each newly added file, which may be 0; z is the deviation value, which is used to adjust the current position of the old file, adding z bytes, which may be a negative number; the difference of the content corresponding to segment lenf is calculated and compressed and written to the difference file; if segment lenf and segment lenb are not connected, the middle area is directly compressed and written to the newly added file, if they are connected, it does not exist; the beginning of segment lenb is used as the beginning of the next round of approximate matching area; the control file, difference file and newly added file are generated into a new file;

[0029] At the shore end, an incremental update file is generated from the old file and the new file through the above steps, and the incremental update file is transmitted to the ship end, where the old file and the incremental update file are generated.

[0030] Furthermore, in the first step, in the EM-SA-DS algorithm, in order to obtain the information of the previous characters of s[sa[i]] from the external memory without incurring a large amount of random I / O, the concept of substring buffer is proposed, which is defined as: a substring buffer is a substring containing d+2 triples<ch,t,dc> are used to store the starting character of a suffix and the related information of its d+1 preceding characters, and the first triplet stores the related information of the character s[sa[i]], the second triplet stores the related information of the character s[sa[i]-1], and so on; among them, ch represents the character value, t represents whether ch is L type or S type, and dc represents whether ch is a d-critical character.

[0031] Furthermore, in the EM-SA-DS algorithm, let xa and ya be two ordered sets, where xa is the set of all L-type d-critical suffixes arranged in descending order, and ya is the set of all S-type d-critical suffixes arranged in ascending order; the specific process of inductive sorting of EM-SA-DS is as follows:

[0032] Step 1.1, calculate strbuf(sa,i), strbuf(xa,i) and strbuf(ya,i) of all non-empty suffixes in sa, xa and ya respectively;

[0033] Step 1.2. Scan each non-empty element sa[i] in sa from left to right, and let j = sa[i] - 1: Case 1: If s[sa[i]] is an L-type d-critical character, then copy the substring buffer of the suffix suf(s, sa[i]) from strbuf(xa) to strbuf(sa), and at the same time delete the substring buffer of the suffix suf(s, sa[i]) from strbuf(xa); Case 2: Obtain the ch, t, and dc information of the character s[sa[j]] from strbuf(sa, i); If s[j] is of L-type, put suf(s, j) into the leftmost empty position of its bucket bucket(sa, s[j]), let the inserted position be j', and update strbuf(sa, j') accordingly according to strbuf(sa, i).

[0034] Step 1.3. Scan each non-empty element sa[i] in sa from right to left, and let j = sa[i] - l: Case 1: If s[sa[i]] is an S-type d-critical character, then copy the substring buffer of the suffix suf(s, sa[i]) from strbuf(ya) to strbuf(sa), and at the same time delete the substring buffer of the suffix suf(s, sa[i]) from strbuf(ya); Case 2: Obtain the ch, t, and dc information of the character s[sa[j]] from strbuf(sa, i); If s[j] is of S-type, put suf(s, j) into the rightmost empty position of its bucket bucket(sa, s[j]), let the inserted position be j', and update strbuf(sa, j') accordingly according to strbuf(sa, i).

[0035] Further, in the first step, all the suffixes are rearranged in ascending order of lexicographical order, and the integer array formed by using their respective starting subscripts to represent the suffix strings is the suffix array.

[0036] Further, in the first step, a suffix refers to the string formed from a certain position to the last position in the string.

[0037] Further, in the first step, the lexicographical order means that for two strings s1 and s2, if s1 < s2, currently only when: for the first m characters, s1[i] == s2[i]; for the length of s1 being m, or for the (m + 1)-th character, s1[m + 1] < s2[m + 1].

[0038] Furthermore, in the first step, the main idea of ​​the EM-SA-DS algorithm is: first select all d-critical substrings in the string s and sort them to obtain the reduced string s1; then according to the size of the reduced string s1, call the suffix array memory construction algorithm or recursively call the EM-SA-DS algorithm to calculate the suffix array sa1 of the string s1 and finally derive the suffix array sa of s through sa1.

[0039] Furthermore, in the first step, the EM-SA-DS algorithm must know the predecessor of each character s[sa[i]] in s during the sorting process and insert it into the starting position of its corresponding bucket; if all operations are performed in memory, the predecessor of the character and the relevant information of the bucket can be obtained directly from the memory.

[0040] Furthermore, when the length n of the string s is very large, all of this information cannot be stored in the memory and must be stored with the help of external memory.

[0041] The beneficial effects of the present invention are:

[0042] 1) Traditional ship-to-shore remote updates usually adopt a full update method, which has a large amount of data transmission. The present invention can effectively reduce the size of the update package and reduce the amount of data transmitted.

[0043] 2) In the water environment, the network signal coverage may not be comprehensive, and there is a weak network environment with low network bandwidth, high latency, and high packet loss, which is prone to packet loss and disconnection. The present invention can reduce the rollback of upgrade failures.

[0044] 3) The EM-SA-DS algorithm is used instead of qSufSort to perform suffix sorting on the binary block array, which reduces the memory required for operation and improves efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 For difference files and newly added files diagram;

[0046] Figure 2 Generating a schematic diagram for the suffix array of the present invention;

[0047] Figure 3 This is the header file diagram of the present invention;

[0048] Figure 4 Incremental update file format diagram for the present invention;

[0049] Figure 5 A flow chart for incrementally updating files of the present invention;

[0050] Figure 6 It is an expansion diagram for comparing the documents of the present invention;

[0051] Figure 7 Generate the process diagram of the incremental update file and restore the new file for the present invention;

[0052] Figure 8 Overall block diagram of the incremental upgrade of the present invention;

[0053] Fig. 9 Schematic diagram of the suffix array after sorting the string "home" of the present invention in lexicographical order. Detailed implementation manners

[0054] The present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and the detailed implementation manners and specific operation processes are given, but the protection scope of the present invention is not limited to the following embodiments.

[0055] Assume that the number of bytes of the old file is n. The Bsdiff algorithm uses the qSufSort method to perform suffix sorting on the old file in lexicographical order to generate a suffix array, as Figure 2 shown. This method requires 17n bytes of memory space in a 64-bit environment, so the server configuration generally cannot perform incremental updates for large applications.

[0056] The present invention uses the EM-SA-DS algorithm to replace the qSufSort method.

[0057] Based on the above considerations, the present invention uses the following method to generate an incremental update file, and the flowchart is as Figure 5 shown:

[0058] The first step is to generate a suffix array based on the old file.

[0059] Construct an array based on the old file, and use the EM-SA-DS algorithm to perform suffix sorting on the array constructed from the old file to generate a suffix array.

[0060] A suffix refers to a string formed from a certain position to the last position in a string. For example, the string "home" has 4 corresponding suffix strings: "e", "me", "ome", "home". Lexicographical order means that for two strings s1 and s2, if s1 < s2, currently only when: for the first m characters, s1[i] == s2[i]; for s1 with a length of m, or for the (m + 1)-th character, s1[m + 1] < s2[m + 1]. Rearrange all the suffixes in ascending lexicographical order, and use their respective starting subscripts to represent the suffix strings, forming an integer array, that is, the suffix array. As Fig. 9 shown.

[0061] The main idea of ​​the EM-SA-DS algorithm is: first select all d-critical substrings in the string s (Note: the specific process of inductive sorting of EM-SA-DS below is a detailed explanation of generating d-critical substrings) and sort them to obtain the reduced string s1; then according to the size of the reduced string s1, call the suffix array memory construction algorithm or recursively call the EM-SA-DS algorithm to calculate the suffix array sa1 of the string s1 and finally deduce the suffix array sa of s through sa1.

[0062] Since the previous character of each character s[sa[i]] in s must be known during the sorting process and inserted into the starting position of the corresponding bucket, if all operations are performed in memory, the previous character and bucket related information can be obtained directly from the memory. However, when the length n of the string s is very large, this information cannot be stored in memory and must be stored in external memory.

[0063] In the EM-SA-DS algorithm, in order to obtain the information of the previous characters of s[sa[i]] from the external memory without causing a large amount of random I / O, the concept of substring buffer is proposed, which is defined as: a substring buffer is a triple containing d+2<ch,t,dc> The arrays are used to store the starting character of a suffix and the related information of its d+1 preceding characters, and the first triple stores the related information of the character s[sa[i]], the second triple stores the related information of the character s[sa[i]-1], and so on. Among them, ch represents the character value, t represents whether ch is L-type or S-type, and dc represents whether ch is a d-critical character. Let xa and ya be two ordered sets, where xa is the set of all L-type d-critical suffixes arranged in descending order, and ya is the set of all S-type d-critical suffixes arranged in ascending order.

[0064] The specific process of EM-SA-DS inductive sorting is as follows:

[0065] 1. Calculate strbuf(sa,i), strbuf(xa,i) and strbuf(ya,i) of all non-empty suffixes in sa, xa and ya respectively;

[0066] 2. Scan each non-empty element sa[i] in sa from left to right, and let j = sa[i]-1: (1) If s[sa[i]] is an L-type d-critical character, copy the substring buffer of the suffix suf(s,sa[i]) from strbuf(xa) to strbuf(sa), and delete the substring buffer of the suffix suf(s,sa[i]) from strbuf(xa); (2) Get the ch, t, and dc information of the character s[sa[j]] from strbuf(sa,i). If s[j] is L-type, put suf(s,j) into the leftmost empty position of its bucket (sa,s[j]), set the insertion position to j', and update strbuf(sa,j') accordingly according to strbuf(sa,i);

[0067] 3. Scan each non-empty element sa[i] in sa from right to left, and let j = sa[i]-l: (1) If s[sa[i]] is an S-type (non-LMS) d-critical character, copy the substring buffer of the suffix suf(s,sa[i]) from strbuf(ya) to strbuf(sa), and delete the substring buffer of the suffix suf(s,sa[i]) from strbuf(ya); (2) Get the ch, t, and dc information of the character s[sa[j]] from strbuf(sa,i). If s[j] is S-type, put suf(s,j) into the rightmost empty position of its bucket (sa,s[j]), set the insertion position to j', and update strbuf(sa,j') accordingly according to strbuf(sa,i).

[0068] Among them, LMS in non-LMS: The substrings in LMS decomposition are divided into L type (Leftmost, that is, a substring is the leftmost part of a suffix), M type (Middle, that is, a substring is neither the leftmost nor the rightmost part), and S type (Suffix, that is, a substring is the rightmost part of a suffix).

[0069] The second step is to initialize the difference file.

[0070] The header file is 32 bytes in total and is divided into four parts. The first part is the file format signature, which is an 8-byte constant "BSDIFF40". The second part is the length of the control file compression stream, which is an 8-byte integer. The third part is the length of the difference file compression stream, which is an 8-byte integer. The fourth part is the length of the new file, which is an 8-byte integer. Figure 3 shown.

[0071] After the header file is the compressed stream of the control file, difference file and newly added file, the specific content will be written later. Figure 4 shown.

[0072] The third step is to compare the new file with the generated suffix array to generate an approximately matching area.

[0073] Initialize an empty approximate matching area in the old file starting from the end of the previous section. Starting from the end of the previous section of the new file, start comparing each byte with the suffix array, and find a section in the old file that is exactly the same as the new file, with a length of len. Compare the fields of length len at the beginning of the approximate matching area of ​​the new and old files. If they are exactly the same, the next time the new file is compared directly from len bytes; if there are more than 8 bytes that are different, or it has reached the end of the new file, this round ends and the approximate matching area of ​​the old file is output; otherwise, the comparison continues from the next byte of the new file. Figure 6 shown.

[0074] The fourth step is to determine the boundaries of the difference files.

[0075] For the approximate matching area generated in the third step, compare the beginning of the area backward with the beginning of the new file to obtain a segment lenf with a 50% byte identity rate. Compare the end of the area forward with the end of the new file to obtain a segment lenb with a 50% byte identity rate. If the two segments overlap, find a position in the overlapping part to maximize the number of bytes in the approximate matching area as the boundary.

[0076] Step 5: Write control files, difference files, and new files.

[0077] The control file is an array, each element of which is a triple (x, y, z) consisting of three 8-byte integers. x is the size of each difference file; y is the size of each newly added file, which may be 0; z is the deviation value, which is used to adjust the current position of the old file, adding z bytes, which may be negative. The difference of the content corresponding to segment lenf is calculated and compressed and written to the difference file; if segment lenf and segment lenb are not connected, the middle area is directly compressed and written to the newly added file, if they are connected, it does not exist; the beginning of segment lenb is used as the beginning of the next round of approximate matching area. Generate a new file with the obtained control file, difference file and newly added file.

[0078] On the shore side, the old file and the new file are used to generate an incremental update file through the above steps, which is then transmitted to the ship side, where the old file and the incremental update file are used to generate a new file. Since the incremental update file is smaller than the original file in most cases, the transmission pressure can be greatly reduced, meeting the requirement of completing the software update with the minimum transmission amount. Figure 7 and Figure 8 As shown, the differential file in the figure is also a difference file.

[0079] Simply put, the old file is application V1.0, and the new file is application V2.0. The purpose of incremental update is to update the V1.0 version file on the ship to V2.0. Originally, the complete V2.0 program package should be sent to the ship. This technology can generate an incremental update package smaller than the V2.0 complete program package based on V2.0 and V1.0 on the shore side, thereby reducing the consumed traffic and increasing the success rate of ocean-going ships upgrading in weak network areas.

[0080] Let's take a simple example to illustrate the specific process of the above steps. Assume that the old file is "cgakxmiszraklzedqliazvukgesogaoxicccroplk\n", with a size of 42 bytes, and the new file is "cgakxmiyyraklzedmufdzwwhlaqliazvukgesog\n", with a size of 40 bytes.

[0081] The suffix array generated based on the old file in the first step is [41,10,2,29,19,33,34,0,35,15,14,25,1,28,24,18,32,6,40,23,11,3,17,39,12,5,27,37,30,38,16,9,36,26,7,22,21,31,4,13,8,20].

[0082] The second step is to initialize the difference file. The header file is 32 bytes in total and is divided into four parts. The first part is the file format signature, which is an 8-byte constant "BSDIFF40". The second, third, and fourth parts are all initialized to 0. After the header file is the compression stream of the control file, difference file, and newly added file. The specific content will be written in the following steps.

[0083] Steps 3 to 5 form a cycle.

[0084] In the first loop, the similar matching area generated in the third step is the first 26 bytes of the new file, "cgakxmiyyraklzedmufdzwwhla", because bytes 27-39 of the new file can completely match the last 17-39 bytes of the old file, but there are more than 8 mismatched bytes corresponding to bytes 28-40 of the corresponding old file.

[0085] The fourth step is to compare from the beginning of the region backwards with the beginning of the new file in this round to obtain a segment lenf with a 50% byte identity rate of "cgakxmiyyraklzed". Compare from the end of the region forwards with the end of the new file in this round to obtain a segment lenb with a 50% byte identity rate of "cgakxmiyyraklzed".

[0086] The fifth step is to write the control file, difference file and new file. The control file is an array. The first triple (x, y, z) x is 16 (the length of the difference file), y is 10 (the length of the new file), z is 0 (the deviation value), and it is written after compression. The difference file is the compression of the difference between the new file "cgakxmiyyraklzed" and the old file "cgakxmiszraklzed" in the lenf part. The new file is the compression of the part "mufdzwwhla" between lenf and lenb.

[0087] The second loop will start from the 27th byte of the new file and the 17th byte of the old file. The similar matching area "qliazvukgesog" generated in the third step. The fourth step segment lenf is "qliazvukgesog" and lenb is empty. The fifth step controls the second triple (x, y, z) of the file. x is 16 (the length of the difference file), y is 1 (the length of the newly added file), and z is 12 (the deviation value). The difference file is the compression of the difference between the new file "qliazvukgesog" and the old file "qliazvukgesog" in the lenf part. The newly added file is compressed with "\n". The end of the new file has been traversed and the loop ends.

[0088] Generate a new file.

[0089] The above-mentioned embodiment only expresses one implementation mode of the present invention, and its description is relatively specific and detailed, but it cannot be understood as limiting the scope of the invention patent. It should be pointed out that for ordinary technicians in this field, several modifications and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be based on the attached claims.

Claims

1. A ship incremental update method based on an improved Bsdiff algorithm, characterized in that: The following steps are involved: The first step is to use the EM-SA-DS algorithm to generate a suffix array based on the old file; The second step is to initialize the difference file; The third step is to compare the new file with the generated suffix array to generate an approximate matching area; Step 4: Determine the boundary of the difference file; Step 5: Write control files, difference files, and new files; At the shore end, an incremental update file is generated from the old file and the new file through the above steps, and the incremental update file is transmitted to the ship end, where a new file is generated from the old file and the incremental update file.

2. The ship incremental update method based on the improved Bsdiff algorithm according to claim 1 is characterized in that: The specific steps are as follows: The first step is to generate a suffix array based on the old file; Build an array based on the old file, use the EM-SA-DS algorithm to sort the suffixes of the array built from the old file, and generate a suffix array; The second step is to initialize the difference file; The header file is 32 bytes in total and is divided into four parts. The first part is the file format signature, which is an 8-byte constant "BSDIFF40". The second part is the length of the control file compression stream, which is an 8-byte integer. The third part is the length of the difference file compression stream, which is an 8-byte integer; the fourth part is the length of the new file, which is an 8-byte integer; After the header file are compressed streams of the control file, difference file, and newly added files; The third step is to compare the new file with the generated suffix array to generate an approximate matching area; Initialize an empty approximate matching area in the old file starting from the end of the previous segment; start from the end of the previous segment of the new file, start to compare with the suffix array at each byte, and find a section in the old file that is exactly the same as the new file, with a length of len; compare the fields of length len at the beginning of the approximate matching area of ​​the new file and the old file. If they are exactly the same, the next time the new file is compared directly from len bytes; if there are more than 8 bytes that are different, or the end of the new file has been reached, then this round ends and the approximate matching area of ​​the old file is output; otherwise, the comparison continues from the next byte of the new file; Step 4: Determine the boundary of the difference file; For the approximate matching area generated in the third step, compare the beginning of the area with the beginning of the new file to obtain a segment lenf with a 50% byte identity rate; compare the end of the area with the end of the new file to obtain a segment lenb with a 50% byte identity rate; if the two segments overlap, find a position in the overlapping part to maximize the number of bytes in the approximate matching area as the boundary; Step 5: Write control files, difference files, and new files; The control file is an array, each element of which is a triple (x, y, z) consisting of three 8-byte integers; x is the size of each difference file; y is the size of each newly added file, which may be 0; z is the deviation value, which is used to adjust the current position of the old file, adding z bytes, which may be a negative number; the difference of the content corresponding to segment lenf is calculated and compressed and written to the difference file; if segment lenf and segment lenb are not connected, the middle area is directly compressed and written to the newly added file, if they are connected, it does not exist; the beginning of segment lenb is used as the beginning of the next round of approximate matching area; the control file, difference file and newly added file are generated into a new file; At the shore end, an incremental update file is generated from the old file and the new file through the above steps, and the incremental update file is transmitted to the ship end, where the old file and the incremental update file are generated.

3. The ship incremental update method based on the improved Bsdiff algorithm according to claim 2 is characterized in that: In the first step, in the EM-SA-DS algorithm, in order to obtain the information of the previous characters of s[sa[i]] from the external memory without causing a large amount of random I / O, the concept of substring buffer is proposed, which is defined as: a substring buffer is a string containing d+2 triples<ch,t,dc> are used to store the starting character of a suffix and the related information of its d+1 preceding characters, and the first triplet stores the related information of the character s[sa[i]], the second triplet stores the related information of the character s[sa[i]-1], and so on; among them, ch represents the character value, t represents whether ch is L type or S type, and dc represents whether ch is a d-critical character.

4. The ship incremental update method based on the improved Bsdiff algorithm according to claim 3 is characterized in that: In the EM-SA-DS algorithm, let xa and ya be two ordered sets, where xa is the set of all L-type d-critical suffixes arranged in descending order, and ya is the set of all S-type d-critical suffixes arranged in ascending order; the specific process of inductive sorting of EM-SA-DS is as follows: Step 1.1, calculate strbuf(sa,i), strbuf(xa,i) and strbuf(ya,i) of all non-empty suffixes in sa, xa and ya respectively; Step 1.

2. Scan each non-empty element sa[i] in sa from left to right, and let j = sa[i] - 1: Case 1: If s[sa[i]] is an L-type d-critical character, then copy the substring buffer of the suffix suf(s, sa[i]) from strbuf(xa) to strbuf(sa), and at the same time delete the substring buffer of the suffix suf(s, sa[i]) from strbuf(xa); Case 2: Obtain the ch, t, and dc information of the character s[sa[j]] from strbuf(sa, i); If s[j] is of L-type, put suf(s, j) into the leftmost empty position of its bucket bucket(sa, s[j]), let the inserted position be j', and update strbuf(sa, j') accordingly according to strbuf(sa, i). Step 1.

3. Scan each non-empty element sa[i] in sa from right to left, and let j = sa[i] - l: Case 1: If s[sa[i]] is an S-type d-critical character, then copy the substring buffer of the suffix suf(s, sa[i]) from strbuf(ya) to strbuf(sa), and at the same time delete the substring buffer of the suffix suf(s, sa[i]) from strbuf(ya); Case 2: Obtain the ch, t, and dc information of the character s[sa[j]] from strbuf(sa, i); If s[j] is of S-type, put suf(s, j) into the rightmost empty position of its bucket bucket(sa, s[j]), let the inserted position be j', and update strbuf(sa, j') accordingly according to strbuf(sa, i).

5. The ship incremental update method based on the improved Bsdiff algorithm according to claim 2 is characterized in that: In the first step, all suffixes are rearranged in ascending lexicographical order, and the integer array formed by using their respective starting subscripts to represent the suffix strings is the suffix array.

6. The ship incremental update method based on the improved Bsdiff algorithm according to claim 5 is characterized in that: In the first step, a suffix refers to the string formed from a certain position to the last position in the string.

7. The ship incremental update method based on the improved Bsdiff algorithm according to claim 5 is characterized in that: In the first step, the lexicographical order means that for two strings s1 and s2, if s1 < s2, currently only when: for the first m characters, s1[i] == s2[i]; for the length of s1 being m, or for the (m + 1)-th character, s1[m + 1] < s2[m + 1].

8. The ship incremental update method based on the improved Bsdiff algorithm according to claim 2 is characterized in that: In the first step, the main idea of the EM-SA-DS algorithm: First, select all d-critical substrings in the string s and sort them to obtain the reduced string s1; then, according to the size of the reduced string s1, call the in-memory construction algorithm of the suffix array or recursively call the EM-SA-DS algorithm to calculate the suffix array sa1 of the string s1. Finally, derive the suffix array sa of s through sa1.

9. The ship incremental update method based on the improved Bsdiff algorithm according to claim 2 is characterized in that: In the first step, the EM-SA-DS algorithm must know the predecessor of each character s[sa[i]] in s during the sorting process and insert it into the starting position of the corresponding bucket. If all operations are performed in memory, the predecessor and bucket related information of the character can be obtained directly from the memory.

10. The ship incremental update method based on the improved Bsdiff algorithm according to claim 9 is characterized in that: When the length n of string s is very large, all this information cannot be stored in the memory and must be stored in external memory.