Transmission data compression method and system based on cloud computing

By obtaining the global redundant features of high-frequency codes and redundant substrings, the sliding window length of the LZ77 compression algorithm is optimized, and the problem of low compression efficiency of the LZ77 compression algorithm in system log data with high global repetition is solved, achieving more efficient data compression and transmission.

CN120342401APending Publication Date: 2025-07-18SHENZHEN ENCYCLOPEDIA ZHIYUN TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510397734.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When the existing LZ77 compression algorithm processes system log data with high global repetition, the compression efficiency is not high enough and fails to fully utilize the global similarity characteristics of the data.

Method used

By obtaining the global and local occurrences of high-frequency codes, calculating the global redundant high frequency and redundant matching density, obtaining the redundant distance of the redundant substrings, global compression is used to optimize the sliding window length, and global compression of system log data is achieved.

Benefits of technology

The compression efficiency of system log data is improved, the problem of insufficient compression efficiency of data with high global repetition in the prior art is solved, and more efficient data transmission is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120342401A_ABST
    Figure CN120342401A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of electric digital data processing, and provides a transmission data compression method and system based on cloud computing, and the method comprises the steps: collecting system log data, obtaining a to-be-compressed character string and a to-be-compressed character string set, obtaining a high-frequency code according to the to-be-compressed character string set, and calculating the global redundancy high frequency of the to-be-compressed character string; according to the global redundancy high frequency and the high frequency code of the to-be-compressed character string, the redundancy matching density and the global compression gain degree of the to-be-compressed character string are obtained, then the to-be-compressed character string is subjected to global compression according to the global compression gain degree of the to-be-compressed character string, and then the global compression character string is compressed by using an LZ77 compression algorithm. And then transmission of the system log data is completed. The method aims at solving the problems that an existing LZ77 compression algorithm only considers local similarity, and the compression efficiency of system log data with high global repeatability is not high enough.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electronic digital data processing, and particularly relates to a method and system for compressing transmission data based on cloud computing. Background Art

[0002] With the continuous improvement of computer hardware structures, computer systems have become increasingly complex, and the operation and maintenance management of cloud service platforms has received much attention. Among them, daily maintenance work such as fault diagnosis, abnormal risk, and security inspection of large computer platform systems and related devices is inseparable from system logs. For modern data centers, the data scale of system logs is extremely large, reaching tens of billions of entries (PB level). The storage and transmission costs of such data are relatively high. Therefore, it is necessary to compress such data quickly and densely.

[0003] Currently, the LZ77 compression algorithm is often used to compress text data by reducing the amount of data transmitted, so as to achieve the purpose of improving transmission efficiency. The advantage of the LZ77 compression algorithm lies in fully considering the local similarity characteristics of data and having a good compression effect. However, the global repeatability of system log data is relatively high, resulting in low compression efficiency of the existing LZ77 compression algorithm for some system logs. Summary of the Invention

[0004] In order to solve the above technical problems, the purpose of the present invention is to provide a method and system for compressing transmission data based on cloud computing, and the specific technical solutions adopted are as follows:

[0005] In a first aspect, an embodiment of the present invention provides a method for compressing transmission data based on cloud computing, and the method includes the following steps:

[0006] Collect system log data to obtain a string to be compressed and a set of strings to be compressed;

[0007] Obtain high-frequency codes according to the set of strings to be compressed;

[0008] Record the number of occurrences of the high-frequency codes in the set of strings to be compressed as the global number of occurrences of the high-frequency codes;

[0009] Obtain the local number of occurrences of the high-frequency codes in the string to be compressed according to the number of occurrences of the high-frequency codes in the string to be compressed;

[0010] Calculate the global redundant frequency of the string to be compressed according to the local number of occurrences of the high-frequency codes in the string to be compressed, the global number of occurrences of the high-frequency codes, and the information entropy of the string to be compressed;

[0011] Obtain global redundant strings according to the global redundant frequencies of all strings to be compressed;

[0012] Obtain a redundant substring according to the high-frequency codes in the global redundant string; record the number of high-frequency codes in the redundant substring as the length of the redundant substring;

[0013] According to the string to be compressed and the redundant substring, obtain the redundant distance of the redundant substring and the redundant matching density of the string to be compressed;

[0014] According to the redundant distance of the redundant substring and the redundant matching density of the string to be compressed, obtain the global compression gain of the string to be compressed;

[0015] Compress the string to be compressed according to the global compression gain of the string to be compressed and the length of the redundant substring, and complete the transmission of the system log data.

[0016] Preferably, the obtaining of the high-frequency codes according to the set of strings to be compressed includes:

[0017] Count the occurrence times of each ASCII code in the set of strings to be compressed, and record the first preset number of ASCII codes with the most occurrence times as the high-frequency codes.

[0018] Preferably, the obtaining of the local occurrence times of the high-frequency codes in the string to be compressed according to the occurrence times of the high-frequency codes in the string to be compressed includes:

[0019] Record the occurrence times of the high-frequency codes in the string to be compressed as the local occurrence times of the high-frequency codes in the string to be compressed.

[0020] Preferably, the calculating of the global redundant frequency of the string to be compressed according to the local occurrence times, global occurrence times of the high-frequency codes in the string to be compressed and the information entropy of the string to be compressed includes:

[0021] Record the maximum value of the information entropy of all strings to be compressed as the maximum string information entropy;

[0022] Record the product of the local occurrence times and the global occurrence times of the high-frequency codes in the string to be compressed as the frequent occurrence times of the high-frequency codes in the string to be compressed;

[0023] Record the sum of the frequent occurrence times of all high-frequency codes in the string to be compressed as the first cumulative sum;

[0024] Record the ratio of the information entropy of the string to be compressed to the maximum string information entropy as the first ratio;

[0025] Record the product of the first cumulative sum and the first ratio as the first product;

[0026] Record the normalized value of the first product as the global redundant frequency of the string to be compressed.

[0027] Preferably, obtaining the global redundant string according to the global redundant high frequency of all strings to be compressed includes:

[0028] Regarding the strings to be compressed with a global redundant high frequency greater than a preset global threshold as global redundant strings.

[0029] Preferably, obtaining the redundant substring according to the high-frequency codes in the global redundant string includes:

[0030] Combining continuously adjacent high-frequency codes in the global redundant string into a redundant substring.

[0031] Preferably, obtaining the redundant distance of the redundant substring and the redundant matching density of the string to be compressed according to the string to be compressed and the redundant substring includes:

[0032] Using a string matching algorithm to match each redundant substring with each string to be compressed, and recording the number of successful matches between the redundant substring and the string to be compressed as the successful match count of the redundant substring;

[0033] Sorting the redundant substrings that successfully match the string to be compressed in descending order of the successful match count to obtain a redundant substring sequence;

[0034] Deleting the redundant substrings in the redundant substring sequence whose successful match count with the string to be compressed is less than a preset match threshold;

[0035] Recording the position where the redundant substring successfully matches the string to be compressed as the matching position of the redundant substring;

[0036] Recording the minimum value of the Euclidean distances between all adjacent two matching positions of the redundant substring as the redundant distance of the redundant substring;

[0037] Recording the product of the length of the redundant substring and the successful match count as the coverage of the redundant substring;

[0038] Recording the sum of the coverages of all redundant substrings in the redundant substring sequence as the redundant coverage of the string to be compressed;

[0039] Recording the product of the redundant coverage of the string to be compressed and the global redundant high frequency as the first numerator;

[0040] Recording the ratio of the first numerator to the length of the string to be compressed as the redundant matching density of the string to be compressed.

[0041] Preferably, obtaining the global compression gain of the string to be compressed according to the redundant distance of the redundant substring and the redundant matching density of the string to be compressed includes:

[0042] Recording the sum of the redundant distances of all redundant substrings in the redundant substring sequence as the second cumulative sum;

[0043] Denote the ratio of the second cumulative sum to the redundancy matching density of the string to be compressed as the second ratio;

[0044] Denote the normalized value of the second ratio as the global compression gain of the string to be compressed.

[0045] Preferably, compressing the string to be compressed according to the global compression gain of the string to be compressed and the length of the redundant substring to complete the transmission of the system log data includes:

[0046] Globally compress the string to be compressed with a global compression gain greater than the global compression threshold;

[0047] The steps of global compression are: first sort all redundant substrings in descending order of the length of the redundant substring, obtain the serial number of each redundant substring, and then replace the redundant substring in the string to be compressed with the serial number of the redundant substring to obtain the globally compressed string;

[0048] Set the sliding window length in the compression algorithm to the maximum value of the lengths of all redundant substrings;

[0049] Use the globally compressed string as the input of the compression algorithm to obtain the compression result of the system log data;

[0050] Encrypt the compression result using a security protocol, and then transmit the encrypted compression result using a transmission protocol.

[0051] In a second aspect, an embodiment of the present invention further provides a transmission data compression system based on cloud computing, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the method described in any one of the above are implemented.

[0052] The present invention has at least the following beneficial effects: First, according to the set of strings to be compressed, high-frequency codes are obtained. Considering both the local similarity and the global similarity of the high-frequency codes, based on the global occurrence times and local occurrence times of the high-frequency codes, and in combination with the information entropy of the strings to be compressed, the global redundant high frequency of the strings to be compressed is obtained to measure the similarity between the strings to be compressed and other strings to be compressed, improving the accuracy of the evaluation of the redundancy degree of the strings to be compressed; According to the global redundant high frequency and high-frequency codes of the strings to be compressed, redundant substrings are obtained. According to the matching result between the strings to be compressed and the redundant substrings, the global compression gain degree of the strings to be compressed is obtained. According to the global compression gain degree of the strings to be compressed, the strings to be compressed are first globally compressed to obtain globally compressed strings, and then, according to the length of the redundant substrings, the sliding window length of the LZ77 compression algorithm is obtained, and the LZ77 compression algorithm is used to compress the globally compressed strings to complete the transmission of system log data, improving the compression efficiency and solving the problem that the existing LZ77 compression algorithm only considers local similarity and has a low compression efficiency for system log data with high global repeatability. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0054] Figure 1 FIG. is a flowchart of the steps of a transmission data compression method based on cloud computing provided by an embodiment of the present invention;

[0055] Figure 2 FIG. is a schematic diagram of redundancy distance. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following, in combination with the drawings and preferred embodiments, details the specific implementation manners, structures, features, and effects of the transmission data compression method and system based on cloud computing proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0057] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.

[0058] The following specifically describes the specific solutions of the transmission data compression method and system based on cloud computing provided by the present invention in conjunction with the accompanying drawings.

[0059] Please refer to Figure 1 , which shows the flowchart of the steps of the transmission data compression method based on cloud computing provided by an embodiment of the present invention. The method includes the following steps:

[0060] Step S001, collect system log data, preprocess the system log data, and obtain the string to be compressed and the set of strings to be compressed.

[0061] The main purpose of this embodiment is to compress and transmit system log data. Therefore, it is necessary to first obtain system log data from the cloud of the data center, including relevant information such as kernel status data, non-core component data, user data, traffic records, etc. To improve the compression efficiency, the system log data is converted into a string to be compressed, and the set containing all strings to be compressed is denoted as the set of strings to be compressed.

[0062] Thus, the string to be compressed and the set of strings to be compressed are obtained.

[0063] Step S002, obtain the high-frequency codes according to the set of strings to be compressed, and obtain the global occurrence times of the high-frequency codes according to the occurrence times of the high-frequency codes in the set of strings to be compressed.

[0064] It should be noted that the LZ77 compression algorithm compresses the string to be compressed by replacing the locally repeated character combinations in the string to be compressed with shorter codes. There may be a lot of duplicate data between different system log data, such as cross or partially overlapping user names, kernel versions, driver protocols, etc. Information will not be compressed because the corresponding character combinations are not in the same string to be compressed, resulting in low compression efficiency. Before compressing the set of strings to be compressed, first analyze the similarity between each string to be compressed in the set of strings to be compressed and the rest of the strings to be compressed, that is, the global redundancy degree of each string to be compressed.

[0065] Specifically, count the occurrence times of each ASCII code in the set of strings to be compressed, and record the top n ASCII codes with the most occurrence times as the high-frequency codes. The empirical value of n is 50. The occurrence times of the high-frequency codes in the set of strings to be compressed are denoted as the global occurrence times of the high-frequency codes.

[0066] Thus, the global occurrence times of the high-frequency codes in the string to be compressed are obtained.

[0067] Step S003, calculate the global redundant frequency of the string to be compressed according to the global occurrence times of the high-frequency codes in the string to be compressed and the information entropy of the string to be compressed.

[0068] It should be noted that the information entropy of the string to be compressed reflects the degree of chaos of ASCII codes in the string to be compressed. If the degree of chaos of ASCII codes in the string to be compressed is greater and the repeatability is lower, it is more difficult to obtain a good compression effect through the local matching mode of the LZ77 compression algorithm.

[0069] The information entropy of the string to be compressed reflects the degree of chaos of characters in the string to be compressed and can be used to measure the redundancy degree of information in the string to be compressed.

[0070] Specifically, the maximum value of the information entropy of all strings to be compressed is denoted as the maximum string information entropy MAXHQ. The information entropy of the string to be compressed reflects the degree of chaos of characters in the string to be compressed, and the maximum string information entropy reflects the degree of chaos of characters in all strings to be compressed. The smaller the degree of chaos, the greater the redundancy degree. By obtaining the ratio of the information entropy of each string to be compressed to the maximum string information entropy, it is used to measure the redundancy degree of each string to be compressed relative to the whole of all strings to be compressed. The sum of the global occurrence times of all high-frequency codes in the string to be compressed reflects the absolute redundancy degree of the string to be compressed. Multiply the ratio of the information entropy of the string to be compressed to the maximum string information entropy by the sum of the global occurrence times of all high-frequency codes in the string to be compressed, comprehensively consider the relative redundancy degree and absolute redundancy degree of the string to be compressed, and obtain the global redundancy high frequency to measure the global redundancy degree of each string to be compressed. Global compression of strings to be compressed with relatively high relative redundancy degree and absolute redundancy degree can obtain a better compression effect. Among them, the acquisition of information entropy is a well-known technology, and this embodiment will not elaborate here.

[0071] According to the global occurrence times of high-frequency codes in the string to be compressed, the information entropy of the string to be compressed, and the maximum string information entropy, the global redundancy high frequency of the string to be compressed is expressed as follows:

[0072]

[0073] Among them, γ i is the global redundancy high frequency of the i-th string to be compressed; exp() is the exponential function with the natural constant as the base; V i a is the global occurrence times of the a-th high-frequency code in the i-th string to be compressed; H i is the information entropy of the i-th string to be compressed; MAXHQ is the maximum string information entropy; x i is the number of high-frequency codes in the i-th string to be compressed.

[0074] It should be further explained that, when the cumulative number of high-frequency codes in the string to be compressed increases, it means that the similarity between the string to be compressed and the other strings to be compressed is higher, the global redundancy high frequency value is larger, and the global redundancy of the string to be compressed is greater; when the ratio of the information entropy of the string to be compressed to the information entropy of the maximum string is smaller, it means that the degree of confusion inside the string to be compressed is relatively lower, the repeatability is relatively higher, the local similarity of the string to be compressed is greater, the better the compression effect can be obtained by local compression, and the global redundancy high frequency value is smaller.

[0075] At this point, the global redundancy high frequency of the character string to be compressed is obtained.

[0076] Step S004, obtaining redundant substrings according to the global redundant high frequency and high frequency code of the character string to be compressed, and obtaining a redundant substring sequence according to the character string to be compressed and the redundant substrings.

[0077] It should be noted that different system log data contain different levels of information redundancy. When compressing system log data, giving priority to analyzing the character strings to be compressed with relatively high redundancy can greatly improve the compression efficiency and make up for the limitations of the LZ77 compression algorithm.

[0078] Specifically, the character string to be compressed whose global redundancy frequency is greater than the global threshold α is recorded as the global redundant character string, and the empirical value of the global threshold α is 0.7. The continuous adjacent high-frequency codes in the global redundant character string form a redundant substring, and the number of high-frequency codes in the redundant substring is recorded as the length of the redundant substring. Using the kmp string matching algorithm, each redundant substring is matched with each character string to be compressed, and the number of successful matches between the redundant substring and the character string to be compressed is obtained. The number of successful matches between the redundant substring and the character string to be compressed is recorded as the number of successful matches of the redundant substring on the character string to be compressed. The redundant substrings that successfully match the character string to be compressed are sorted from most to least according to the number of successful matches to obtain a redundant substring sequence, and the redundant substrings in the redundant substring sequence whose number of successful matches with the character string to be compressed is less than the matching threshold ω are eliminated, and the empirical value of the matching threshold ω is 2.

[0079] Furthermore, the position where the redundant substring successfully matches the string to be compressed is recorded as the matching position of the redundant substring, and the minimum value of the Euclidean distance between all two adjacent matching positions of the redundant substring is recorded as the redundant distance of the redundant substring. The redundant distance diagram is shown in FIG. Figure 2 As shown, where D(W b ) represents the redundant substring W b The redundant distance, redundant substring W b The Euclidean distance between matching position 1 and matching position 2 is the smallest.

[0080] So far, the redundant substring sequence is obtained.

[0081] Step S005: According to the redundant substring sequence and the global redundant high frequency of the string to be compressed, obtain the redundant matching density of the string to be compressed, and calculate the global compression gain of the string to be compressed.

[0082] Specifically, according to the redundant substring sequence and the global redundant high frequency of the string to be compressed, the redundant matching density of the string to be compressed is expressed as follows:

[0083]

[0084] where ρ i is the redundant matching density of the i-th string to be compressed; γ i is the global redundant high frequency of the i-th string to be compressed; k i is the number of redundant substrings in the i-th redundant substring sequence; is the length of the y-th redundant substring in the i-th redundant substring sequence; P i y is the successful matching times of the y-th redundant substring in the i-th redundant substring sequence; Ld i is the length of the i-th string to be compressed.

[0085] It should be noted that when the proportion of the total length of all redundant substrings in the redundant substring sequence to the length of the string to be compressed is larger and the global redundant high frequency is larger, it indicates that there are more redundant substrings in the string to be compressed, and the redundant matching density value is larger.

[0086] Furthermore, according to the redundant matching density of the string to be compressed and the redundant distance of the redundant substring, the global compression gain of the string to be compressed is expressed as follows:

[0087]

[0088] where Z i is the global compression gain of the i-th string to be compressed; ρ i is the redundant matching density of the i-th string to be compressed; is the redundant distance of the y-th redundant substring in the i-th redundant substring sequence; exp() is the exponential function with the natural constant as the base; k i is the number of redundant substrings in the i-th redundant substring sequence.

[0089] It should be noted that when the redundancy matching density of the string to be compressed is higher and the distance between redundant substrings is smaller, it indicates that the local repetition degree of the string to be compressed is higher, and the global compression gain is smaller. In the LZ77 compression algorithm, a better compression effect can be achieved by using a smaller sliding window, and the gain effect of global compression on the string to be compressed is not obvious.

[0090] Thus, the global compression gain of the string to be compressed is obtained.

[0091] Step S006: According to the global compression gain of the string to be compressed, perform global compression on the string to be compressed to obtain a globally compressed string. According to the length of the redundant substring, obtain the length of the sliding window of the LZ77 compression algorithm, and use the LZ77 compression algorithm to compress the globally compressed string to complete the transmission of the system log data.

[0092] Specifically, first perform global compression on the strings to be compressed whose global compression gain is greater than the global compression threshold t. The empirical value of the global compression threshold t is 0.6. The steps of global compression are as follows: First, sort all redundant substrings in descending order of their lengths to obtain the serial numbers of each redundant substring, and then replace the corresponding redundant substrings in the string to be compressed with the serial numbers of the redundant substrings to obtain a globally compressed string.

[0093] It should be noted that in the LZ77 compression algorithm, the compression efficiency of the string to be compressed is directly related to the size of the sliding window. A larger sliding window can better find the locally repeated substrings in the string to be compressed to achieve a better compression effect, but it will increase the time required for compression and reduce the compression efficiency. Based on the analysis of the principle of the LZ77 compression algorithm, the closer the distance between the repeatedly occurring redundant substrings in the string to be compressed, the smaller the sliding window used, and the higher the local compression efficiency. If the distance between the repeatedly occurring redundant substrings in the string to be compressed is farther, it is better to first perform global compression encoding on the redundant substrings in the string to be compressed.

[0094] Furthermore, set the length of the sliding window in the LZ77 compression algorithm to the maximum value of the lengths of all redundant substrings, use the globally compressed string as the input of the LZ77 compression algorithm to obtain the compression result of the system log data, encrypt the compression result using the SSL security protocol to ensure data security, and then use the HTTP transmission protocol to transmit the data.

[0095] Thus, the transmission of the system log data is completed.

[0096] Based on the same inventive concept as the above method, this embodiment also provides a transmission data compression system based on cloud computing, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of any one of the above methods for transmitting data compression based on cloud computing.

[0097] In summary, this embodiment solves the problem that the existing LZ77 compression algorithm only considers local similarity and the compression efficiency of system log data with high global repeatability is not high enough. In this embodiment, the system log data is converted into a string to be compressed, the redundancy degree of the string to be compressed is analyzed, the redundancy matching density of the string to be compressed and the redundancy distance of the redundant substring are obtained, the global compression gain of the string to be compressed is calculated, the string to be compressed greater than the global compression threshold is first globally compressed to obtain a globally compressed string, then the sliding window length in the LZ77 compression algorithm is obtained according to the lengths of all redundant substrings, the globally compressed string is further compressed, and then the obtained compression result is encrypted and transmitted to complete the transmission of the system log data.

[0098] It should be noted that the above sequence of this embodiment is only for description and does not represent the superiority or inferiority of the embodiment. In addition, the above specific embodiments of this specification have been described. Also, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0099] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for compressing transmitted data based on cloud computing, characterized in that The method includes the following steps: Collect system log data, and obtain the string to be compressed and the set of strings to be compressed; Obtain high-frequency codes according to the set of strings to be compressed; Record the number of occurrences of the high-frequency codes in the set of strings to be compressed as the global number of occurrences of the high-frequency codes; Obtain the local number of occurrences of the high-frequency codes in the string to be compressed according to the number of occurrences of the high-frequency codes in the string to be compressed; Calculate the global redundant frequency of the string to be compressed according to the local number of occurrences of the high-frequency codes in the string to be compressed, the global number of occurrences, and the information entropy of the string to be compressed; Obtain the global redundant string according to the global redundant frequency of all strings to be compressed; Obtain the redundant substring according to the high-frequency codes in the global redundant string; record the number of high-frequency codes in the redundant substring as the length of the redundant substring; Obtain the redundant distance of the redundant substring and the redundant matching density of the string to be compressed according to the string to be compressed and the redundant substring; Obtain the global compression gain of the string to be compressed according to the redundant distance of the redundant substring and the redundant matching density of the string to be compressed; Compress the string to be compressed according to the global compression gain of the string to be compressed and the length of the redundant substring, and complete the transmission of the system log data.

2. The transmission data compression method based on cloud computing according to claim 1, wherein The obtaining of high-frequency codes according to the set of strings to be compressed includes: Count the number of occurrences of each ASCII code in the set of strings to be compressed, and record the first preset number of ASCII codes with the most occurrences as the high-frequency codes.

3. The transmission data compression method based on cloud computing according to claim 1, wherein The obtaining of the local number of occurrences of the high-frequency codes in the string to be compressed according to the number of occurrences of the high-frequency codes in the string to be compressed includes: Record the number of occurrences of the high-frequency codes in the string to be compressed as the local number of occurrences of the high-frequency codes in the string to be compressed.

4. The transmission data compression method based on cloud computing according to claim 1, wherein The calculating of the global redundant frequency of the string to be compressed according to the local number of occurrences of the high-frequency codes in the string to be compressed, the global number of occurrences, and the information entropy of the string to be compressed includes: Record the maximum value of the information entropy of all strings to be compressed as the maximum string information entropy; Record the product of the local number of occurrences and the global number of occurrences of the high-frequency codes in the string to be compressed as the frequent occurrence number of the high-frequency codes in the string to be compressed; Record the sum of the frequent occurrence numbers of all high-frequency codes in the string to be compressed as the first cumulative sum; Record the ratio of the information entropy of the string to be compressed to the maximum string information entropy as the first ratio; Record the product of the first cumulative sum and the first ratio as the first product; Record the normalized value of the first product as the global redundant frequency of the string to be compressed.

5. The transmission data compression method based on cloud computing according to claim 1, characterized in that The obtaining of the global redundant string according to the global redundant frequency of all strings to be compressed includes: Record the strings to be compressed with a global redundant frequency greater than the preset global threshold as the global redundant strings.

6. The transmission data compression method based on cloud computing according to claim 1, wherein The obtaining of the redundant substring according to the high-frequency codes in the global redundant string includes: Form a redundant substring by the continuously adjacent high-frequency codes in the global redundant string.

7. The method for compressing transmitted data based on cloud computing according to claim 1, wherein The obtaining of the redundant distance of the redundant substring and the redundant matching density of the string to be compressed according to the string to be compressed and the redundant substring includes: Use a string matching algorithm to match each redundant substring with each string to be compressed, and record the number of successful matches between the redundant substring and the string to be compressed as the number of successful matches of the redundant substring; Sort the redundant substrings that successfully match the string to be compressed from most to least according to the number of successful matches to obtain a redundant substring sequence; Deleting redundant substrings in the redundant substring sequence whose number of successful matches with the character string to be compressed is less than a preset matching threshold; The position where the redundant substring successfully matches the character string to be compressed is recorded as the matching position of the redundant substring; The minimum value of the Euclidean distance between all two adjacent matching positions of the redundant substring is recorded as the redundant distance of the redundant substring; The product of the length of the redundant substring and the number of successful matches is recorded as the coverage of the redundant substring; The sum of the coverage of all redundant substrings in the redundant substring sequence is recorded as the redundant coverage of the string to be compressed; The product of the redundancy coverage of the string to be compressed and the global redundancy high frequency is recorded as the first numerator; The ratio of the length of the first numerator to the length of the string to be compressed is recorded as the redundant matching density of the string to be compressed.

8. The transmission data compression method based on cloud computing according to claim 1, characterized in that, The method obtains the global compression gain of the string to be compressed according to the redundant distance of the redundant substrings and the redundant matching density of the string to be compressed, include: The sum of the redundant distances of all redundant substrings in the redundant substring sequence is recorded as the second cumulative sum; Recording the ratio of the second accumulated sum to the redundant matching density of the character string to be compressed as a second ratio; The normalized value of the second ratio is recorded as the global compression gain of the character string to be compressed.

9. The method for compressing transmission data based on cloud computing according to claim 1, characterized in that, The method of compressing the string to be compressed according to the global compression gain of the string to be compressed and the length of the redundant substring to complete the transmission of the system log data includes: Perform global compression on the character string to be compressed whose global compression gain is greater than the global compression threshold; The steps of global compression are: first, sort all redundant substrings from large to small according to the length of the redundant substrings, obtain the sequence number of each redundant substring, and then replace the redundant substrings in the string to be compressed with the sequence number of the redundant substring to obtain the global compressed string; Set the sliding window length in the compression algorithm to the maximum length of all redundant substrings; Use the global compression string as the input of the compression algorithm to obtain the compression result of the system log data; The compression result is encrypted using a security protocol, and then the encrypted compression result is transmitted using a transmission protocol.

10. A transmission data compression system based on cloud computing, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 are implemented.

Citation Information

Cited By

  • Genome sequencing data analysis method based on next-generation sequencing technology

    CN121075424A

  • A method for analyzing genomic sequencing data based on next-generation sequencing technology

    CN121075424B