Data segmentation compression method, module and electronic equipment

The compression unit value is determined through differential calculation, the sliding compression window length is adjusted, and segmented compression is performed for fragments with high similarity to data streams in the chip. The LM77 and the run encoding algorithm are used to solve the problem of low data compression rate in the chip and improve the compression rate and efficiency.

CN120389756APending Publication Date: 2025-07-29SHANGHAI ANLOGIC INFOTECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510524018.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In the prior art, in the process of compressing data information in the chip, there is a problem of low compression rate, especially since the high-position identical parts of adjacent data information are not effectively utilized.

Method used

The compression unit value is determined through differential calculation, the length of the sliding compression window is adjusted, and the segmented compression is performed for data streams with high similarity is performed, and the encoding and compression is performed using LM77 encoding and run encoding algorithms.

Benefits of technology

It improves data compression rate and compression calculation efficiency, is suitable for raw data streams in various dense situations, and achieves higher data compression rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120389756A_ABST
    Figure CN120389756A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data processing, and relates to a data segmentation compression method and module and electronic equipment, and the method comprises the steps: respectively carrying out the difference calculation of a timestamp, a storage address, an order code and a mantissa of data information in an original data stream, obtaining a difference calculation result, and determining a compression unit value; according to a sliding compression window corresponding to the compression unit value, performing first coding compression on a timestamp, a storage address, an order code and a mantissa of data information in the original data stream to obtain first compressed data; performing second coding compression on the data type in the original data stream to obtain second compressed data; and obtaining compressed data of the data information according to the first compressed data and the second compressed data corresponding to the same data information in the original data stream, and outputting the compressed data. The compression unit value is determined through the difference calculation result so as to determine the length of the sliding compression window, so that the similarity of segmented data information of the sliding compression window is improved, segmented self-compression is facilitated, and the compression rate is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to a data segmentation compression method, module and electronic equipment. Background Art

[0002] With the development of modern information technology, the number of processor cores in chip design is increasing, chip technology is becoming more and more sophisticated, and the complexity of chip calculations is also increasing. This has led to a sharp increase in the amount of data that needs to be processed when the chip is running, and the importance of data compression has become increasingly prominent.

[0003] The data information collected by the chip usually includes data type, timestamp, storage address, and data content. The data content can be floating-point data or fixed-point data. Generally, the data types of the previous and next data information in the data stream are the same or similar, the timestamps are continuous, and the storage addresses are adjacent. And when the data content is floating-point data, the exponents of the floating-point data of the previous and next data information in the data stream are also continuous. Based on this, the process of compressing these data information often involves comparing each data information before and after, and judging the compression type based on the comparison results. Then, the data information is compressed accordingly according to the judgment result of the compression type. However, this method essentially uses the compression type to compress the data information itself. The compression process also compresses the same high-bit parts of adjacent data information in binary (i.e., redundant information), resulting in a low compression rate. Summary of the invention

[0004] The purpose of this application is to provide a data segmentation compression method, which can segment fragments of a data stream with high similarity and compress each segment independently to improve the overall compression efficiency.

[0005] According to a first aspect of an embodiment of the present application, a data segmentation compression method is provided, comprising: Obtain the original data stream obtained by collecting the target signal; Performing differential calculations on the timestamp, storage address, exponent, and mantissa of the data information in the original data stream to obtain differential calculation results; Determining a compression unit value according to a sparse coefficient obtained from the difference calculation result; According to the sliding compression window corresponding to the compression unit value, the timestamp, storage address, exponent and mantissa of the data information in the original data stream are respectively subjected to first coding compression to obtain first compressed data; Performing second coding compression on the data type in the original data stream to obtain second compressed data; Obtain the compressed data of the data information according to the first compressed data and the second compressed data corresponding to the same data information in the original data stream, and output the compressed data.

[0006] Through the above technical solution, since the data types of the data information before and after in the original data stream of the target signal transmitted in the chip are the same or similar, the timestamps are continuous, the storage addresses are adjacent, and the data content is that the exponent of the floating-point data is continuous, the compression method in this case is feasible. In this case, the compression unit value is determined through the differential calculation result, and then the length of the sliding compression window is changed, so as to adjust the sliding compression window according to the sparsity of the data information distribution in the original data stream, ensuring that the sliding compression window is applicable to the original data stream with various density situations.

[0007] Segment the original data stream through the sliding compression window. The similarity of the data information within each segment corresponding to the sliding compression window each time is higher. The timestamps, storage addresses, exponents, and mantissas of the data information within the segment corresponding to the sliding compression window are compressed by the first encoding to obtain the first compressed data. Compared with directly compressing the data information in the original data stream, the first compressed data has a higher data compression ratio.

[0008] In one implementation manner, the differential calculations are respectively performed on the timestamps, storage addresses, exponents, and mantissas of the data information in the original data stream to obtain differential calculation results, including: Calculate the difference between the timestamp of the latter data information and the timestamp of the former data information in the original data stream to obtain the timestamp difference value corresponding to the latter data information; Calculate the difference between the storage address of the latter data information and the storage address of the former data information in the original data stream to obtain the storage address difference value corresponding to the latter data information; Calculate the difference between the exponent of the latter data information and the exponent of the former data information in the original data stream to obtain the exponent difference value corresponding to the latter data information; Calculate the difference between the mantissa of the latter data information and the mantissa of the former data information in the original data stream to obtain the mantissa difference value corresponding to the latter data information; Among them, the differential calculation results include the timestamp difference value, storage address difference value, storage address difference value, exponent difference value, and mantissa difference value corresponding to the same data information.

[0009] Through the above technical solution, the timestamps, storage addresses, exponents, and mantissas of data information are classified, and corresponding differential calculations are performed according to their respective categories, so as to obtain the timestamp difference values, storage address difference values, exponent difference values, and mantissa difference values of data information other than the first data information, and obtain the differential calculation results of these other data.

[0010] In one embodiment, determining the compression unit value according to the sparse coefficient obtained from the differential calculation result includes: Count the proportion of zero values in the differential calculation result; According to the preset region division rule, obtain the sparse coefficient corresponding to the proportion of zero values; According to the sparse coefficient, obtain the compression unit value.

[0011] Through the above technical solution, by counting the proportion of zero values in the differential calculation result, the similarity between the data information to be counted is judged, and this similarity shows the sparse situation of the distribution of data information in the original data stream. The region division rule is preset by the staff, and this region division rule shows the correlation between the proportion of zero values and the sparse coefficient. Furthermore, according to this sparse coefficient, the compression unit value is obtained. Here, the compression unit value is used as the length of the subsequent sliding compression window, so as to ensure that there is a high similarity between the data information within the sliding compression window when performing a single compression of the sliding compression window according to the sparse situation, thereby improving the compression ratio.

[0012] In one embodiment, determining the compression unit value according to the sparse coefficient obtained from the differential calculation result includes: Extract k differential calculation results as a sample differential set, k ∈ [0, +∞]; count the number of zero values in the sample differential set to obtain the zero value quantity j; according to the ratio between j and k, obtain the proportion of zero values q; Find the corresponding sparse coefficient n from the preset sparse comparison table according to the proportion of zero values q; Take the sparse coefficient n as the power value of 2 to obtain the compression unit value N, L = 2 n ; Among them, the sparse coefficient n in the sparse comparison table increases as the proportion of zero values q increases.

[0013] Through the above solution, the region division rule is embodied as a region comparison table, which reflects the relationship between the proportion of zero values and the sparse coefficient through this region comparison table, and then obtains the compression unit value, thereby determining the length of the sliding compression window.

[0014] In one embodiment, the first encoding compression is respectively performed on the timestamps, storage addresses, exponents, and mantissas of the data information in the original data stream according to the sliding compression window corresponding to the compression unit value to obtain first compressed data, including: Obtain a first timestamp compression set of the data information according to the timestamp of the data information at the head in the original data stream and the timestamp difference values of other data information; use the compression unit value as the length of the sliding compression window, and perform first encoding compression on the first timestamp compression set to obtain a second timestamp compression set; Obtain a first storage address compression set of the data information according to the storage address of the data information at the head in the original data stream and the storage address difference values of other data information; use the compression unit value as the length of the sliding compression window, and perform first encoding compression on the first storage address compression set to obtain a second storage address compression set; Obtain a first exponent compression set of the data information according to the exponent of the data information at the head in the original data stream and the exponent difference values of other data information; use the compression unit value as the length of the sliding compression window, and perform first encoding compression on the first exponent compression set to obtain a second exponent compression set; Obtain a first mantissa compression set of the data information according to the mantissa of the data information at the head in the original data stream and the mantissa difference values of other data information; use the compression unit value as the length of the sliding compression window, and perform first encoding compression on the first mantissa compression set to obtain a second mantissa compression set; Sort out the second timestamp compression set, the second storage address compression set, the second exponent compression set, and the second mantissa compression set to obtain the first compressed data.

[0015] Through the above technical solution, differential compression is performed on the data information other than the head data information, and the sliding compression window of this case is applied to the differentially compressed original data stream for further compression to obtain the first compressed data regarding timestamps, storage addresses, decoding, and mantissas, improving the compression rate. In this example, two rounds of compression are performed.

[0016] In one embodiment, the first encoding compression is respectively performed on the timestamps, storage addresses, exponents, and mantissas of the data information in the original data stream according to the sliding compression window corresponding to the compression unit value to obtain first compressed data, including: Use the compression unit value as the length of the sliding compression window, and perform first encoding compression on the timestamps of the data information in the original data stream to obtain timestamp compression data; Taking the compressed unit value as the length of the sliding compression window, perform first encoding compression on the storage addresses of the data information in the original data stream to obtain storage address compressed data; Taking the compressed unit value as the length of the sliding compression window, perform first encoding compression on the exponent of the data information in the original data stream to obtain exponent compressed data; Taking the compressed unit value as the length of the sliding compression window, perform first encoding compression on the mantissa of the data information in the original data stream to obtain mantissa compressed data; Sort out the timestamp compressed data, storage address compressed data, exponent compressed data, and mantissa compressed data to obtain the first compressed data.

[0017] In this case, the original data stream is directly segmented and compressed by category through a sliding compression window. Since the similarity of the data information before and after within the segment is relatively high, the compression ratio of the obtained first compressed data is relatively increased. In this example, single-round compression is performed.

[0018] In one implementation, the first encoding compression is the LM77 encoding compression algorithm, and the second encoding compression is the run-length encoding compression algorithm.

[0019] In one implementation, the calculating the proportion of zero values in the differential calculation results includes: Traverse the differential calculation results and classify the differential calculation results into multiple sub-regions according to the clustering distribution principle; Calculate the proportion of zero values in the differential calculation results in each sub-region for calculating the compressed unit value corresponding to the sub-region, so as to perform first encoding compression on the timestamps, storage addresses, exponents, and mantissas of the data information in the sub-region according to the sliding compression window corresponding to the compressed unit value to obtain the compressed data corresponding to the sub-region; Wherein, the first compressed data includes the compressed data of each sub-region.

[0020] Through the above technical solution, the original data stream is divided into regions according to the clustering distribution principle. Subsequently, during the execution of step 103, samples are taken for each region and the compressed unit value is calculated, and then step 104 is executed for each region as a unit to obtain the first compressed data corresponding to each data information within the region.

[0021] According to the second aspect of the embodiments of the present application, there is provided a data segmentation compression module, including: A data receiving unit, configured to obtain the original data stream obtained by collecting the target signal; A compression preprocessing unit, connected to the data receiving unit, configured to perform differential calculations on the timestamps, storage addresses, exponents, and mantissas of the data information in the original data stream to obtain differential calculation results; A compression control unit, connected to the compression preprocessing unit, is configured to determine a compression unit value according to the sparse coefficient obtained from the differential calculation result and generate a compression start instruction; A first compression unit, connected to the compression control unit and the compression preprocessing unit, is configured to adjust to a working mode according to the compression start instruction, and in the working mode, perform first encoding compression on the timestamp, storage address, exponent, and mantissa of the data information in the original data stream respectively according to the sliding compression window corresponding to the compression unit value to obtain first compressed data; A second compression unit, connected to the compression control unit and the data receiving unit, is configured to adjust to a working mode according to the compression start instruction, and in the working mode, perform second encoding compression on the data type in the original data stream to obtain second compressed data; An arrangement unit, connected to the first compression unit and the second compression unit, is configured to obtain the compressed data of the data information according to the first compressed data and the second compressed data corresponding to the same data information in the original data stream; An output unit, connected to the arrangement unit, is configured to output the compressed data.

[0022] Through the above technical solution, the compression unit value is determined through the differential calculation result, and then the length of the sliding compression window is changed, so as to realize adjusting the sliding compression window according to the sparse situation of the data information distribution in the original data stream, and ensure that the sliding compression window is applicable to the original data stream in various density situations.

[0023] According to the third aspect of the embodiments of the present application, an electronic device is provided, including a memory and a processor, where the memory is used to store a computer program executable by the processor; the processor is used to execute the computer program in the memory to implement the above method. Description of the Drawings

[0024] Figure 1 is a flowchart of a data segmentation compression method shown according to an exemplary embodiment.

[0025] Figure 2 is Figure 1 a flowchart of step 102 in

[0026] Figure 3 is Figure 1 a flowchart of step 103 in

[0027] Figure 4 is Figure 1 a flowchart of a certain step 104 in

[0028] Figure 5 is Figure 1Another flowchart of step 104 in the present invention

[0029] Figure 6 It is a block diagram of a data segment compression module shown according to another exemplary embodiment.

[0030] Figure 7 It is a block diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners

[0031] Unless otherwise defined, the technical terms or scientific terms used in this specification and the claims should have the ordinary meanings understood by those of ordinary skill in the technical field to which the present invention belongs. The following will describe the detailed implementation manners of the present invention with reference to the accompanying drawings. It should be noted that in the process of the detailed description of these implementation manners, in order to make a concise description, this specification may not describe all the features of the actual implementation manners in detail. Without departing from the spirit and scope of the present invention, those skilled in the art can modify and replace the implementation manners of the present invention, and the obtained implementation manners are also within the protection scope of the present invention.

[0032] An embodiment of the present application provides a data segment compression method. This data segment compression method can be applied to a chip with multiple processors. Please refer to Figure 1 , the data segment compression method may include the following steps 101 to 106: Step 101, obtaining the original data stream obtained by collecting the target signal. Specifically, the target signal is the signal that needs to perform compressed data transmission during the operation of the chip, usually the signal collected during the operation of the chip. After being collected by the chip, these signals are converted into an original data stream. The original data stream includes at least one piece of data information. These data information include data type, time stamp, storage address, and data content (fixed-point data or floating-point data). When the data content is floating-point data, the data content includes an exponent and a mantissa.

[0033] This case mainly focuses on the situation where the data content of the data information in the original data stream is floating-point data. If the data content of the data information is fixed-point data, then in step 102, only the time stamp and storage address of the data information are differentially calculated to obtain the corresponding differential result. In step 103, the compression unit value is determined. And in the classification compression of step 104, the sliding compression window corresponding to the compression unit value is used to perform the first encoding compression on the time stamp and storage address of the data information to obtain the first compressed data. In step 105, the second encoding compression is performed on the data type to obtain the second compressed data. In step 106, the first compressed data, the second compressed data of the data information, and the fixed-point data as the data content are combined into the compressed data and output.

[0034] Step 102: Perform differential calculations on the timestamps, storage addresses, exponents, and mantissas of the data information in the original data stream respectively to obtain the corresponding differential calculation results.

[0035] Specifically, referring to Figure 2 as shown, the implementation of Step 102 includes: Step 2-1: Calculate the difference between the timestamp of the latter data information and the timestamp of the former data information in the original data stream to obtain the timestamp difference value corresponding to the latter data information.

[0036] Step 2-2: Calculate the difference between the storage address of the latter data information and the storage address of the former data information in the original data stream to obtain the storage address difference value corresponding to the latter data information.

[0037] Step 2-3: Calculate the difference between the exponent of the latter data information and the exponent of the former data information in the original data stream to obtain the exponent difference value corresponding to the latter data information.

[0038] Step 2-4: Calculate the difference between the mantissa of the latter data information and the mantissa of the former data information in the original data stream to obtain the mantissa difference value corresponding to the latter data information.

[0039] Step 2-5: Obtain the differential calculation result based on the timestamp difference value, storage address difference value, storage address difference value, exponent difference value, and mantissa difference value corresponding to the same data information.

[0040] In this case, the timestamps, storage addresses, exponents, and mantissas of the data information at the head of the original data stream are retained. After classifying other data information according to timestamps, storage addresses, exponents, and mantissas, corresponding differential calculations are performed respectively to obtain the timestamp difference value, storage address difference value, storage address difference value, exponent difference value, and mantissa difference value. The timestamp difference value, storage address difference value, storage address difference value, exponent difference value, and mantissa difference value corresponding to the same data information are used as the differential calculation result of this data information. At this time, the differential calculation result can be used to identify the continuity and repeatability between the front and back data in the original data stream, and further identify the high-order same part (i.e., redundant information) of adjacent data information in binary.

[0041] Step 103: Determine the compression unit value according to the sparsity coefficient obtained from the differential calculation result.

[0042] Specifically, referring to Figure 3 as shown, the execution of Step 103 includes: Step 3-1: Statistically calculate the proportion of zero values in the differential calculation result.

[0043] Step 3-2: Obtain the sparse coefficient corresponding to the zero-value ratio according to the preset region division rule.

[0044] Step 3-3: Obtain the compression unit value according to the sparse coefficient.

[0045] The zero values in the differential calculation results in Step 3-1 represent that the data information before and after is the same in terms of timestamp, storage address, exponent, or mantissa part, that is, there is redundancy. A high zero-value ratio means a high similarity of the data within the original data stream (or original data stream segment) counted in Step 3-1. The region division rule in Step 3-2 is preset by the staff. For example, when the zero-value ratio is in [0, 20%], the sparse coefficient is 1. Then, based on this sparse coefficient, the compression unit value is calculated. This compression unit value is used as the length of the subsequent sliding compression window, so as to ensure a high similarity between the data information within the sliding compression window when performing single compression of the sliding compression window according to the sparse situation, thereby improving the compression ratio.

[0046] In one example, Step 3-1 is as follows: Extract k differential calculation results as a sample differential set, where k ∈ [0, +∞]; count the number of zero values in the sample differential set to obtain the zero-value quantity j; obtain the zero-value ratio q according to the ratio between j and k. In this example, sampling inspection is performed during the statistics in Step 3-1.

[0047] Step 3-2 is as follows: Find the corresponding sparse coefficient n from the preset sparse look-up table according to the zero-value ratio q; the sparse coefficient n in the sparse look-up table increases as the zero-value ratio q increases. For example, when the zero-value ratio q in the sparse look-up table is in [0, 20%], the sparse coefficient n is 1; when the zero-value ratio q is in (20%, 40%], the sparse coefficient n is 2; when the zero-value ratio q is in (40%, 60%], the sparse coefficient n is 3; when the zero-value ratio q is in (60%, 80%], the sparse coefficient n is 4; when the zero-value ratio q is in (80%, 100%], the sparse coefficient n is 5; in Step 3-1, if the zero-value ratio q is 30%, then the corresponding sparse coefficient n is 2. Among them, the region division rule can be embodied as a region look-up table or can also be embodied as a calculation formula between the zero-value ratio and the sparse coefficient, etc.

[0048] Step 3-3 is as follows: Take the sparse coefficient n as the power value of 2 to obtain the compression unit value N, L = 2 n . For example, if the sparse coefficient n is 2, then the compression unit value N = 2 3 = 8. Then when this compression unit value is applied to the sliding compression window, the length of the sliding compression window is 8 bit.

[0049] In one example, Step 3-1, counting the zero-value ratio in the differential calculation results, includes: Traverse the differential calculation results and classify the differential calculation results into multiple sub-regions according to the clustering distribution principle; Statistically calculate the proportion of zero values in the differential calculation results for each sub-region, so as to calculate the compression unit value corresponding to the sub-region, and perform first-encoding compression on the timestamps, storage addresses, exponents, and mantissas of the data information in the sub-region according to the sliding compression window corresponding to the compression unit value to obtain the compression data corresponding to the sub-region; among them, the first compressed data includes the compressed data of each sub-region.

[0050] In this example, the original data stream is divided into regions according to the clustering distribution principle. Subsequently, during the execution of step 103, samples are taken for each region and the compression unit value is calculated, and then step 104 is executed for each region to obtain the first compressed data corresponding to each data information in the region. For example, the clustering distribution principle divides the original data stream into region A, region B, and region C. For region A, execute step 3-1 to obtain the proportion of zero values, execute step 3-2 to obtain the sparsity coefficient, execute step 3-3 to obtain the compression unit value, and execute step 104 to obtain the first compressed data of this region A; for region B, execute step 3-1 to obtain the proportion of zero values, execute step 3-2 to obtain the sparsity coefficient, execute step 3-3 to obtain the compression unit value, and execute step 104 to obtain the first compressed data of this region B; for region C, execute step 3-1 to obtain the proportion of zero values, execute step 3-2 to obtain the sparsity coefficient, execute step 3-3 to obtain the compression unit value, and execute step 104 to obtain the first compressed data of this region C; execute step 104, and when executing step 105, use the first compressed data and the second compressed data corresponding to the data information as part of the compressed data.

[0051] Step 104, according to the sliding compression window corresponding to the compression unit value, perform first-encoding compression on the timestamps, storage addresses, exponents, and mantissas of the data information in the original data stream to obtain the first compressed data.

[0052] In one example, the sliding compression window performs differential compression on other data information except the first data information (i.e., step 102), and then performs the compression of this step. Refer to Figure 4 As shown, step 104 includes: Step 4-1, according to the timestamp of the data information at the head of the original data stream and the timestamp difference value of other data information, obtain the first timestamp compression set of the data information; use the compression unit value as the length of the sliding compression window, and perform first-encoding compression on the first timestamp compression set to obtain the second timestamp compression set.

[0053] Step 4-2: Obtain the first storage address compression set of the data information based on the storage address of the data information at the head in the original data stream and the differential value of the storage addresses of other data information; use the compression unit value as the length of the sliding compression window, and perform first encoding compression on the first storage address compression set to obtain the second storage address compression set.

[0054] Step 4-3: Obtain the first exponent compression set of the data information based on the exponent of the data information at the head in the original data stream and the differential value of the exponents of other data information; use the compression unit value as the length of the sliding compression window, and perform first encoding compression on the first exponent compression set to obtain the second exponent compression set.

[0055] Step 4-4: Obtain the first mantissa compression set of the data information based on the mantissa of the data information at the head in the original data stream and the differential value of the mantissas of other data information; use the compression unit value as the length of the sliding compression window, and perform first encoding compression on the first mantissa compression set to obtain the second mantissa compression set.

[0056] Step 4-5: Organize the second timestamp compression set, the second storage address compression set, the second exponent compression set, and the second mantissa compression set to obtain the first compressed data.

[0057] In this example, first, differential compression is performed on the original data stream through step 102, and then the length of the sliding compression window (i.e., the compression unit value) is confirmed using step 103, so as to segment the original data stream by similarity. In the part with higher similarity, the length of the sliding compression window is larger, and more data information can be accommodated, realizing the aggregation of data information with higher similarity before and after for this round of compression, improving the compression rate. This example performs two rounds of compression.

[0058] In another example, the sliding compression window compresses the original data stream. As shown in Figure 5 Step 104 includes: Step 4-1': Use the compression unit value as the length of the sliding compression window, and perform first encoding compression on the timestamps of the data information in the original data stream to obtain timestamp compressed data.

[0059] Step 4-2': Use the compression unit value as the length of the sliding compression window, and perform first encoding compression on the storage addresses of the data information in the original data stream to obtain storage address compressed data.

[0060] Step 4-3': Use the compression unit value as the length of the sliding compression window, and perform first encoding compression on the exponents of the data information in the original data stream to obtain exponent compressed data.

[0061] Step 4-4’, using the compression unit value as the length of the sliding compression window, perform first encoding compression on the mantissa of the data information in the original data stream to obtain mantissa compressed data.

[0062] Step 4-5’, organize the timestamp compressed data, storage address compressed data, exponent compressed data, and mantissa compressed data to obtain first compressed data.

[0063] In this example, after using Step 103 to confirm the length of the sliding compression window (i.e., the compression unit value), the length of the sliding compression window in the area with high similarity of front and back data information in the original data stream is larger, and the data information it can accommodate is more, ensuring that the data information within the sliding compression window is the aggregation of data information with high front and back similarity. In this case, the original data stream is directly segmented and compressed by category through the sliding compression window. Since the similarity of front and back data information within the segment is relatively high, the compression ratio of the obtained first compressed data is relatively increased. This example performs single-round compression.

[0064] In one example, the first encoding compression in Step 104 uses the LM77 encoding compression algorithm.

[0065] Step 105, perform second encoding compression on the data types in the original data stream to obtain second compressed data.

[0066] Specifically, perform second encoding compression on the data types of the data information in the original data stream. The second encoding compression uses the run-length encoding compression algorithm.

[0067] Step 106, based on the first compressed data and the second compressed data corresponding to the same data information in the original data stream, obtain the compressed data of the data information and output the compressed data.

[0068] Specifically, when the data content of the data information in the original data stream is floating-point data, the first compressed data and the second compressed data corresponding to the data information are used as the compressed data of the data information. This compressed data is the compression result of the data information, and this compressed data is stored in the database or sent to the transmission end of the target data.

[0069] When the data content of the data information in the original data stream is fixed-point data, the first compressed data includes the compression results of the timestamp and storage address of the data information, the second compressed data includes the compression result of the data type, and the data should include the first compressed data, the second compressed data, and the data content of the data information.

[0070] In summary, the technical solution provided by this application has the following advantages: By determining the compression unit value based on the differential calculation result, and then changing the length of the sliding compression window, it is possible to adjust the sliding compression window according to the sparsity of the data information distribution in the original data stream, ensuring a higher similarity between the segmented data information within the sliding compression window, and improving the compression ratio and compression calculation efficiency of the compressed data obtained from the sliding compression window; This method of configuring the length of the sliding compression window is applicable to original data streams with various degrees of sparsity. And this solution is easy to implement through hardware, and the setting of the region division rules has strong flexibility.

[0071] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as they include the same logical relationship, they are within the protection scope of this patent; Adding insignificant modifications to the algorithm or process or introducing insignificant designs, but not changing the core design of its algorithm and process, are within the protection scope of this patent.

[0072] Another exemplary embodiment of this application also provides a data segmentation and compression module. As Figure 6 shown, in this embodiment, the data segmentation and compression module 60 includes: A data receiving unit 601, configured to obtain the original data stream obtained by collecting the target signal; A compression preprocessing unit 602, connected to the data receiving unit 601, and configured to perform differential calculations on the timestamps, storage addresses, exponents, and mantissas of the data information in the original data stream respectively to obtain differential calculation results; A compression control unit 603, connected to the compression preprocessing unit, and configured to determine the compression unit value according to the sparsity coefficient obtained from the differential calculation result and generate a compression start instruction; A first compression unit 604, connected to the compression control unit 603 and the compression preprocessing unit 602, and configured to adjust to the working mode according to the compression start instruction, and in the working mode, perform first encoding compression on the timestamps, storage addresses, exponents, and mantissas of the data information in the original data stream respectively according to the sliding compression window corresponding to the compression unit value to obtain first compressed data; A second compression unit 605, connected to the compression control unit 603 and the data receiving unit 601, and configured to adjust to the working mode according to the compression start instruction, and in the working mode, perform second encoding compression on the data types in the original data stream to obtain second compressed data; An organizing unit 606, connected to the first compression unit 604 and the second compression unit 605, and configured to obtain the compressed data of the data information according to the first compressed data and the second compressed data corresponding to the same data information in the original data stream; An output unit 607, connected to the sorting unit 606, is configured to output compressed data.

[0073] In this embodiment, the compression unit value is determined based on the differential calculation result, and then the length of the sliding compression window is changed, so as to adjust the sliding compression window according to the sparsity of the data information distribution in the original data stream, ensuring that the sliding compression window is applicable to the original data streams with various density situations.

[0074] It is not difficult to find that this embodiment is a structural embodiment corresponding to the first embodiment, and this embodiment can be implemented in cooperation with the first embodiment. The relevant technical details mentioned in the first embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied in the first embodiment.

[0075] It is worth mentioning that each module involved in this embodiment is a logical module. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, to highlight the innovative part of the present invention, units that are not closely related to solving the technical problems proposed by the present invention are not introduced in this embodiment, but this does not mean that there are no other units in this embodiment.

[0076] An embodiment of the present application also provides an electronic device, including a processor and a memory; the memory is configured to store a computer program executable by the processor; the processor is configured to execute the computer program in the memory to implement the data segmentation compression method described in any one of the above embodiments.

[0077] An embodiment of the present application also provides a computer-readable storage medium, when the executable computer program in the storage medium is executed by a processor, it can implement the image processing method described in any one of the above embodiments.

[0078] An embodiment of the present application also provides a computer program product, including a computer program, which when executed by a processor, implements the driving method of the universal serial bus interface described in any one of the above embodiments.

[0079] Regarding the device in the above embodiments, the specific manner in which the processor performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.

[0080] Figure 7 is a block diagram of an electronic device shown according to an exemplary embodiment. For example, the electronic device 700 may be provided as a server. Refer to Figure 7, Device 700 includes a processing component 722, which further includes one or more processors and memory resources represented by memory 932 for storing instructions executable by processing component 722, such as application programs. The application programs stored in memory 732 may include one or more modules each corresponding to a set of instructions. In addition, processing component 722 is configured to execute instructions to perform the above-described method for image processing.

[0081] Device 700 may also include a power component 726 configured to perform power management of device 700, a wired or wireless network interface 750 configured to connect device 700 to a network, and an input / output (I / O) interface 758. Device 700 may operate based on an operating system stored in memory 732, such as Windows ServerTM, MacOS XTM, UnixTM, LinuxTM, FreeBSDTM or the like.

[0082] In an exemplary embodiment, there is also provided a non-transitory computer-readable storage medium including instructions, such as memory 732 including instructions, the above instructions being executable by processing component 722 of device 700 to complete the above method. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0083] In the present invention, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance. The term "plurality" refers to two or more, unless otherwise clearly defined.

[0084] The above description of the embodiments is for the convenience of those of ordinary skill in the art to understand and apply the present application. It is obvious that those skilled in the art can easily make various modifications to these embodiments and apply the general principles described herein to other embodiments without creative effort. Therefore, the present application is not limited to the embodiments herein, and the improvements and modifications made by those skilled in the art without departing from the scope and spirit of the present application are within the scope of the present application.

Claims

1. A data segmentation and compression method, characterized in that, Including: Obtaining the original data stream obtained by collecting the target signal; Performing differential calculations on the timestamps, storage addresses, exponents, and mantissas of the data information in the original data stream respectively to obtain differential calculation results; Determining the compression unit value according to the sparse coefficient obtained from the differential calculation results; Performing first encoding compression on the timestamps, storage addresses, exponents, and mantissas of the data information in the original data stream respectively according to the sliding compression window corresponding to the compression unit value to obtain first compressed data; Performing second encoding compression on the data types in the original data stream to obtain second compressed data; Obtaining the compressed data of the data information according to the first compressed data and the second compressed data corresponding to the same data information in the original data stream, and outputting the compressed data.

2. The data segment compression method according to claim 1, wherein The performing differential calculations on the timestamps, storage addresses, exponents, and mantissas of the data information in the original data stream respectively to obtain differential calculation results includes: Calculating the difference between the timestamp of the latter data information and the timestamp of the former data information in the original data stream to obtain the timestamp difference value corresponding to the latter data information; Calculating the difference between the storage address of the latter data information and the storage address of the former data information in the original data stream to obtain the storage address difference value corresponding to the latter data information; Calculating the difference between the exponent of the latter data information and the exponent of the former data information in the original data stream to obtain the exponent difference value of the latter data information; Calculating the difference between the mantissa of the latter data information and the mantissa of the former data information in the original data stream to obtain the mantissa difference value of the latter data information; Wherein, the differential calculation results include the timestamp difference value, storage address difference value, storage address difference value, exponent difference value, and mantissa difference value corresponding to the same data information.

3. The data segment compression method according to claim 1, wherein The determining the compression unit value according to the sparse coefficient obtained from the differential calculation results includes: Counting the proportion of zero values in the differential calculation results; Obtaining the sparse coefficient corresponding to the proportion of zero values according to the preset region division rule; Obtaining the compression unit value according to the sparse coefficient.

4. The data segment compression method according to claim 3, wherein The determining the compression unit value according to the sparse coefficient obtained from the differential calculation results includes: Extracting k differential calculation results as a sample differential set, k ∈ [0, +∞]; counting the number of zero values in the sample differential set to obtain the zero value quantity j; obtaining the proportion of zero values q according to the ratio between j and k; Finding the corresponding sparse coefficient n from the preset sparse look-up table according to the proportion of zero values q; Take the sparse coefficient n as a power value of 2 to obtain a compression unit value N, where L = 2 n ; Wherein, the sparse coefficient n in the sparse look-up table increases as the proportion of zero values q increases.

5. The data segmentation compression method according to claim 2, wherein The performing first encoding compression on the timestamps, storage addresses, exponents, and mantissas of the data information in the original data stream respectively according to the sliding compression window corresponding to the compression unit value to obtain first compressed data includes: Obtain a first timestamp compression set of data information based on the timestamp of the data information at the head in the original data stream and the timestamp difference values of other data information; Use the compression unit value as the length of the sliding compression window, and perform first encoding compression on the first timestamp compression set to obtain a second timestamp compression set; Obtain a first storage address compression set of data information based on the storage address of the data information at the head in the original data stream and the storage address difference values of other data information; Use the compression unit value as the length of the sliding compression window, and perform first encoding compression on the first storage address compression set to obtain a second storage address compression set; Obtain a first exponent compression set of data information based on the exponent of the data information at the head in the original data stream and the exponent difference values of other data information; Use the compression unit value as the length of the sliding compression window, and perform first encoding compression on the first exponent compression set to obtain a second exponent compression set; Obtain a first mantissa compression set of data information based on the mantissa of the data information at the head in the original data stream and the mantissa difference values of other data information; Use the compression unit value as the length of the sliding compression window, and perform first encoding compression on the first mantissa compression set to obtain a second mantissa compression set; Sort out the second timestamp compression set, the second storage address compression set, the second exponent compression set, and the second mantissa compression set to obtain the first compressed data.

6. The data segment compression method according to claim 1, wherein The first encoding compression according to the sliding compression window corresponding to the compression unit value for the timestamps, storage addresses, exponents, and mantissas of the data information in the original data stream to obtain the first compressed data includes: Use the compression unit value as the length of the sliding compression window, and perform first encoding compression on the timestamps of the data information in the original data stream to obtain timestamp compression data; Use the compression unit value as the length of the sliding compression window, and perform first encoding compression on the storage addresses of the data information in the original data stream to obtain storage address compression data; Use the compression unit value as the length of the sliding compression window, and perform first encoding compression on the exponents of the data information in the original data stream to obtain exponent compression data; Use the compression unit value as the length of the sliding compression window, and perform first encoding compression on the mantissas of the data information in the original data stream to obtain mantissa compression data; Sort out the timestamp compression data, the storage address compression data, the exponent compression data, and the mantissa compression data to obtain the first compressed data.

7. The data segment compression method according to claim 5 or 6, characterized in that The first encoding compression is the LM77 encoding compression algorithm, and the second encoding compression is the run-length encoding compression algorithm.

8. The data segment compression method according to claim 3, characterized in that, The counting of the proportion of zero values in the difference calculation results includes: Traverse the difference calculation results, and classify the difference calculation results into multiple sub-regions according to the clustering distribution principle; Count the proportion of zero values in the differential calculation results in each of the sub-regions for calculating the compression unit value corresponding to the sub-region, and perform first encoding compression on the timestamps, storage addresses, exponents, and mantissas of the data information in the sub-region according to the sliding compression window corresponding to the compression unit value to obtain the compressed data corresponding to the sub-region; Among them, the first compressed data includes the compressed data of each of the sub-regions.

9. A data segmentation and compression module, characterized in that, It includes: A data receiving unit for obtaining the original data stream obtained by collecting the target signal; A compression preprocessing unit connected to the data receiving unit for performing differential calculations on the timestamps, storage addresses, exponents, and mantissas of the data information in the original data stream respectively to obtain differential calculation results; A compression control unit connected to the compression preprocessing unit for determining the compression unit value according to the sparse coefficient obtained from the differential calculation results and generating a compression start instruction; A first compression unit connected to the compression control unit and the compression preprocessing unit for adjusting to the working mode according to the compression start instruction, and performing first encoding compression on the timestamps, storage addresses, exponents, and mantissas of the data information in the original data stream respectively according to the sliding compression window corresponding to the compression unit value in the working mode to obtain the first compressed data; A second compression unit connected to the compression control unit and the data receiving unit for adjusting to the working mode according to the compression start instruction, and performing second encoding compression on the data types in the original data stream in the working mode to obtain the second compressed data; An organizing unit connected to the first compression unit and the second compression unit for obtaining the compressed data of the data information according to the first compressed data and the second compressed data corresponding to the same data information in the original data stream; An output unit connected to the organizing unit for outputting the compressed data.

10. An electronic device, characterized in that, It includes a memory and a processor, the memory is used for storing the computer program executable by the processor; the processor is used for executing the computer program in the memory to implement the method according to any one of claims 1 to 8.