Data compression transmission method, device and equipment and readable storage medium

By constructing dynamic character sets and multiple compression algorithms, efficient compression and transmission are carried out on the data characteristics of the Beidou satellite system, solving the problems of large data volume and high frequency of change, and improving data transmission efficiency and communication efficiency.

CN120416347AActive Publication Date: 2025-08-01CRRC INDUSTRAIL ACADEMY (QINGDAO) CO LTD

Patent Information

Application Number
CN202510617746.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-01
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

In the prior art, data compression methods with low correlation, high frequency of change and large amount of data are inefficient, and the limited communication resources of the Beidou satellite system cannot be effectively utilized.

Method used

A dynamic character set is constructed based on the characteristics of the data to be transmitted. By analyzing the frequency distribution and time series characteristics of the fault data, combining time information and wind field numbering, a variety of character sets are constructed, and algorithms such as pre-filled dictionary, probability interval coding, head compression mechanism and prime base quantization are used for compression, and finally transmitted according to the preset single packet data utilization rate.

Benefits of technology

It improves the reliability and transmission efficiency of data compression, ensures single packet data utilization, optimizes the efficiency of Beidou short message communication, especially in the data transmission of wind turbines, and significantly improves the efficiency and accuracy of fault diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120416347A_ABST
    Figure CN120416347A_ABST
Patent Text Reader

Abstract

The invention discloses a data compression transmission method, device and equipment and a readable storage medium, and is applied to the field of data compression transmission, and the data compression transmission method comprises the following steps: respectively constructing corresponding character sets according to data characteristics of to-be-transmitted data; performing data evaluation on the character set, determining a corresponding compression algorithm according to an evaluation result, and performing compression to obtain compressed data; and transmitting the compressed data according to a preset single packet data utilization rate. According to the invention, for Beidou short message communication limitation, a character set is constructed for data characteristics, an adaptive compression algorithm is automatically selected based on a data evaluation result, and the compressed data is transmitted according to a preset single packet data utilization rate. Compared with a traditional fixed character set, the data frequency is improved; the data transmission efficiency is greatly improved while the data compression reliability is ensured; and the utilization rate of single-packet data can be ensured, and the communication efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data compression and transmission, and in particular to a data compression and transmission method, device, equipment and computer-readable storage medium. Background Art

[0002] The short message communication resources of the Beidou satellite system are limited. The communication bandwidth and processing capacity of the satellite are fixed. When a large number of users request short message communication services simultaneously, the situation of resource tension will occur. Therefore, generally, data needs to be compressed, such as compressing all data to be transmitted using binary coding. This compression method is applicable to data with small changes in a short period of time; or compressing PVT (position, velocity, and time) information using a compression protocol and the corresponding transmission protocol. This method is applicable to the transmission of data with continuous correlation, etc. For data with less or no correlation, or data with large changes in a short period of time, the existing compression methods cannot achieve good compression effects.

[0003] Therefore, how to provide an efficient compression method for data with low correlation, high change frequency, and large data volume is a technical problem that needs to be solved urgently at present. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a data transmission method, device, equipment and computer-readable storage medium, which solves the problem that there is no efficient compression and data transmission method for data with low correlation, high change frequency, and large data volume in the prior art.

[0005] To solve the above technical problem, the present invention provides a data compression and transmission method, including: respectively constructing corresponding character sets according to the data characteristics of the data to be transmitted; performing data evaluation on the character sets, determining corresponding compression algorithms according to the evaluation results and performing compression to obtain compressed data; and transmitting the compressed data according to a preset single-pack data utilization rate.

[0006] Optionally, respectively constructing corresponding character sets according to the data characteristics of the data to be transmitted includes: segmenting and optimizing the fault data by analyzing the frequency distribution and time series characteristics of the fault data to construct a first character set; constructing a second character set based on time information; the time information includes the start time, end time, and current system time of the fault; and constructing a third character set based on the wind farm number, wind turbine number, and whether it is the first-occurring fault.

[0007] Optionally, data evaluation is performed on the character set, and a corresponding compression algorithm is determined according to the evaluation result and compression is performed to obtain compressed data, including: calculating the repeated substring length, repeated density, and data entropy of the character set; using a pre-filled dictionary and probability interval coding to compress a fourth character set whose repeated substring length is greater than a first threshold, repeated density is greater than a second threshold, and data entropy is not greater than a third threshold to obtain first compressed data; using a header compression mechanism and prime base quantization to compress a fifth character set whose repeated substring length is not greater than the first threshold, repeated density is not greater than the second threshold, and data entropy is greater than the third threshold to obtain second compressed data; using the Deflate compression method to compress a sixth character set with a uniform character frequency distribution to obtain third compressed data.

[0008] Optionally, using a pre-filled dictionary and probability interval coding to compress a fourth character set whose repeated substring length is greater than a first threshold, repeated density is greater than a second threshold, and data entropy is not greater than a third threshold to obtain first compressed data, including: constructing the pre-filled dictionary based on historical data; for the characters in the fourth character set that are in the pre-filled dictionary, output a dictionary match flag bit and an index; for the characters in the fourth character set that are not in the pre-filled dictionary, use the LZ77 algorithm for matching. If the match is successful, output a sliding window flag bit, a distance, and a length; if the match fails, output an original value flag bit and an original value; use the dictionary match flag bit, the sliding window flag bit, and the original value flag bit as first data to be encoded; use the distance, the length, and the original value as second data to be encoded, third data to be encoded, and fourth data to be encoded respectively; perform probability interval coding on the first data to be encoded, the second data to be encoded, the third data to be encoded, and the fourth data to be encoded respectively to obtain the first compressed data.

[0009] Optionally, performing probability interval coding on the first data to be encoded, the second data to be encoded, the third data to be encoded, and the fourth data to be encoded respectively to obtain the first compressed data, including: using the first data to be encoded, the second data to be encoded, the third data to be encoded, and the fourth data to be encoded as input data respectively; defining symbols for the input data and calculating the frequencies of the symbols, setting initial symbol intervals according to the frequencies of the symbols; updating the symbol intervals based on the initial symbol intervals and the symbols until the final symbol intervals are obtained; using the shortest binary number within the final symbol intervals as the first compressed data.

[0010] Optionally, the fifth character set with a repeated substring length not greater than the first threshold, a repeated density not greater than the second threshold, and a data entropy greater than the third threshold is compressed using a header compression mechanism and prime base quantization to obtain second compressed data, including: statistically analyzing the symbol frequencies of the fifth character set to generate a frequency dictionary and a frequency sequence; the frequency dictionary includes a symbol sequence and a quantized frequency difference sequence; calculating the total frequency based on the frequencies of all symbols in the fifth character set, and using a prime number greater than or equal to the total frequency as the quantization base; quantizing the frequency sequence using the quantization base to obtain a quantized frequency sequence, constructing a state space based on the quantized frequency sequence, generating a coding table based on the state space, and converting the fifth character set into a compressed bit stream using the coding table; using dynamic Bit-Packing and fixed Bit-Packing methods to compress the header information to obtain a header compression result; the header information includes the frequency dictionary and the quantization base.

[0011] Optionally, the compressed data is transmitted according to a preset single-packet data utilization rate, including: if the data volume of the compressed data is less than the minimum value of the target range, calculating a first difference according to the data volume of the compressed data and the minimum value, and increasing the data volume of the data to be transmitted according to the first difference; the target range is determined according to the preset single-packet data utilization rate; if the data volume of the compressed data is greater than the maximum value of the target range, transmitting a part of the compressed data, and the data volume of the part of the data is the maximum value of the target range; and calculating a second difference according to the data volume of the compressed data and the maximum value, and reducing the data volume of the data to be transmitted according to the second difference; if the data volume of the compressed data is within the target range, transmitting the compressed data.

[0012] The present invention also provides a data compression and transmission device, including: a character set construction module, configured to respectively construct corresponding character sets according to the data characteristics of the data to be transmitted; a compression module, configured to perform data evaluation on the character sets, determine corresponding compression algorithms according to the evaluation results, and perform compression to obtain compressed data; and a transmission module, configured to transmit the compressed data according to a preset single-packet data utilization rate.

[0013] The present invention also provides a data compression and transmission device, including: a memory, configured to store a computer program; and a processor, configured to implement the steps of the data compression and transmission method as described above when executing the computer program.

[0014] The present invention also provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are loaded and executed by a processor, the steps of the data compression and transmission method as described above are implemented.

[0015] It can be seen that the present invention constructs corresponding character sets according to the data characteristics of the data to be transmitted; performs data evaluation on the character sets, determines the corresponding compression algorithm based on the evaluation results and performs compression to obtain compressed data; and transmits the compressed data according to the preset single-packet data utilization rate. The present invention addresses the communication limitations of Beidou short messages, constructs character sets based on data characteristics, and automatically selects an adaptive compression algorithm based on the data evaluation results, and transmits the compressed data according to the preset single-packet data utilization rate. Compared with traditional fixed character sets, the data frequency is increased; and the data transmission efficiency is greatly improved while ensuring the reliability of data compression; it can also ensure the single-packet data utilization rate, thereby improving communication efficiency.

[0016] In addition, the present invention also provides a data compression transmission device, equipment and readable storage medium, which also have the above-mentioned beneficial effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0018] Figure 1 A flowchart of a data compression transmission method provided by an embodiment of the present invention.

[0019] Figure 2 This is an example diagram of a dynamic priority screening mechanism provided by an embodiment of the present invention.

[0020] Figure 3 This is a flowchart illustrating an example of a compression algorithm provided by an embodiment of the present invention.

[0021] Figure 4 This is a flowchart illustrating another compression algorithm provided by an embodiment of the present invention.

[0022] Figure 5 This is an example diagram of an adaptive data screening method provided by an embodiment of the present invention.

[0023] Figure 6 This is a flowchart illustrating an overall data compression method provided by an embodiment of the present invention.

[0024] Figure 7 A structural diagram of a data compression transmission device provided by an embodiment of the present invention.

[0025] Figure 8Schematic diagram of a data compression and transmission device provided by an embodiment of the present invention. Detailed implementation manners

[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0027] Please refer to Figure 1 , Figure 1 Flowchart of a data compression and transmission method provided by an embodiment of the present invention. The method may include step S101-step S103.

[0028] S101: Construct corresponding character sets respectively according to the data characteristics of the data to be transmitted.

[0029] The execution subject of this embodiment is a terminal. The type of the terminal is not limited in this embodiment, as long as it can complete the operations of the data compression and transmission method. The data to be transmitted in this embodiment is data with a large amount, low correlation, and high change frequency. Exemplarily, the data to be transmitted may be wind power data, and the wind turbine data may include: fault data, time data, and other data. The construction of the character set is not limited in this embodiment. It can be understood that the corresponding character set can be constructed according to the characteristics of the real-time data to be transmitted. Therefore, the character set can be a dynamic character set.

[0030] Further, the above-mentioned constructing corresponding character sets respectively according to the data characteristics of the data to be transmitted may specifically include step 11-step 13.

[0031] Step 11: Segment and optimize the fault data by analyzing the frequency distribution and time series characteristics of the fault data, and construct a first character set.

[0032] Specifically, the fault data is usually a long field. Based on the fault data, by analyzing the data frequency distribution and time series characteristics, a frequency adjustment scheme is formulated, high-frequency and low-frequency patterns are identified, and segmentation and optimization are performed to construct the corresponding character set. Exemplarily, the fault data within a period of time is called for counting the data occurrence frequency, analyzing the data frequency, and formulating a suitable frequency adjustment method. Analyze the frequency distribution and time series characteristics of the data, and identify high-frequency and low-frequency patterns. Then the fault data is segmented and decomposed into several independent fields, and each field contains different values. Since some fields in the fault data appear frequently and some fields are always 0, according to the characteristics of the fault code.

[0033] Step 12: Construct a second character set based on time information; the time information includes the start time of the fault, the end time of the fault, and the current system time.

[0034] Specifically, based on the time information, including the fault time (start time and end time of the fault) and the current system time, construct the corresponding character set. Among them, the time information can be preprocessed. The specific process can include: (1) Convert the time to a UNIX timestamp. Convert all fault times (such as the start time and end time of the fault) to UNIX timestamps, which can ensure that all times are unified into a standard time representation, facilitating subsequent differential calculations. However, 32-bit UNIX can only represent the time from 00:00:00 UTC on January 1, 1970 to 03:14:07 UTC on January 19, 2038. To extend the representable time range and reduce the converted UNIX timestamp value, calculate the difference between the converted UNIX timestamp and 00:00:00 on January 1, 2023, so that the UNIX timestamp value of 00:00:00 on January 1, 2023 is represented by 0, and the improved available time limit is extended to 03:14:07 UTC on January 19, 2091. (2) Differential encoding. Take the earliest occurrence time of the first-occurring fault as the reference time, and use the differential encoding method to calculate the time difference for the end time of the first-occurring fault and the start and end times of other faults, and convert them into binary time codes for storage. For times with a large difference from the reference time, special bytes are used for transmission. For example: The occurrence time of fault 1 is 12:00:00 on December 21, 2024, which is 62563200, and the occurrence time of fault 2 is 12:00:40 on December 21, 2024, which is 62563240. Taking fault 1 as the reference time, fault 2 is represented as +40, which can be represented by the byte 00101000. When data is backlogged and the time difference between faults is large, dynamic byte encoding is used to store the differential encoding value. A mode identification code is introduced for processing. If the mode identification code is 1, it represents the normal data compression mode, and if it is 0, it represents the data backlog mode. At the same time, based on different data differences, dynamic bytes are introduced for joint encoding, as shown in Table 1.

[0035] Table 1 Joint Encoding in Data Backlog Mode

[0036]

[0037] (3) Time update: Each time data is transmitted, the reference time is updated according to the latest occurrence time of the current fault event. Suppose that after a certain fault data transmission, the latest fault start time is 12:03:20 on December 21, 2024, which is 1703155400. Then the next reference time will be updated to 1703155400, and the time difference between adjacent events will continue to be calculated.

[0038] It should be noted that the above first character set, second character set, and third character set are different character sets obtained according to different data characteristics.

[0039] Step 13: Construct a third character set based on the wind farm number, the wind turbine number, and whether it is a first-occurrence fault.

[0040] Specifically, a byte stream is constructed in units of bits. Since the number of wind turbine numbers is no more than 127 and the data of whether it is a first-occurrence fault is represented by Boolean characters, the 1-bit binary data of whether it is a first-occurrence fault and the 7-bit binary wind turbine number are combined according to a certain rule to form an 8-bit binary encoding string.

[0041] Furthermore, before constructing the corresponding character sets according to the data characteristics of the data to be transmitted, the following steps may further be included: Step 1: Assign weights to each data in the database according to evaluation indicators, and calculate the total weighted score of the faults of each data by using the weighted method; the evaluation indicators include at least one of the fault type, fault time, fault level, and first-occurrence mark; Step 2: Determine the transmission priority of each data according to the total weighted score of the faults, and determine the transmission order of each data according to the transmission priority; Step 3: Determine the preset initial amount size according to the historical compression ratio; Step 4: Determine the data to be transmitted according to the transmission order and the preset initial amount size.

[0042] Exemplarily, reference can be made to Figure 2 , Figure 2 , which is an example diagram of a dynamic priority screening mechanism provided by an embodiment of the present invention. For the wind turbine data in the wind turbine database, the weighted method is used to assign index weights to the fault type, the time of fault occurrence, the fault level, etc., calculate the total weighted score of the faults and sort them. The higher the total weighted score, the higher the priority. The data is selected for transmission in order of priority. The size of the selected data volume is determined according to the historical compression ratio. Specifically, the system makes a preliminary estimate according to the designed efficient compression method, preset the initial data volume to be compressed, and the preset initial amount can be estimated according to the historical compression ratio to determine the initial data volume.

[0043] S102: Perform data evaluation on the character set, determine the corresponding compression algorithm according to the evaluation result, and perform compression to obtain compressed data.

[0044] In this embodiment, data evaluation is performed on each of the above character sets. The data evaluation criteria in this embodiment are not specifically limited. For example, they can be the character frequency distribution and the repeatability of substrings, etc. Select the corresponding compression algorithm according to the data evaluation structure. For example, for evenly distributed characters or high-frequency characters, the Deflate compression method is used; for unevenly distributed characters, other efficient and adaptable compression methods are used.

[0045] Further, the above data evaluation of the character set is performed, and the corresponding compression algorithm is determined according to the evaluation result and compression is performed to obtain compressed data, which may specifically include steps 21-step 24.

[0046] Step 21: Calculate the length of repeated substrings, repetition density, and data entropy of the character set.

[0047] Specifically, assume that the data sequence is , and the frequency of each symbol is represented as follows: .

[0048] Among them, represents the number of times the symbol appears, represents the frequency of the symbol appears, that is, the character frequency distribution; n represents the total number of symbols.

[0049] Define the substring as the substring from position to . For any substring , define the number of times it appears as , the substring length is , the total length of all repeated substrings and the repetition density are: .

[0050] Data entropy is an important indicator to measure data randomness. Based on the character frequency distribution and the Shannon entropy formula, the data entropy can be expressed as follows: .

[0051] If the frequencies of some symbols are significantly higher than those of other symbols, it indicates that the data has high redundancy and is suitable for entropy coding; if the data entropy is low, it indicates that there is a lot of redundant information in the data and is suitable for entropy coding. If the data entropy is close to the maximum value, it indicates that the data is close to a uniform distribution and the compression effect may be poor; if the repetition density is high, it indicates that there are a large number of repeated patterns in the data and is suitable for dictionary coding. If the repetition density is low, it indicates that there are few repeated patterns in the data and other compression methods may be required.

[0052] Step 22: Compress the fourth character set with a repeated substring length greater than the first threshold, a repetition density greater than the second threshold, and a data entropy not greater than the third threshold by using a pre-filled dictionary and probability interval coding to obtain the first compressed data.

[0053] It should be noted that steps 22 and 23 in this embodiment are directed to non-uniformly distributed data. This embodiment does not limit the first threshold, the second threshold, and the third threshold. Specifically, for a large number of repeated substrings (with a relatively long length of the repeated substring), medium and low entropy , and high repetition density data, pre-filled dictionaries and probability interval coding are used for compression. To address the problem of the inefficiency of the initial stage of the LZ77 (Lempel-Ziv 1977, a lossless compression algorithm) algorithm, a pre-filled dictionary is designed to capture high-frequency patterns using prior knowledge, reduce the output of redundant characters in the early stage, improve the compression efficiency in the initial stage, avoid inefficiency when the field is empty, and accelerate compression convergence. Enhance the ability to capture short repetition patterns, improve the compression efficiency for locally repeated data, reduce the output of redundant characters in the early stage, and increase the compression ratio. At the same time, to address the problem of residual symbol redundancy in the output of the LZ77 algorithm, a coding strategy (i.e., probability interval coding) is added to eliminate statistical redundancy.

[0054] Furthermore, the above-mentioned compression of the fourth character set with a repeated substring length greater than the first threshold, a repetition density greater than the second threshold, and a data entropy not greater than the third threshold using a pre-filled dictionary and probability interval coding to obtain the first compressed data specifically may include the following steps: Step 221: Construct a pre-filled dictionary based on historical data; Step 222: For the characters in the fourth character set that are in the pre-filled dictionary, output the dictionary match flag bit and the index; Step 223: For the characters in the fourth character set that are not in the pre-filled dictionary, perform matching using the LZ77 algorithm. If the match is successful, output the sliding window flag bit, the distance, and the length; if the match fails, output the raw value flag bit and the raw value; Step 224: Use the dictionary match flag bit, the sliding window flag bit, and the raw value flag bit as the first data to be encoded; Step 225: Use the distance, the length, and the raw value as the second data to be encoded, the third data to be encoded, and the fourth data to be encoded respectively; Step 226: Perform probability interval coding on the first data to be encoded, the second data to be encoded, the third data to be encoded, and the fourth data to be encoded respectively to obtain the first compressed data.

[0055] Specifically, reference can be made to Figure 3 , Figure 3 which is a flowchart example of a compression algorithm provided by an embodiment of the present invention. Figure 3 The data input in Figure 3The [flag bit 1] [index] in it, and at the same time add the processed symbols to the sliding window. (3) LZ77 sliding window compression: If the pre-filled dictionary matching fails, search for the longest repeated field of the symbol in the sliding window. If the match is successful in the sliding window, output [sliding window flag bit] [distance] [length], that is Figure 3 The [flag bit 2] [distance] [length] in it. If the match fails, output [raw value flag] [raw value], that is Figure 3 The [flag bit 3] [distance] [raw value] in it. When the LZ77 sliding window matches the data stream (referring to the fourth character set), add the processed symbols to the sliding window. When the window size is exceeded, the earliest symbol is eliminated. Further, the pre-filled dictionary can also be updated. The specific conditions are as follows: Statistically update the pre-filled dictionary based on the frequency of the current value. If it exceeds the fourth threshold, add it to the pre-filled dictionary. (4) Probability interval coding. Construct the above [dictionary matching flag bit], [sliding window flag bit], and [raw value flag] [raw value] into unified data to be encoded. At the same time, construct [index], [distance], and [length] into unified data to be encoded respectively. Perform probability interval coding on the above four groups of data to be encoded respectively, and output four groups of compressed data.

[0056] Further, the above-mentioned probability interval coding is performed on the first data to be encoded, the second data to be encoded, the third data to be encoded, and the fourth data to be encoded respectively to obtain the first compressed data. Specifically, it can include the following steps: Step 2261: Respectively use the first data to be encoded, the second data to be encoded, the third data to be encoded, and the fourth data to be encoded as input data; Step 2262: Define symbols for the input data and calculate the frequencies of each symbol, and set the initial symbol intervals according to the frequencies of each symbol; Step 2263: Update the symbol intervals based on the initial symbol intervals and each symbol until the final symbol intervals are obtained; Step 2264: Use the shortest binary number within the final symbol intervals as the first compressed data.

[0057] Specifically, reference can be made to Figure 3 The probability interval coding part in it. (1) Define symbols: Define multiple symbol types such as [dictionary matching flag bit] (i.e., A), [raw value flag] [raw value] (i.e., B + raw value), and [sliding window flag bit]. (2) Set the initial symbol frequencies and initial symbol intervals: Obtain the frequencies of each symbol as the initial frequencies, and map the symbols to the interval [0, 1) to divide the initial intervals. (3) Read the current field and encode it according to the current interval. (4) Read the data stream (referring to the input data): For each input symbol, further divide the sub-intervals within the current interval (assumed to be [L, H]) according to its probability, and update the current interval to the sub-interval corresponding to the symbol. The updated sub-interval rules are as follows: .

[0058] wherein, and are respectively the probabilities of symbols and ; represents the new lowest, represents the new highest; represents the previous highest, represents the previous lowest. (5) Repeat the above process until all symbols are processed. Finally, select a shortest binary fraction within the final coding interval as the coding result.

[0059] To better understand steps 2261 - 2264, the following example can be referred to: Divide the total interval [0, 1) according to the symbol occurrence probabilities. For example: The probability of symbol A is 50%, corresponding to the interval [0, 0.5); the probability of symbol B is 30%, corresponding to the interval [0.5, 0.8); the probability of symbol C is 20%, corresponding to the interval [0.8, 1). At the beginning, the coding range is the entire interval: lowest point = 0, highest point = 1. Coding "BAC": The first symbol B: B is in the [0.5, 0.8) segment of the original interval; new interval: lowest point = 0.5, highest point = 0.8. The second symbol A: Locate according to the proportion of A in the current interval [0.5, 0.8) (original 0 - 50%): new lowest = 0.5 + (0.8 - 0.5) × 0 = 0.5, new highest = 0.5 + (0.8 - 0.5) × 0.5 = 0.65; at this time, the interval is reduced to 0.5 - 0.65. The third symbol C: Locate according to the proportion of C in the current interval (0.5 - 0.65) (original 80% - 100%): new lowest = 0.5 + (0.65 - 0.5) × 0.8 = 0.5 + 0.12 = 0.62, new highest = 0.5 + (0.65 - 0.5) × 1 = 0.65; final interval 0.62 - 0.65. Output the shortest code: Find a shortest binary fraction within the final range (0.62 - 0.65), for example: 0.625 (binary 0.101) is just within the interval, and the final coding result: 101.

[0060] Step 23: Compress the fifth character set with repeated substring lengths not greater than the first threshold, repeated densities not greater than the second threshold, and data entropy greater than the third threshold by using the header compression mechanism and prime base quantization to obtain the second compressed data.

[0061] For fewer repeated substrings (i.e., shorter repeated substring lengths), high data entropy , low repeated density For the data, a compression method using a header compression mechanism and prime base quantization is adopted. For the convenience of binary operations, the traditional ANS (Asymmetric Numeral Systems) algorithm quantizes symbols into powers of 2 (such as 1 / 2, 1 / 4, 1 / 8). However, for small data, the actual probability distribution may deviate significantly from powers of 2, resulting in significant quantization errors. At the same time, its state partition is based on uniform probability, and multiple symbols may be mapped to the same state interval, resulting in the need to handle conflicts additionally during encoding, thus reducing the compression efficiency. Therefore, an improved ANS algorithm based on small data is designed, which uses prime base quantization to divide the probability interval more finely and flexibly, making the probability approach the actual frequency, reducing quantization errors, avoiding state overlap, improving the compression efficiency, and fully exploiting the potential of ANS. For small data, the symbol table stored by the ANS compression algorithm will occupy a large amount of memory. Therefore, a header compression method is designed to reduce the data overhead.

[0062] Furthermore, the above-mentioned fifth character set with a repeated substring length not greater than the first threshold, a repeated density not greater than the second threshold, and a data entropy greater than the third threshold is compressed using the header compression mechanism and prime base quantization to obtain the second compressed data, which may specifically include the following steps: Step 231: Statistically analyze the symbol frequencies of the fifth character set to generate a frequency dictionary and a frequency sequence; the frequency dictionary includes a symbol sequence and a quantized frequency difference sequence; Step 232: Calculate the total frequency based on the frequencies of all symbols in the fifth character set, and use a prime number greater than or equal to the total frequency as the quantization base; Step 233: Quantize the frequency sequence using the quantization base to obtain a quantized frequency sequence, construct a state space based on the quantized frequency sequence, generate a coding table based on the state space, and convert the fifth character set into a compressed bit stream using the coding table; Step 234: Compress the header information using the dynamic Bit-Packing and fixed Bit-Packing methods to obtain the header compression result; the header information includes the frequency dictionary and the quantization base.

[0063] Specifically, reference can be made to Figure 4 , Figure 4 which is a flowchart example of another compression algorithm provided by an embodiment of the present invention. Figure 4 The left side is the mainstream data compression process, which includes the construction of a frequency dictionary, the selection of a quantization base, the generation of a state space and a coding table, and the compression of data in the fifth character set; Figure 4 The right side is the header information compression process. (1) Generate a frequency dictionary and a frequency sequence: Statistically analyze the symbol frequencies of the input data stream (referring to the fifth character set) to generate a frequency dictionary, sort the symbols in the frequency dictionary from low to high according to the frequencies, and convert them into a difference sequence. The conversion formula is as follows: . (2) Quantization base selection: A predefined list of prime candidates is used for quantifying probabilities. Calculate the actual frequency and the actual probability of all symbols, find the total frequency of all symbols, and select a prime number greater than or equal to the total frequency as the quantization base . Quantify the actual probability and actual frequency based on the quantization base. The quantization formula is as follows: .

[0064] Where, is the character frequency; is the quantized character frequency; is the character probability; is the quantized character probability.

[0065] (3) Data encoding: Define the global state space of the state table and the frequency subspace of each symbol as follows: ; Define the initial state as (agreed rule with the decompression end), l is the scaling factor. When reading the current character, to ensure that the current state is always within the state subspace of the character, the current state needs to be scaled and the bitstream is output. The output bitstream and the scaled state are calculated as follows: .

[0066] Calculate the state and the corresponding number of output bits and bit bitstream , and complete the corresponding state transition table (i.e., the encoding table). The rule for updating the new state based on the previous state is as follows: .

[0067] Where, represents the cumulative frequency .

[0068] During encoding, find the corresponding new state and output bitstream in the encoding table according to the current state and the input symbol. Finally, output a string of bitstreams as the main body of data compression.

[0069] (4) Header compression: The header information mainly includes the frequency dictionary and the quantization base . The frequency dictionary includes the symbol sequence and the sequence of quantized frequency differences. The quantization base is represented by a fixed number of bytes. Compress the symbol sequence based on the dynamic Bit-Packing (a technique for reducing storage space by compressing data bits) method, and calculate the number of bytes required for storing each symbol in the symbol sequence , if the number of bytes required for storing all symbols is the same, i.e., , then identification code 1 is used; if the number of bytes required for storing all symbols is different, then identification code 0 is used, and the number of bytes required for storing each symbol, i.e., the bit stream , by default, each symbol can be represented by up to 4 bytes and is represented by 2 bits , and the final storage length is bits of the bit stream. Based on the fixed Bit-Packing method, the frequency difference sequence is compressed and stored with a fixed number of bits. Calculate the maximum number of bits required for storage in the frequency difference sequence , and all elements in the frequency difference sequence are stored with the maximum number of bits , and finally a bit stream with a length of is obtained.

[0070] Step 24: Compress the sixth character set with uniform character frequency distribution by using the Deflate compression method to obtain the third compressed data.

[0071] Both Step 22 and Step 23 are for data with non-uniform character frequency distribution. For data with uniform character frequency distribution or high-frequency characters, the Deflate (DEFLATE Compression Method) compression method can be used. Deflate is a widely used data compression algorithm and will not be elaborated here. It should be noted that the above fourth, fifth, and sixth character sets are different character sets obtained after data evaluation of the first, second, and third character sets, and corresponding efficient compression algorithms are adopted for the fourth and fifth character sets respectively. Among them, whether the character frequency is uniform can be determined according to the entropy ratio, and the specific entropy ratio formula is: , ; n is the size of the character set.

[0072] S103: Transmit the compressed data according to the preset single-packet data utilization rate.

[0073] Since the communication capacity of the Beidou short message is N, to improve the communication efficiency, based on the adaptive screening adjustment and compression cache mechanism, ensure that the single-packet data utilization rate reaches more than M (such as 96%). Therefore, in this embodiment, the compressed data is dynamically adjusted through the preset single-packet data rate to achieve high-efficiency communication.

[0074] Further, the transmission of the compressed data according to the preset single-packet data utilization rate may specifically include the following steps: Step 31: If the data volume of the compressed data is less than the minimum value of the target range, calculate a first difference based on the data volume of the compressed data and the minimum value, and increase the data volume of the data to be transmitted according to the first difference; the target range is determined according to the preset single-packet data utilization rate; Step 32: If the data volume of the compressed data is greater than the maximum value of the target range, transmit part of the data in the compressed data, and the data volume of the part of the data is the maximum value of the target range; and calculate a second difference based on the data volume of the compressed data and the maximum value, and reduce the data volume of the data to be transmitted according to the second difference; Step 33: If the data volume of the compressed data is within the target range, transmit the compressed data.

[0075] For a better understanding of Steps 31 - 33, reference can be made to Figure 5 , Figure 5 which is an example diagram of an adaptive data screening method provided by an embodiment of the present invention. The initial data is compressed according to an efficient compression algorithm, and the size of the compressed data is calculated in real time. . By compressing the current data and calculating the data volume of the compressed data, the system can evaluate the actual compression effect and provide a basis for the next screening and adjustment. Specifically, (1) Judge the adaptability of the compressed data size to the target range: Judge whether the size of the compressed data is within the target range bit ~ bit. If it indicates that the data volume to be compressed is insufficient, then additional data needs to be added. At this time, the system will continue to select data with higher priority from the remaining uncompressed data for supplementation to increase the data volume to the target range; if it indicates that the data volume is excessive and some data needs to be reduced. (2) Dynamically adjust the data volume: Calculate the difference between the currently compressed data and the target range, evaluate the data volume that needs to be added or reduced. If data needs to be added, select high-priority data from the priority list and append it to the data to be compressed. If data needs to be reduced, preferentially remove low-priority data. (3) Continuously iterate and optimize: Continuously repeat the above steps until the size of the compressed data meets the requirements of the target range.

[0076] Further, a character set caching mechanism is also provided: Cache the constructed character set, reduce duplicate data processing operations, and greatly improve the processing efficiency of the system. Especially when processing a large amount of data, it can significantly reduce the calculation time and resource consumption.

[0077] Apply the data compression and transmission method provided by the embodiments of the present invention. According to the data characteristics of the data to be transmitted, construct corresponding character sets respectively; perform data evaluation on the character sets, determine corresponding compression algorithms according to the evaluation results and perform compression to obtain compressed data; transmit the compressed data according to the preset single-packet data utilization rate. The present invention aims at the limitations of Beidou short message communication, constructs character sets according to data characteristics, automatically selects an adapted compression algorithm based on the data evaluation results, and transmits the compressed data according to the preset single-packet data utilization rate. Compared with the traditional fixed character set, it improves the data frequency; while ensuring the reliability of data compression, it greatly improves the data transmission efficiency; it can also ensure the single-packet data utilization rate and improve the communication efficiency. For data characteristics, construct a dynamic character set based on mechanisms such as packet grouping optimization, segmentation optimization, and differential coding. Compared with the traditional fixed character set, it can improve the data frequency and greatly improve the data compression efficiency; propose a compression algorithm of pre-filled dictionary and probability interval coding for data with a large number of repeated substrings, medium-low entropy, and high repetition density. Based on the pre-filled dictionary and probability interval coding strategy, it can improve the data compression speed and compression efficiency; for data with fewer repeated substrings, high entropy, and low repetition density, propose a header compression mechanism and a compression algorithm of prime base quantization. Using prime base quantization to divide the probability interval more finely and flexibly, making the probability approach the actual frequency, reducing the quantization error, and improving the compression efficiency; at the same time, design a header compression mechanism for the data header to reduce the header transmission resources and greatly save the data storage space; for the problem that the Beidou short message communication has a low frequency and a small capacity and it is difficult to transmit a large amount of wind field data, design a dynamic priority screening mechanism to divide the priorities of wind turbine data, ensuring that key information is preferentially transmitted under limited communication resources, and improving the efficiency and accuracy of fault diagnosis; for the problem that the single transmission capacity of Beidou short message transmission is limited and the compression ratio is unknown, propose an adaptive data screening mechanism. Based on the adaptive screening adjustment and caching mechanism, ensure that the single-packet data utilization rate reaches M (such as 96%) or more, and improve the communication efficiency. Moreover, the present invention designs a pre-filled dictionary, uses prior knowledge to capture high-frequency patterns, reduces the output of early redundant characters, improves the compression efficiency in the initial stage, avoids inefficiency when the field is empty, and accelerates the compression convergence; enhances the ability to capture short repeated patterns, improves the compression efficiency for locally repeated data, reduces the output of early redundant characters, and improves the compression ratio; at the same time, for the problem of residual symbol redundancy in the output of the LZ77 algorithm, add an encoding strategy to eliminate statistical redundancy; and, based on the improved ANS algorithm for small data, use prime base quantization to divide the probability interval more finely and flexibly, making the probability approach the actual frequency, reducing the quantization error, avoiding state overlap, improving the compression efficiency, and giving full play to the potential of ANS. For small data, the symbol table stored by the ANS compression algorithm will occupy a large amount of memory, so a header compression method is provided to reduce the data overhead.

[0078] In view of the limitations of Beidou short message communication, a data compression and transmission method is constructed based on wind turbine data, which greatly improves the data transmission efficiency while ensuring the reliability of data compression. For a better understanding of the present invention, please refer specifically to Figure 6 , Figure 6 which is a flow example diagram of an overall data compression method provided by an embodiment of the present invention, and specifically may include: Construction of a dynamic character set: For data characteristics, a dynamic SCS (Super-Character Set) is constructed based on composite mechanisms such as packet assembly optimization mechanism, segmentation optimization, and differential coding. Data compression mechanism: Based on data with a large number of repeated substrings, medium and low entropy, and high repetition density, a compression algorithm of pre-filled dictionary and probability interval coding is proposed; based on data with fewer repeated substrings, high entropy, and low repetition density, a header compression mechanism and a prime base quantization compression algorithm are proposed; based on data with uniform character frequency distribution, high entropy, lack of obvious repetition patterns or high-frequency characters, the Deflate compression method is adopted. Data evaluation and preprocessing mechanism: Since the single transmission capacity of Beidou short message transmission is limited, an adaptive data screening mechanism is proposed in the case of unknown compression ratio. Based on the adaptive screening adjustment and compression cache mechanism, it is ensured that the utilization rate of single-packet data reaches more than M (such as 96%).

[0079] Next, the data compression and transmission device provided by the embodiment of the present invention will be introduced. The data compression and transmission described below can be mutually corresponding and referred to the data compression and transmission method described above.

[0080] Please refer specifically to Figure 7 , Figure 7 which is a structural schematic diagram of a data compression and transmission device provided by an embodiment of the present invention, and may include: a character set construction module 100, configured to construct corresponding character sets respectively according to the data characteristics of the data to be transmitted; a compression module 200, configured to evaluate the data of the character set, determine corresponding compression algorithms according to the evaluation results and perform compression to obtain compressed data; a transmission module 300, configured to transmit the compressed data according to a preset single-packet data utilization rate.

[0081] Based on the above embodiment, the character set construction module 100 may include: a first construction unit, configured to perform segmentation optimization on the fault data by analyzing the frequency distribution and time series characteristics of the fault data to construct a first character set; a second construction unit, configured to construct a second character set based on time information; the time information includes the fault start time, end time, and current system time; a third construction unit, configured to construct a third character set based on the wind farm number, wind turbine number, and whether it is a first-occurrence fault.

[0082] Based on the above embodiments, the compression module 200 may include: a calculation unit for calculating the length of repeated substrings, the repetition density, and the data entropy of the character set; a first compression unit for compressing a fourth character set with a length of repeated substrings greater than a first threshold, a repetition density greater than a second threshold, and a data entropy not greater than a third threshold by using a pre-filled dictionary and probability interval coding to obtain first compressed data; a second compression unit for compressing a fifth character set with a length of repeated substrings not greater than the first threshold, a repetition density not greater than the second threshold, and a data entropy greater than the third threshold by using a header compression mechanism and prime base quantization to obtain second compressed data; and a third compression unit for compressing a sixth character set with a uniform character frequency distribution by using the Deflate compression method to obtain third compressed data.

[0083] Based on the above embodiments, the first compression unit may include: a dictionary construction subunit for constructing the pre-filled dictionary based on historical data; a first output subunit for outputting a dictionary match flag bit and an index for characters in the fourth character set that are in the pre-filled dictionary; a second output subunit for, for characters in the fourth character set that are not in the pre-filled dictionary, performing matching by using the LZ77 algorithm, and if the matching is successful, outputting a sliding window flag bit, a distance, and a length; and if the matching fails, outputting an original value flag bit and an original value; a first encoded data determination subunit for using the dictionary match flag bit, the sliding window flag bit, and the original value flag bit as first data to be encoded; a second encoded data determination subunit for using the distance, the length, and the original value as second data to be encoded, third data to be encoded, and fourth data to be encoded respectively; and a probability interval coding subunit for performing probability interval coding on the first data to be encoded, the second data to be encoded, the third data to be encoded, and the fourth data to be encoded respectively to obtain the first compressed data.

[0084] Based on the above embodiments, the probability interval coding subunit includes: an input data determination subunit for using the first data to be encoded, the second data to be encoded, the third data to be encoded, and the fourth data to be encoded as input data respectively; an initial symbol interval determination subunit for defining symbols for the input data and calculating the frequencies of the symbols, and setting initial symbol intervals according to the frequencies of the symbols; a symbol interval update subunit for updating the symbol intervals based on the initial symbol intervals and the symbols until final symbol intervals are obtained; and a first compression subunit for using the shortest binary number within the final symbol intervals as the first compressed data.

[0085] Based on the above embodiments, the second compression unit includes: a statistical subunit, configured to statistically analyze the symbol frequencies of the fifth character set to generate a frequency dictionary and a frequency sequence; the frequency dictionary includes a symbol sequence and a quantized frequency difference sequence; a quantization base determination subunit, configured to calculate the total frequency based on the frequencies of all symbols in the fifth character set, and use a prime number greater than or equal to the total frequency as the quantization base; a second compression subunit, configured to quantize the frequency sequence using the quantization base to obtain a quantized frequency sequence, construct a state space based on the quantized frequency sequence, generate a coding table based on the state space, and convert the fifth character set into a compressed bit stream using the coding table; a third compression subunit, configured to compress the header information using dynamic Bit-Packing and fixed Bit-Packing methods to obtain a header compression result; the header information includes the frequency dictionary and the quantization base.

[0086] Based on any of the above embodiments, the transmission module 300 may include: a data volume increase unit, configured to, if the data volume of the compressed data is less than the minimum value of the target range, calculate a first difference based on the data volume of the compressed data and the minimum value, and increase the data volume of the data to be transmitted according to the first difference; the target range is determined according to the preset single-packet data utilization rate; a data volume decrease unit, configured to, if the data volume of the compressed data is greater than the maximum value of the target range, transmit a part of the compressed data, the data volume of the part of the data being the maximum value of the target range; and calculate a second difference based on the data volume of the compressed data and the maximum value, and decrease the data volume of the data to be transmitted according to the second difference; a transmission unit, configured to, if the data volume of the compressed data is within the target range, transmit the compressed data.

[0087] It should be noted that the order of the modules and units in the above data compression and transmission device can be changed before and after without affecting the logic.

[0088] Apply the data compression and transmission device provided by the embodiment of the present invention. The character set construction module 100 is used to construct corresponding character sets respectively according to the data characteristics of the data to be transmitted; the compression module 200 is used to perform data evaluation on the character sets, determine corresponding compression algorithms according to the evaluation results and perform compression to obtain compressed data; the transmission module 300 is used to transmit the compressed data according to the preset single-packet data utilization rate. This device constructs a character set according to the data characteristics in view of the limitations of Beidou short message communication, automatically selects an adapted compression algorithm based on the data evaluation results, and transmits the compressed data according to the preset single-packet data utilization rate. Compared with the traditional fixed character set, it improves the data frequency; while ensuring the reliability of data compression, it greatly improves the data transmission efficiency; it can also ensure the single-packet data utilization rate and improve the communication efficiency. For the data characteristics, a dynamic character set is constructed based on mechanisms such as packet assembly optimization, segmentation optimization, and differential coding. Compared with the traditional fixed character set, it can improve the data frequency and greatly improve the data compression efficiency; for data with a large number of repeated substrings, medium-low entropy, and high repetition density, a compression algorithm of pre-filled dictionary and probability interval coding is proposed. Based on the pre-filled dictionary and probability interval coding strategy, it can improve the data compression speed and efficiency; for data with fewer repeated substrings, high entropy, and low repetition density, a header compression mechanism and a prime base quantization compression algorithm are proposed. The prime base quantization is used to divide the probability interval more finely and flexibly, making the probability approach the actual frequency, reducing the quantization error, and improving the compression efficiency; at the same time, a header compression mechanism is designed for the data header to reduce the header transmission resources and greatly save the data storage space; for the problem that the Beidou short message communication has a low frequency and small capacity and it is difficult to transmit a large amount of wind field data, a dynamic priority screening mechanism is designed to divide the priorities of wind turbine data, ensuring that key information is preferentially transmitted under limited communication resources, and improving the efficiency and accuracy of fault diagnosis; for the problem that the single transmission capacity of Beidou short message transmission is limited and the compression ratio is unknown, an adaptive data screening mechanism is proposed. Based on the adaptive screening adjustment and caching mechanism, it ensures that the single-packet data utilization rate reaches more than M (such as 96%), improving the communication efficiency. Moreover, the present invention designs a pre-filled dictionary, uses prior knowledge to capture high-frequency patterns, reduces the output of early redundant characters, improves the compression efficiency in the initial stage, avoids inefficiency when the field is empty, and accelerates the compression convergence; enhances the ability to capture short repeated patterns, improves the compression efficiency of local repeated data, reduces the output of early redundant characters, and improves the compression ratio; at the same time, for the problem of residual symbol redundancy in the output of the LZ77 algorithm, an encoding strategy is added to eliminate statistical redundancy; and, based on the improved ANS algorithm for small data, the prime base quantization is used to divide the probability interval more finely and flexibly, making the probability approach the actual frequency, reducing the quantization error, avoiding state overlap, improving the compression efficiency, and giving full play to the potential of ANS. For small data, the symbol table stored by the ANS compression algorithm will occupy a large amount of memory, so a header compression method is provided to reduce the data overhead.

[0089] The data compression and transmission device provided in the embodiments of the present invention will be introduced below. The data compression and transmission device described below can be correspondingly referred to the data compression and transmission method described above.

[0090] Please refer to Figure 8 , Figure 8 , which is a schematic structural diagram of a data compression and transmission device provided in the embodiments of the present invention, and may include: a memory 10 for storing computer programs; a processor 20 for executing the computer programs to implement the above data compression and transmission method.

[0091] The memory 10, the processor 20, and the communication interface 31 all complete mutual communication through the communication bus 32.

[0092] In the embodiments of the present invention, one or more programs are stored in the memory 10. The programs may include program codes, and the program codes include computer operation instructions. In the embodiments of the present invention, the programs for implementing the following functions may be stored in the memory 10: respectively constructing corresponding character sets according to the data characteristics of the data to be transmitted; performing data evaluation on the character sets, determining corresponding compression algorithms according to the evaluation results and performing compression to obtain compressed data; transmitting the compressed data according to the preset single-packet data utilization rate.

[0093] In a possible implementation manner, the memory 10 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function, etc.; the data storage area may store the data created during use.

[0094] In addition, the memory 10 may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include NVRAM. The memory stores an operating system and operation instructions, executable modules or data structures, or subsets thereof, or extended sets thereof. Among them, the operation instructions may include various operation instructions for implementing various operations. The operating system may include various system programs for implementing various basic tasks and processing hardware-based tasks.

[0095] The processor 20 may be a central processing unit (CPU), an application-specific integrated circuit, a digital signal processor, a field programmable gate array, or other programmable logic devices. The processor 20 may be a microprocessor or any conventional processor, etc. The processor 20 may call the programs stored in the memory 10.

[0096] The communication interface 31 may be an interface of a communication module for connecting to other devices or systems.

[0097] Of course, it should be noted that Figure 8 the structure shown does not constitute a limitation on the data compression and transmission device in the embodiments of the present invention. In practical applications, the data compression and transmission device may include more or fewer components than Figure 8 those shown, or combine certain components.

[0098] Next, the readable storage medium provided by the embodiments of the present invention will be introduced. The readable storage medium described below can be correspondingly referred to the data compression transmission method described above.

[0099] The present invention also provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above data compression transmission method are implemented.

[0100] The computer-readable storage medium may include: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0101] In the present specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method part.

[0102] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in the form of hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0103] Finally, it should also be noted that in this text, relationships such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "including", "comprising" or any other variant is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device.

[0104] The above has introduced in detail a data compression and transmission method, apparatus, device and computer-readable storage medium provided by the present invention. Specific examples are used in this text to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A data compression and transmission method, characterized in that, Including: Construct corresponding character sets respectively according to the data characteristics of the data to be transmitted; Perform data evaluation on the character sets, determine corresponding compression algorithms according to the evaluation results and perform compression to obtain compressed data; Transmit the compressed data according to a preset single-packet data utilization rate.

2. The data compression and transmission method according to claim 1, wherein Construct corresponding character sets respectively according to the data characteristics of the data to be transmitted, including: Segment and optimize the fault data by analyzing the frequency distribution and time series characteristics of the fault data, and construct a first character set; Construct a second character set based on time information; the time information includes the start time of the fault, the end time, and the current system time; Construct a third character set based on the wind farm number, the wind turbine number, and whether it is a first-occurrence fault.

3. The data compression and transmission method according to claim 1, characterized in that, Perform data evaluation on the character sets, determine corresponding compression algorithms according to the evaluation results and perform compression to obtain compressed data, including: Calculate the repeated substring length, repeated density, and data entropy of the character sets; Compress a fourth character set with a repeated substring length greater than a first threshold, a repeated density greater than a second threshold, and a data entropy not greater than a third threshold using a pre-filled dictionary and probability interval coding to obtain first compressed data; Compress a fifth character set with a repeated substring length not greater than the first threshold, a repeated density not greater than the second threshold, and a data entropy greater than the third threshold using a header compression mechanism and prime base quantization to obtain second compressed data; Compress a sixth character set with a uniform character frequency distribution using the Deflate compression method to obtain third compressed data.

4. The data compression and transmission method according to claim 3, characterized in that, Compress a fourth character set with a repeated substring length greater than a first threshold, a repeated density greater than a second threshold, and a data entropy not greater than a third threshold using a pre-filled dictionary and probability interval coding to obtain first compressed data, including: Construct the pre-filled dictionary based on historical data; For the characters in the fourth character set that are in the pre-filled dictionary, output a dictionary match flag bit and an index; For the characters in the fourth character set that are not in the pre-filled dictionary, perform matching using the LZ77 algorithm. If the matching is successful, output a sliding window flag bit, a distance, and a length; if the matching fails, output an original value flag bit and an original value; Take the dictionary match flag bit, the sliding window flag bit, and the original value flag bit as first data to be encoded; Take the distance, the length, and the original value as second data to be encoded, third data to be encoded, and fourth data to be encoded respectively; Perform probability interval coding on the first data to be encoded, the second data to be encoded, the third data to be encoded, and the fourth data to be encoded respectively to obtain the first compressed data.

5. The data compression and transmission method according to claim 4, wherein Perform probability interval coding on the first data to be encoded, the second data to be encoded, the third data to be encoded, and the fourth data to be encoded respectively to obtain the first compressed data, including: Take the first data to be encoded, the second data to be encoded, the third data to be encoded, and the fourth data to be encoded as input data respectively; Define symbols for the input data and calculate the frequencies of each symbol, and set initial symbol intervals according to the frequencies of each symbol; Update the symbol interval based on the initial symbol interval and each symbol until the final symbol interval is obtained; Use the shortest binary number within the final symbol interval as the first compressed data.

6. The data compression and transmission method according to claim 3, characterized in that Compress the fifth character set with a repeated substring length not greater than the first threshold, a repeated density not greater than the second threshold, and a data entropy greater than the third threshold using a header compression mechanism and prime base quantization to obtain second compressed data, including: Count the symbol frequencies of the fifth character set to generate a frequency dictionary and a frequency sequence; the frequency dictionary includes a symbol sequence and a quantized frequency difference sequence; Calculate the total frequency based on the frequencies of all symbols in the fifth character set, and use a prime number greater than or equal to the total frequency as the quantization base; Quantize the frequency sequence using the quantization base to obtain a quantized frequency sequence, construct a state space based on the quantized frequency sequence, generate a coding table based on the state space, and convert the fifth character set into a compressed bit stream using the coding table; Compress the header information using dynamic Bit-Packing and fixed Bit-Packing methods to obtain a header compression result; the header information includes the frequency dictionary and the quantization base.

7. The data compression and transmission method according to any one of claims 1 to 6, characterized in that, Transmit the compressed data according to a preset single-packet data utilization rate, including: If the data volume of the compressed data is less than the minimum value of the target range, calculate a first difference based on the data volume of the compressed data and the minimum value, and increase the data volume of the data to be transmitted according to the first difference; the target range is determined according to the preset single-packet data utilization rate; If the data volume of the compressed data is greater than the maximum value of the target range, transmit a part of the compressed data, and the data volume of the part of the data is the maximum value of the target range; and calculate a second difference based on the data volume of the compressed data and the maximum value, and reduce the data volume of the data to be transmitted according to the second difference; If the data volume of the compressed data is within the target range, transmit the compressed data.

8. A data compression and transmission device, characterized in that, Including: A character set construction module for constructing corresponding character sets respectively according to the data characteristics of the data to be transmitted; A compression module for evaluating the data of the character set, determining a corresponding compression algorithm according to the evaluation result and performing compression to obtain compressed data; A transmission module for transmitting the compressed data according to a preset single-packet data utilization rate.

9. A data compression and transmission device, characterized in that, Including: A memory for storing computer programs; A processor for implementing the steps of the data compression and transmission method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are loaded and executed by the processor, the steps of the data compression and transmission method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Data compression method, data compression device, computer equipment and storage medium

    CN115208414A

  • Data transmission method, device and equipment

    CN116915363A

  • RapidIO protocol stack-based compression algorithm optimization method and apparatus, and electronic device

    CN119906760A

Cited By

  • Electronic cigarette data transmission method, electronic cigarette equipment and storage medium

    CN120857189A

  • Data sending method and device, electronic equipment and storage medium

    CN121173824A