Data Compression Method and Device Based on Genetic Algorithm Optimized Sparse Autoencoder

Through the genetic algorithm, the network parameters of the sparse autoencoder are optimized, and the nonlinear mapping capability and convergence speed of shallow neural network models in data compression is solved, achieving more efficient data compression and dimensionality reduction effects.

CN115796268BActive Publication Date: 2025-07-18TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211527709.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-01
Publication Date
2025-07-18
Estimated Expiration
2042-12-01

AI Technical Summary

Technical Problem

In the prior art, shallow neural network models have problems such as poor nonlinear mapping capabilities, slow convergence speed and highly dependence on label data when data compression.

Method used

The data compression method based on genetic algorithm is adopted to optimize the sparse autoencoder, and by determining the network topology of the sparse autoencoder, the network is initialized and the weights and thresholds are optimized, and lossless compression is performed in combination with improved genetic algorithms and arithmetic coding.

Benefits of technology

It improves the nonlinear feature mapping capability of sparse autoencoder, enhances the scope of application of data compression, and improves the search speed and compression accuracy of the network without relying on label data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115796268B_ABST
    Figure CN115796268B_ABST
Patent Text Reader

Abstract

The present invention provides a data compression method and apparatus based on a genetic algorithm-optimized sparse autoencoder, which relates to the technical field of data processing. The method includes: determining the network topology of the sparse autoencoder, initializing the sparse autoencoder network to obtain the initial weights and initial thresholds of the sparse autoencoder network; optimizing the initial weights and initial thresholds of the sparse autoencoder network through an improved genetic algorithm; performing normalization processing and data correction on the uncompressed data to obtain new uncompressed data; inputting the new uncompressed data into the improved sparse autoencoder network, and taking the output data of the hidden layer as the compressed data; performing lossless compression on the compressed data through arithmetic coding to obtain new compressed data. It effectively reduces the dimension of the sample data. The multi-layer network structure of the sparse autoencoder enables the sparse autoencoder to have a powerful non-linear feature mapping ability, and does not require the input data to have labels, thereby improving the applicable range of the compression algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a data compression method and device based on a genetic algorithm for optimizing a sparse autoencoder. Background Art

[0002] In a complex network system covering multiple data sources, the data transmitted to the data center through a wireless sensor network is characterized by a large scale and a variety of types. It has become increasingly important to learn the complex internal features of the data from the massive sensor data, remove the redundant information in the data, obtain a reduced-dimensional representation of the original data, and thus achieve a compressed representation of the data. Currently, data compression through neural networks is achieved by using a shallow neural network model. Using a shallow neural network model for data compression has the deficiencies of poor network non-linear mapping ability, slow convergence speed, and high dependence on labeled data. Summary of the Invention

[0003] Aiming at the problems of poor network non-linear mapping ability, slow convergence speed, and high dependence on labeled data in the existing shallow neural network model for data compression, the present invention proposes a data compression method and device based on a genetic algorithm for optimizing a sparse autoencoder.

[0004] To solve the above technical problems, the present invention provides the following technical solutions:

[0005] On the one hand, a data compression method based on a genetic algorithm for optimizing a sparse autoencoder is provided. The method is applied to an electronic device and includes the following steps:

[0006] S1: Determine the network topology of the sparse autoencoder, initialize the sparse autoencoder network, and obtain the initial weights and initial thresholds of the sparse autoencoder network;

[0007] S2: Optimize the initial weights and initial thresholds of the sparse autoencoder network through an improved genetic algorithm to obtain the optimal weights, optimal thresholds, and an improved sparse autoencoder network of the sparse autoencoder network;

[0008] S3: Perform normalization processing and data correction on the uncompressed data to obtain new uncompressed data;

[0009] S4: Input the new uncompressed data into the improved sparse autoencoder network, and use the output data of the hidden layer as the compressed data;

[0010] S5: Perform lossless compression on the compressed data through arithmetic coding to obtain new compressed data, and complete the compression of the sensor data.

[0011] Optionally, in step S1, determining the network topology of the sparse autoencoder includes:

[0012] Determine the number of input nodes, the number of output nodes, the number of hidden layers, and the number of hidden layer nodes of the sparse autoencoder network;

[0013] Among them, the number of input nodes is the same as the number of output nodes, which is determined by the feature dimension of the input data; the number of hidden layers is set to 1; the number of hidden layer nodes is determined according to the following formula (1):

[0014]

[0015] Among them, h represents the number of hidden layer nodes of the network, m represents the number of input nodes of the network, n represents the number of output nodes of the network, and a is any integer between 1 and 10.

[0016] Optionally, in step S2, the initial weights and initial thresholds of the sparse autoencoder network are optimized by an improved genetic algorithm to obtain the optimal weights, optimal thresholds, and improved sparse autoencoder network of the sparse autoencoder network, including:

[0017] S21: Encode the initial weights and initial thresholds of the sparse autoencoder network by an improved genetic algorithm to obtain an initial population; among them, the encoding of all initial weights and initial thresholds is concatenated as the encoding of an individual;

[0018] S22: Calculate the fitness value of each individual in the population. Among them, the average mean square error of the training set and test set data is the fitness function of the improved genetic algorithm, and the fitness function can be written as:

[0019]

[0020] Among them, N is the total number of samples in the training set and test set, M is the number of samples in the training set, is the predicted output of the sparse autoencoder network, y i is the expected output of the sparse autoencoder network;

[0021] S23: Perform a selection operation by allocating probabilities to individuals through roulette. The probability of an individual being selected in roulette is proportional to the magnitude of its fitness value. Obtain the optimal weights, optimal thresholds, and improved sparse autoencoder network of the sparse autoencoder network. The probability p i of the i-th individual being selected is:

[0022]

[0023] Among them, K is the number of the population, F i is the fitness value of the i-th individual.

[0024] Optionally, the improved genetic algorithm includes:

[0025] Improve the crossover operation and adaptively adjust the crossover probability during the algorithm iteration. The improved crossover probability p c is shown in the following formula (4):

[0026]

[0027] where f * is the maximum fitness of two crossover individuals in the population, and f max is the average fitness of the entire population, and f avg is the average fitness of the entire population, and p cmax is the maximum crossover probability, and p cmin is the minimum crossover probability.

[0028] Optionally, step S2 further includes improving the mutation operation in the improved genetic algorithm according to the following formula (5):

[0029]

[0030] where f is the fitness value of the mutant individual, and p mmax is the maximum mutation probability, and p mmin is the minimum mutation probability.

[0031] Optionally, in step S3, normalize and correct the uncompressed data to obtain new uncompressed data, including:

[0032] Normalize the input parameters of the sparse autoencoder network in their respective dimensions according to the following formula (6):

[0033]

[0034] where x is the obtained uncompressed data; x min is the minimum value of the uncompressed data; x max is the maximum value of the uncompressed data; x * is the new uncompressed data.

[0035] Optionally, in step S4, input the new uncompressed data into the improved sparse autoencoder network, and use the output data of the hidden layer as the compressed data, including:

[0036] The sparse autoencoder network structure includes an input layer, a hidden layer, and an output layer; the input x of the input layer is the uncompressed data, the output h of the hidden layer is the compressed data, and the output of the output layer is y.

[0037] Optionally, in step S5, perform lossless compression on the compressed data through arithmetic coding to obtain new compressed data and complete the sensing data compression, including:

[0038] S51: Statistically count all characters and their occurrences in the compressed data obtained in step S4, continuously divide the interval [0, 1) into multiple sub - intervals, each sub - interval represents a character, and the size of the sub - interval is proportional to the probability of the character appearing in the compressed data obtained in step S4;

[0039] S52: The encoding starts from an initial interval [0, 1), continuously reads the characters in the compressed data obtained in step S4, and finds the interval [L, H) where the character is located;

[0040] S53: Output any decimal number in the obtained interval [L, H) in binary form to obtain the encoded data, and this encoded data is the new compressed data, completing the compression of the sensing data.

[0041] On the one hand, a data compression device based on a genetic - algorithm - optimized sparse auto - encoder is provided. This device is applied to an electronic device and includes:

[0042] A network topology determination module, which is used to determine the network topology of the sparse auto - encoder, initialize the sparse auto - encoder network, and obtain the initial weights and initial thresholds of the sparse auto - encoder network;

[0043] An optimization module, which is used to optimize the initial weights and initial thresholds of the sparse auto - encoder network through an improved genetic algorithm, and obtain the optimal weights, optimal thresholds of the sparse auto - encoder network, and the improved sparse auto - encoder network;

[0044] A data pre - processing module, which is used to perform normalization processing and data correction on the uncompressed data to obtain new uncompressed data;

[0045] A compressed - data acquisition module, which inputs the new uncompressed data into the improved sparse auto - encoder network and uses the output data of the hidden layer as the compressed data;

[0046] A lossless compression module, which is used to perform lossless compression on the compressed data through arithmetic coding to obtain new compressed data, completing the compression of the sensing data.

[0047] Optionally, the network topology determination module is further used for:

[0048] Determine the number of input nodes, output nodes, number of hidden layers, and number of hidden - layer nodes of the sparse auto - encoder network;

[0049] Among them, the number of input nodes is the same as the number of output nodes, which is determined by the feature dimension of the input data; the number of hidden layers is set to 1; the number of hidden - layer nodes is determined according to the following formula (1):

[0050]

[0051] Among them, h represents the number of nodes in the hidden layer of the network, m represents the number of input nodes of the network, n represents the number of output nodes of the network, and a is any integer between 1 and 10.

[0052] On the one hand, an electronic device is provided. The electronic device includes a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the above data compression method based on a genetic algorithm-optimized sparse autoencoder.

[0053] On the one hand, a computer-readable storage medium is provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the above data compression method based on a genetic algorithm-optimized sparse autoencoder.

[0054] The above technical solutions of the embodiments of the present invention have at least the following beneficial effects:

[0055] In the above solution, 1. The sparse autoencoder used in the present invention can learn better features for expressing samples in a harsh environment and can effectively reduce the dimensionality of sample data by imposing restrictions on the hidden layer.

[0056] 2. The multi-layer network structure of the sparse autoencoder used in the present invention endows the sparse autoencoder with a powerful non-linear feature mapping ability, and does not require the input data to have labels, improving the applicable range of the compression algorithm.

[0057] 3. The improved genetic algorithm used in the present invention optimizes the network parameters of the sparse autoencoder, further improving the search speed and compression accuracy of the network. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0059] Figure 1 is a flowchart of a data compression method based on a genetic algorithm-optimized sparse autoencoder provided by an embodiment of the present invention;

[0060] Figure 2 is a flowchart of a data compression method based on a genetic algorithm-optimized sparse autoencoder provided by an embodiment of the present invention;

[0061] Figure 3It is the structural diagram of the sparse autoencoder network provided by the embodiments of the present invention;

[0062] Figure 4 It is the structural diagram of the sparse autoencoder network optimized by the improved genetic algorithm provided by the embodiments of the present invention;

[0063] Figure 5 It is the block diagram of the data compression device based on the genetic algorithm optimized sparse autoencoder provided by the embodiments of the present invention;

[0064] Figure 6 It is the structural schematic diagram of an electronic device provided by the embodiments of the present invention. Detailed implementation manners

[0065] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0066] The embodiments of the present invention provide a data compression method based on a genetic algorithm optimized sparse autoencoder. This method can be implemented by an electronic device, and the electronic device can be a terminal or a server. As Figure 1 shown in the data compression method flowchart, the processing flow of this method can include the following steps:

[0067] S101: Determine the network topology of the sparse autoencoder, initialize the sparse autoencoder network, and obtain the initial weights and initial thresholds of the sparse autoencoder network;

[0068] S102: Optimize the initial weights and initial thresholds of the sparse autoencoder network through an improved genetic algorithm to obtain the optimal weights, optimal thresholds of the sparse autoencoder network, and the improved sparse autoencoder network;

[0069] S103: Perform normalization processing and data correction on the uncompressed data to obtain new uncompressed data;

[0070] S104: Input the new uncompressed data into the improved sparse autoencoder network, and use the output data of the hidden layer as the compressed data;

[0071] S105: Perform lossless compression on the compressed data through arithmetic coding to obtain new compressed data, and complete the sensing data compression.

[0072] Optionally, in step S101, determining the network topology of the sparse autoencoder includes:

[0073] Determine the number of input nodes, output nodes, number of hidden layers, and number of hidden layer nodes of the sparse autoencoder network;

[0074] Among them, the number of input nodes is the same as the number of output nodes, which is determined by the feature dimension of the input data; the number of hidden layers is set to 1; the number of hidden layer nodes is determined according to the following formula (1):

[0075]

[0076] Among them, h represents the number of nodes in the hidden layer of the network, m represents the number of input nodes of the network, n represents the number of output nodes of the network, and a is any integer between 1 and 10.

[0077] Optionally, in step S102, the initial weights and initial thresholds of the sparse autoencoder network are optimized by an improved genetic algorithm to obtain the optimal weights, optimal thresholds of the sparse autoencoder network, and the improved sparse autoencoder network, including:[[]]

[0078] S121: Encode the initial weights and initial thresholds of the sparse autoencoder network by an improved genetic algorithm to obtain an initial population; among them, the encoding of all initial weights and initial thresholds is concatenated as the encoding of an individual.

[0079] S122: Calculate the fitness value of each individual in the population. Among them, the average mean square error of the training set and the test set data is the fitness function of the improved genetic algorithm, and the fitness function can be written as:

[0080]

[0081] Among them, N is the total number of samples in the training set and the test set, M is the number of samples in the training set, is the predicted output of the sparse autoencoder network, y i is the expected output of the sparse autoencoder network;

[0082] S123: Perform a selection operation by allocating probabilities to individuals through roulette. In roulette, the probability of an individual being selected is proportional to the magnitude of its fitness value. Obtain the optimal weights, optimal thresholds of the sparse autoencoder network, and the improved sparse autoencoder network. The probability p i that the i-th individual is selected is:[[]]

[0083]

[0084] Among them, K is the number of the population, F i is the fitness value of the i-th individual.

[0085] Optionally, the improved genetic algorithm includes:[[]]

[0086] Improve the crossover operation, and adaptively adjust the magnitude of the crossover probability during the algorithm iteration. The improved crossover probability p c is shown in the following formula (4):

[0087]

[0088] Among them, f * is the maximum fitness of two crossover individuals in the population, and f max is the average fitness of the entire population, and f avg is the average fitness of the entire population, and p cmax is the maximum crossover probability, and p cmin is the minimum crossover probability.

[0089] Optionally, step S102 further includes improving the mutation operation in the improved genetic algorithm according to the following formula (5):

[0090]

[0091] Among them, f is the fitness value of the mutant individual, and p mmax is the maximum mutation probability, and p mmin is the minimum mutation probability.

[0092] Optionally, in step S103, normalizing and data correcting the uncompressed data to obtain new uncompressed data, including:

[0093] Normalizing the input parameters of the sparse autoencoder network in their respective dimensions according to the following formula (6):

[0094]

[0095] Among them, x is the obtained uncompressed data; x min is the minimum value of the uncompressed data; x max is the maximum value of the uncompressed data; x * is the new uncompressed data.

[0096] Optionally, in step S104, inputting the new uncompressed data into the improved sparse autoencoder network and using the output data of the hidden layer as the compressed data, including:

[0097] The sparse autoencoder network structure includes an input layer, a hidden layer, and an output layer; the input x of the input layer is the uncompressed data, the output h of the hidden layer is the compressed data, and the output of the output layer is y.

[0098] Optionally, in step S105, losslessly compressing the compressed data through arithmetic coding to obtain new compressed data and completing the compression of the sensing data, including:

[0099] S151: Count all characters and their occurrence times in the compressed data obtained in step S104, continuously divide the interval [0, 1) into multiple sub-intervals, where each sub-interval represents a character, and the size of the sub-interval is proportional to the probability of the character appearing in the compressed data obtained in step S104;

[0100] S152: The encoding starts from an initial interval [0, 1), continuously reads the characters in the compressed data obtained in step S4, and finds the interval [L, H) where the character is located;

[0101] S153: Output any decimal number in the obtained interval [L, H) in binary form to obtain the encoded data, and this encoded data is the new compressed data, completing the compression of the sensing data.

[0102] In the embodiment of the present invention, a data compression method based on a genetic algorithm optimized sparse autoencoder is designed. The sparse autoencoder can learn better features for expressing samples in a harsh environment by imposing restrictions on the hidden layer, and can effectively reduce the dimension of the sample data. At the same time, the multi-layer network structure of the sparse autoencoder enables the sparse autoencoder to have a powerful non-linear feature mapping ability, without the need for the input data to have labels, improving the applicable range of the compression algorithm. In addition, the improved genetic algorithm is used to optimize the network parameters of the sparse autoencoder, further improving the network search speed and compression accuracy.

[0103] The embodiment of the present invention provides a data compression method based on a genetic algorithm optimized sparse autoencoder. This method can be implemented by an electronic device, and the electronic device can be a terminal or a server. As Figure 2 shown in the flowchart of the data compression method, the processing flow of this method can include the following steps:

[0104] S201: Determine the network topology of the sparse autoencoder, initialize the sparse autoencoder network, and obtain the initial weights and initial thresholds of the sparse autoencoder network.

[0105] In a feasible implementation, three elements need to be determined to construct a sparse autoencoder: the number of input nodes, the number of output nodes, the number of hidden layers, and the number of hidden layer nodes. Among them, the sparse autoencoder network is as Figure 3 shown.

[0106] Among them, the number of input nodes is the same as the number of output nodes, which is determined by the feature dimension of the input data; the number of hidden layers is set to 1; the number of hidden layer nodes is determined according to the following formula (1):

[0107]

[0108] Among them, h represents the number of nodes in the hidden layer of the network, m represents the number of input nodes of the network, n represents the number of output nodes of the network, and a is an arbitrary integer between 1 and 10.

[0109] S202: Encode the initial weights and initial thresholds of the sparse autoencoder network through an improved genetic algorithm to obtain an initial population; among them, the encoding of all initial weights and initial thresholds is concatenated to form the encoding of an individual.

[0110] In a feasible implementation manner, Figure 4 The structural schematic diagram of optimizing the sparse autoencoder network by the improved genetic algorithm is shown. First, use the improved genetic algorithm to encode the initial values of the sparse autoencoder network to obtain an initial population. The individual encoding uses binary encoding. Each individual is a binary real number string, which consists of four parts: the connection weights between the input layer and the hidden layer, the hidden layer threshold, the connection weights between the hidden layer and the output layer, and the output layer threshold. The encoding of all weights and thresholds is concatenated to form the encoding of an individual.

[0111] S203: Calculate the fitness value of each individual in the population. Among them, the average mean square error of the training set and test set data is the fitness function of the improved genetic algorithm.

[0112] In a feasible implementation manner, the fitness function can be written as:

[0113]

[0114] Among them, N is the total number of samples in the training set and test set, M is the number of samples in the training set, is the predicted output of the sparse autoencoder network, y i is the expected output of the sparse autoencoder network.

[0115] In a feasible implementation manner, obtain the initial weights and thresholds of the sparse autoencoder network according to the individual. After training the sparse autoencoder network with the training data, predict the network output. Use the average mean square error of the training set and test set data as the fitness function of the network. The smaller the fitness function value, the more accurate the training, and the better the prediction accuracy of the model is taken into account.

[0116] S204: Perform a selection operation by allocating probabilities to individuals through roulette wheel. In roulette wheel, the probability of an individual being selected is proportional to the magnitude of its fitness value, and obtain the optimal weights, optimal thresholds of the sparse autoencoder network and the improved sparse autoencoder network.

[0117] In a feasible implementation manner, perform a selection operation, use roulette wheel to allocate probabilities to individuals, and select them to create the i-th individual of the next generation proportional to its fitness value. The probability p i is:

[0118]

[0119] Among them, K is the number of populations, and F i is the fitness value of the i-th individual;

[0120] After selecting individuals using the selection operation, it is necessary to generate a new generation using these selected individuals. The crossover operation is the most prominent feature that differentiates genetic algorithms from traditional optimization algorithms. It mimics the process of gene recombination in natural sexual reproduction, which can not only optimize neural networks but also generate new individuals, thus ensuring the diversity of population individuals. Since individuals use real number coding, the real number crossover method is adopted for the crossover operation. In the working process of traditional genetic algorithms, the crossover probability is generally set as a constant between 0.3 and 0.8. To improve the global search ability and convergence speed of genetic algorithms, the crossover operation is improved, and the size of the crossover probability is adaptively adjusted during the algorithm iteration process. For the improved genetic algorithm, the improved crossover probability p c is shown in the following formula (4):

[0121]

[0122] Among them, f * is the maximum fitness of the two crossover individuals in the population, f max is the average fitness of the entire population, f avg is the average fitness of the entire population, p cmax is the maximum crossover probability, and p cmin is the minimum crossover probability.

[0123] In a feasible implementation, the maximum crossover probability is set to 0.8, and the minimum crossover probability is set to 0.5;

[0124] In a feasible implementation, in the working process of traditional genetic algorithms, the mutation probability is generally set as a constant between 0.001 and 0.1. To improve the local search ability of genetic algorithms, the mutation operation is improved, and the size of the mutation probability is adaptively adjusted during the algorithm iteration process. According to the following formula (5), the mutation operation in the improved genetic algorithm is improved:

[0125]

[0126] Among them, f is the fitness value of the mutant individual, and p mmax is the maximum mutation probability, and p mmin is the minimum mutation probability.

[0127] In a feasible implementation, the maximum mutation probability is set to 0.1, and the minimum mutation probability is set to 0.01.

[0128] S205: Normalize and correct the uncompressed data to obtain new uncompressed data;

[0129] In a feasible implementation, since there may be incorrect data in the obtained uncompressed data, the data with obvious errors is removed, and interpolation is performed according to the real data to obtain reliable data;

[0130] Before training the model, data preprocessing is performed by normalizing the data. Specifically, the input parameters of the sparse autoencoder network are normalized in their respective dimensions according to the following formula (6):

[0131]

[0132] where x is the obtained uncompressed data; x min is the minimum value of the uncompressed data; x max is the maximum value of the uncompressed data; x * is the new uncompressed data. All normalized data is distributed in the interval [0, 1].

[0133] S206: Input the new uncompressed data into the improved sparse autoencoder network, and use the output data of the hidden layer as the compressed data.

[0134] In a feasible implementation, the sparse autoencoder network structure includes an input layer, a hidden layer, and an output layer; the input x of the input layer is the uncompressed data, the output h of the hidden layer is the compressed data, and the output of the output layer is y.

[0135] S207: Count all the characters and the number of occurrences of the compressed data obtained in step S206, continuously divide the interval [0, 1) into multiple sub-intervals, each sub-interval represents a character, and the size of the sub-interval is proportional to the probability of the character appearing in the compressed data obtained in step S206;

[0136] S209: Encoding starts from an initial interval [0, 1), continuously reads the characters in the compressed data obtained in step S206, and finds the interval [L, H) where the character is located;

[0137] S210: Output any decimal number in the obtained interval [L, H) in binary form to obtain the encoded data, and this encoded data is the new compressed data, completing the compression of the sensing data.

[0138] In an embodiment of the present invention, a data compression method based on a genetic algorithm-optimized sparse autoencoder is designed. The sparse autoencoder can learn better features for expressing samples in a harsh environment by imposing restrictions on the hidden layer, and can effectively reduce the dimension of sample data. At the same time, the multi-layer network structure of the sparse autoencoder endows it with a powerful non-linear feature mapping ability, without the need for the input data to have labels, improving the applicable range of the compression algorithm. In addition, the improved genetic algorithm is used to optimize the network parameters of the sparse autoencoder, further improving the network search speed and compression accuracy.

[0139] Figure 5 is a block diagram of a data compression device based on a genetic algorithm-optimized sparse autoencoder shown according to an exemplary embodiment. Referring to Figure 5 , the device 300 includes:

[0140] A network topology structure determination module 310, configured to determine the network topology structure of the sparse autoencoder, initialize the sparse autoencoder network, and obtain the initial weight and initial threshold of the sparse autoencoder network;

[0141] An optimization module 320, configured to optimize the initial weight and initial threshold of the sparse autoencoder network through an improved genetic algorithm, and obtain the optimal weight, optimal threshold of the sparse autoencoder network, and the improved sparse autoencoder network;

[0142] A data preprocessing module 330, configured to perform normalization processing and data correction on the uncompressed data to obtain new uncompressed data;

[0143] A compressed data acquisition module 340, configured to input the new uncompressed data into the improved sparse autoencoder network, and use the output data of the hidden layer as the compressed data;

[0144] A lossless compression module 350, configured to perform lossless compression on the compressed data through arithmetic coding to obtain new compressed data, and complete the sensing data compression.

[0145] Optionally, the network topology structure determination module 310 is further configured to:

[0146] Determine the number of input nodes, the number of output nodes, the number of hidden layers, and the number of hidden layer nodes of the sparse autoencoder network;

[0147] Wherein, the number of input nodes is the same as the number of output nodes, which is determined by the feature dimension of the input data; the number of hidden layers is set to 1; the number of hidden layer nodes is determined according to the following formula (1):

[0148]

[0149] Among them, h represents the number of nodes in the hidden layer of the network, m represents the number of input nodes of the network, n represents the number of output nodes of the network, and a is an arbitrary integer between 1 and 10.

[0150] Optionally, the optimization module 320 is further configured to encode the initial weights and initial thresholds of the sparse autoencoder network through an improved genetic algorithm to obtain an initial population; among them, the encoding of all the initial weights and initial thresholds is concatenated to form the encoding of an individual.

[0151] Calculate the fitness value of each individual in the population. Among them, the average mean square error of the training set and the test set data is the fitness function of the improved genetic algorithm, and the fitness function can be written as:

[0152]

[0153] Among them, N is the total number of samples in the training set and the test set, and M is the number of samples in the training set. is the predicted output of the sparse autoencoder network, and y i is the expected output of the sparse autoencoder network.

[0154] Perform a selection operation by allocating probabilities to individuals through roulette. In roulette, the probability of an individual being selected is proportional to the magnitude of its fitness value, and obtain the optimal weights, optimal thresholds of the sparse autoencoder network, and the improved sparse autoencoder network. The probability p of the i-th individual being selected i is:

[0155]

[0156] Among them, K is the number of the population, and F i is the fitness value of the i-th individual.

[0157] Optionally, the improved genetic algorithm includes:

[0158] Improve the crossover operation and adaptively adjust the magnitude of the crossover probability during the algorithm iteration. The improved crossover probability p c is shown in the following formula (4):

[0159]

[0160] Among them, f * is the maximum fitness of the two crossover individuals in the population, f max is the average fitness of the entire population, f avg is the average fitness of the entire population, p cmax is the maximum crossover probability, and p cmin is the minimum crossover probability.

[0161] Optionally, the optimization module 320 is further configured to improve the mutation operation in the improved genetic algorithm according to the following formula (5):

[0162]

[0163] where f is the fitness value of the mutant individual, p mmax is the maximum mutation probability, and p mmin is the minimum mutation probability.

[0164] Optionally, the data preprocessing module 330 is further configured to

[0165] normalize the input parameters of the sparse autoencoder network in their respective dimensions according to the following formula (6):

[0166]

[0167] where x is the obtained uncompressed data; x min is the minimum value of the uncompressed data; x max is the maximum value of the uncompressed data; x * is the new uncompressed data.

[0168] Optionally, the compressed data acquisition module 340 is further configured that the sparse autoencoder network structure includes an input layer, a hidden layer, and an output layer; the input x of the input layer is uncompressed data, the output h of the hidden layer is compressed data, and the output of the output layer is y.

[0169] Optionally, the lossless compression module 350 is further configured to: count all the characters and the number of occurrences of the compressed data obtained in the compressed data acquisition module 340, continuously divide the interval [0, 1) into multiple sub-intervals, each sub-interval represents a character, and the size of the sub-interval is proportional to the probability of the character appearing in the compressed data obtained in step S104;

[0170] The encoding starts from an initial interval [0, 1), continuously reads the characters in the compressed data obtained in the compressed data acquisition module 340, and finds the interval [L, H) where the character is located;

[0171] Output any decimal number in the obtained interval [L, H) in binary form to obtain the encoded data, and this encoded data is the new compressed data, completing the compression of the sensing data.

[0172] In the above - mentioned manner, a data compression method based on a genetic - algorithm - optimized sparse auto - encoder is designed. The sparse auto - encoder can, by imposing restrictions on the hidden layer, learn better features for expressing samples in a harsh environment, effectively reduce the dimensionality of sample data, and at the same time, the multi - layer network structure of the sparse auto - encoder endows it with a powerful non - linear feature mapping ability, without the need for the input data to have labels, thus expanding the applicable scope of the compression algorithm. In addition, the improved genetic algorithm is used to optimize the network parameters of the sparse auto - encoder, further improving the network's search speed and compression accuracy.

[0173] Figure 6 FIG. 4 is a schematic structural diagram of an electronic device 400 provided by an embodiment of the present invention. The electronic device 400 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 401 and one or more memories 402. Among them, at least one instruction is stored in the memory 402, and the at least one instruction is loaded and executed by the processor 401 to implement the steps of the following data compression method based on a genetic - algorithm - optimized sparse auto - encoder:

[0174] S1: Determine the network topology of the sparse auto - encoder, initialize the sparse auto - encoder network, and obtain the initial weights and initial thresholds of the sparse auto - encoder network;

[0175] S2: Optimize the initial weights and initial thresholds of the sparse auto - encoder network through the improved genetic algorithm to obtain the optimal weights, optimal thresholds of the sparse auto - encoder network, and the improved sparse auto - encoder network;

[0176] S3: Perform normalization processing and data correction on the uncompressed data to obtain new uncompressed data;

[0177] S4: Input the new uncompressed data into the improved sparse auto - encoder network, and use the output data of the hidden layer as the compressed data;

[0178] S5: Perform lossless compression on the compressed data through arithmetic coding to obtain new compressed data, and complete the sensing data compression.

[0179] In an exemplary embodiment, a computer - readable storage medium is also provided, such as a memory including instructions. The above - mentioned instructions can be executed by a processor in a terminal to complete the above - mentioned sensing data compression method. For example, the computer - readable storage medium can be a ROM, a random - access memory (RAM), a CD - ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0180] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. The said program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a disk, an optical disc, etc.

[0181] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A data compression method based on a genetic algorithm-optimized sparse autoencoder, characterized in that Including the following steps: S1: Determine the network topology of the sparse autoencoder, initialize the sparse autoencoder network, and obtain the initial weights and initial thresholds of the sparse autoencoder network; S2: Optimize the initial weights and initial thresholds of the sparse autoencoder network through an improved genetic algorithm to obtain the optimal weights, optimal thresholds, and the improved sparse autoencoder network of the sparse autoencoder network; Optimizing the initial weights and initial thresholds of the sparse autoencoder network through an improved genetic algorithm to obtain the optimal weights, optimal thresholds, and the improved sparse autoencoder network of the sparse autoencoder network, including: S21: Encode the initial weights and initial thresholds of the sparse autoencoder network through an improved genetic algorithm to obtain an initial population; wherein, the concatenation of the encodings of all initial weights and initial thresholds is the encoding of an individual; S22: Calculate the fitness value of each individual in the population. Among them, the average mean square error of the training set and the test set data is the fitness function of the improved genetic algorithm, and the fitness function can be written as: where N is the total number of samples in the training set and the test set, and M is the number of samples in the training set. is the predicted output of the sparse autoencoder network, and y i is the expected output of the sparse autoencoder network; S23: Perform a selection operation by allocating probabilities to individuals through roulette wheel selection. In roulette wheel selection, the probability of an individual being selected is proportional to the magnitude of its fitness value, obtaining the optimal weights, optimal thresholds of the sparse autoencoder network, and the improved sparse autoencoder network. The probability p that the i-th individual is selected is i as follows: Among them, K is the number of populations, and F i is the fitness value of the i-th individual; S3: Normalize and correct the uncompressed data to obtain new uncompressed data; S4: Input the new uncompressed data into the improved sparse autoencoder network, and use the output data of the hidden layer as the compressed data; S5: Losslessly compress the compressed data through arithmetic coding to obtain new compressed data, and complete the sensing data compression.

2. The method according to claim 1, wherein In the step S1, determining the network topology of the sparse autoencoder includes: Determine the number of input nodes, the number of output nodes, the number of hidden layers, and the number of hidden layer nodes of the sparse autoencoder network; Among them, the number of input nodes is the same as the number of output nodes, which is determined by the feature dimension of the input data; the number of hidden layers is set to 1; the number of hidden layer nodes is determined according to the following formula (1): Where h represents the number of hidden layer nodes in the network, m represents the number of input nodes in the network, n represents the number of output nodes in the network, and a is any integer between 1 and 10.

3. The method according to claim 1, characterized in that The improved genetic algorithm includes: Improve the crossover operation, adaptively adjust the size of the crossover probability during the algorithm iteration process, and the improved crossover probability p c is shown in the following formula (4): Among them, f * is the maximum fitness of two crossover individuals in the population, f max is the average fitness of the entire population, f avg is the average fitness of the entire population, p cmax is the maximum crossover probability, p cmin is the minimum crossover probability.

4. The method according to claim 3, wherein The step S2 further includes improving the mutation operation in the improved genetic algorithm according to the following formula (5): Among them, f is the fitness value of the mutant individual, and p mmax is the maximum mutation probability, and p mmin is the minimum mutation probability.

5. The method according to claim 1, characterized in that, In the step S3, normalizing and correcting the uncompressed data to obtain new uncompressed data, including: Normalize the input parameters of the sparse autoencoder network in their respective dimensions according to the following formula (6): Among them, x is the obtained uncompressed data; x min is the minimum value of the uncompressed data; x max is the maximum value of the uncompressed data; x * is the new uncompressed data.

6. The method according to claim 1, characterized in that, In the step S4, inputting the new uncompressed data into the improved sparse autoencoder network and using the output data of the hidden layer as the compressed data, including: The sparse autoencoder network structure includes an input layer, a hidden layer, and an output layer; the input x of the input layer is the uncompressed data, the output h of the hidden layer is the compressed data, and the output of the output layer is y.

7. The method according to claim 1, wherein In the step S5, losslessly compress the compressed data through arithmetic coding to obtain new compressed data, and complete the sensing data compression, including: S51: Statistically count all characters in the compressed data obtained in step S4 and the number of occurrences thereof, continuously divide the interval [0, 1) into multiple sub-intervals, each sub-interval representing a character, and the size of the sub-interval being proportional to the probability of the character occurring in the compressed data obtained in step S4; S52: Encoding starts from an initial interval [0, 1), continuously reads the characters in the compressed data obtained in step S4, and finds the interval [L, H) where the character is located; S53: Output any decimal number in the obtained interval [L, H) in binary form to obtain the encoded data, and this encoded data is the new compressed data, completing the compression of the sensing data.

8. A data compression device based on a genetic algorithm-optimized sparse autoencoder, characterized in that, The device is applicable to the method of any one of the above claims 1-7, and the device includes: A network topology determination module, configured to determine the network topology of the sparse autoencoder, initialize the sparse autoencoder network, and obtain the initial weights and initial thresholds of the sparse autoencoder network; An optimization module, configured to optimize the initial weights and initial thresholds of the sparse autoencoder network through an improved genetic algorithm to obtain the optimal weights, optimal thresholds of the sparse autoencoder network, and the improved sparse autoencoder network; A data preprocessing module, configured to perform normalization processing and data correction on the uncompressed data to obtain new uncompressed data; A compressed data acquisition module, which inputs the new uncompressed data into the improved sparse autoencoder network, and uses the output data of the hidden layer as the compressed data; A lossless compression module, configured to perform lossless compression on the compressed data through arithmetic coding to obtain new compressed data, completing the compression of the sensing data.

9. The device according to claim 8, characterized in that, The network topology determination module, is further configured to: determine the number of input nodes, the number of output nodes, the number of hidden layers, and the number of hidden layer nodes of the sparse autoencoder network; wherein, the number of input nodes is the same as the number of output nodes, and is determined by the feature dimension of the input data; the number of hidden layers is set to 1; the number of hidden layer nodes is determined according to the following formula (1): where h represents the number of hidden layer nodes of the network, m represents the number of input nodes of the network, n represents the number of output nodes of the network, and a is any integer between 1 and 10.

Citation Information

Patent Citations

  • Genetic algorithm based on stacked denoising sparse autocoder

    CN107609648A

  • Ultra-long linear annular wireless network data compression method

    CN111711970A