A multi-image comprehensive quality-preserving compression method based on entropy value analysis
Through information entropy analysis and reinforcement learning technology, the image compression ratio is dynamically adjusted, and the problem of effective visual information loss after image compression in the prior art is solved, and the effect of efficient quality-saving compression in machine vision applications is achieved.
Patent Information
- Application Number
- CN202510240111.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-03-03
AI Technical Summary
Existing image compression technology fails to effectively balance the compression efficiency with the retention of image effective visual information, especially in machine vision applications, where the fixed compression ratio leads to the loss of effective visual information after image compression.
By introducing information entropy analysis and reinforcement learning technology, the compression ratio of images is dynamically adjusted, and the best compression strategies for different images are independently learned and adapted to different images according to the information entropy, original size and device load of the image.
While reducing image volume, the effective visual data information of the image is retained to the maximum extent, significantly improving the effectiveness and practicality of compressed images, especially in application scenarios where high-precision image recognition and processing are required.
Smart Images

Figure CN119743627B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of enhancing deep learning image processing, and particularly to a multi-image comprehensive quality-preserving compression method based on entropy value analysis. Background Art
[0002] In machine vision applications, it is usually necessary to compress the original images. Firstly, machine vision systems often need to process a large amount of image data. For example, in industrial automation, production line cameras capture hundreds of images per second. Compression can significantly reduce the storage space requirements and lower the hardware costs. Secondly, in real-time monitoring and security systems, compressing images can speed up the data transmission speed, ensure the system responds quickly, and enhance the monitoring ability of security cameras in the network. In addition, in autonomous driving vehicles, compression technology reduces the computational burden during subsequent processing, making image recognition and object detection algorithms more efficient and improving driving safety. At the same time, in low-bandwidth environments, such as drone surveillance missions, compression technology can effectively reduce the network bandwidth occupancy and ensure the transmission of important information. In summary, image compression in machine vision not only improves the storage and transmission efficiency but also optimizes the computational performance, and is a key step in achieving efficient and intelligent processing.
[0003] However, in machine vision applications, such as autonomous vehicle recognition and security monitoring, the effectiveness of the visual data information in images is extremely crucial. The loss of information during the compression process may lead to a significant decline in the practical value of the images. For example, when an object recognition model (such as YOLO algorithm) processes the compressed image, due to the loss of image details, it may not be able to accurately identify or locate the object. Without image compression, since there is a large amount of redundant information in the visual data of the images, this redundant information not only does not contribute to the practical value of the images but also greatly increases the volume of the images, resulting in unnecessary consumption of storage resources and transmission time. In addition, traditional compression methods often use a fixed compression strategy for all images, without considering the complexity of the image content and the actual requirements in specific applications. This "one-size-fits-all" method cannot effectively balance the compression efficiency and the retention of effective visual information in the images, resulting in unsatisfactory effects in important image scenarios.
[0004] To solve these problems, there is an urgent need for an innovative image compression method. Summary of the Invention
[0005] To solve the above problems, this application proposes an innovative image compression method. The core of this method is to introduce the information entropy of the image into the scope of consideration for image compression, and dynamically adjust the compression ratio according to information such as the information entropy of the image, the original size of the image, and the processing load of autonomous driving, so as to reduce the image volume while retaining as much effective visual data information in the image as possible. By introducing reinforcement learning technology, this method can autonomously learn and adapt to the optimal compression strategy for different images, thereby maximizing the practical value of the image. Especially in application scenarios that require high-precision image recognition and processing, it can significantly improve the effectiveness and practicality of the compressed image.
[0006] A multi-picture comprehensive quality-preserving compression method based on entropy value analysis provided by the present invention is specifically as follows:
[0007] S1. Perform information entropy analysis and calculation on the pictures generated in machine vision applications to obtain the information entropy of the pictures;
[0008] S2. Combine the information entropy of the pictures, the picture size, and the machine vision device with reinforcement learning technology, and through reinforcement learning training, obtain the corresponding DQN network. The machine vision device uses the DQN network to process the generated pictures to obtain the target compression ratio, and obtain the target compression scheme for the pictures in machine vision applications.
[0009] Preferably, the expression for obtaining the information entropy of the pictures by performing information entropy analysis and calculation on the pictures generated in machine vision applications is:
[0010] ;
[0011] where is the information entropy of the picture, represents the probability of the pixel pair appearing in the picture.
[0012] Preferably, among them, the calculation formula of is as follows:
[0013] ;
[0014] where, in the pixel pair inside, represents the gray value of the central pixel in the sliding window, b is the average gray value of the pixels in the sliding window excluding the central pixel, is the frequency of the pixel pair appearing in all pixels of the entire picture, is the number of pixels, is a specific visual data processing task, equivalent to a picture, in represents the order of the picture itself.
[0015] Preferably, the specific content of S2 includes defining DQN the state state , defining DQN the action action and DQN the reward reward ;
[0016] Among them, the specific content of defining DQN the state state is:
[0017] Define the state state as , where e represents the information entropy of the image, reflecting the complexity and information volume of the image, s represents the original size of the image, represents the number of pictures to be processed generated by the machine vision device at the same moment;
[0018] Among them, the specific content of defining DQN the action action is:
[0019] For an original picture, DQN the action performed on it
[0020] ;
[0021] Preferably, DQN the reward reward is:
[0022] For the picture processing task , define the picture processing task from generation, calculate the information entropy of the picture, perform compression until it is processed by a specific application YOLO the time taken to complete the processing is , and the accuracy rate of the processing result is ;
[0023] Define that after the picture processing task is processed and completed, the total number of pictures processed by the specific application is ;
[0024] Furthermore, define that after a specific picture processing task is processed and completed, DQN the reward obtained
[0025] 。
[0026] Preferably, DQN the network training steps are as follows:
[0027] Step 1: Initialize DQN the Q network and the experience replay pool D ;
[0028] Step 2: Start multiple rounds of training;
[0029] The specific content of starting multiple rounds of training in Step 2 is as follows:
[0030] Step 2.1: Obtain the first state , which is ([[]] , , ) from the external environment;
[0031] Step 2.2: DQN With a probability of ε, randomly select an action, or with a probability of 1 - ε, select an action with a value of argmax_a'Q ( state,a' ), and record the selected action as ;
[0032] Step 2.3: DQN Execute the action , obtain the reward , then obtain the new state next_state from the external environment, and then store the ([[]] state, , ,next_state ) tuple into the experience replay pool D , set the new state next_ state obtained from the external environment as the current state , and at this time, check the experience replay pool D ;
[0033] If the size of the experience replay pool is sufficient, randomly extract a small batch of transitions ([[]] D ) from state, , , next_state ), calculate = + for each transition, and then use to update the network Q with the loss;
[0034] If this is not the last round of training, decrease the value of ε and repeat steps 2.1 to 2.3.
[0035] In summary, a multi-image comprehensive quality-preserving compression method based on entropy value analysis according to the present invention is more effective than traditional techniques in balancing time consumption and accuracy under various working intensities. At low working intensities, it maintains a low processing time and acceptable accuracy, while at high working intensities, it adapts to the increased working intensity, ensuring an efficient processing time while improving accuracy.
[0036] The technical method of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings
[0037] Figure 1 Relates to the relationship among the information entropy, compression ratio, and accuracy rate of a multi-image comprehensive quality-preserving compression method based on entropy value analysis according to the present invention;
[0038] Figure 2 In a multi-image comprehensive quality-preserving compression method based on entropy value analysis according to the present invention and reward Relates to the relationship among the accuracy rate and time;
[0039] Figure 3 Is the cumulative distribution function of time ( CDF ) graph under low working intensity of a multi-image comprehensive quality-preserving compression method based on entropy value analysis according to the present invention;
[0040] Figure 4 Is the cumulative distribution function of accuracy rate ( CDF ) graph under low working intensity of a multi-image comprehensive quality-preserving compression method based on entropy value analysis according to the present invention;
[0041] Figure 5 Is the cumulative distribution function of time ( CDF ) graph under high working intensity of a multi-image comprehensive quality-preserving compression method based on entropy value analysis according to the present invention;
[0042] Figure 6 Is the cumulative distribution function of accuracy rate ( CDF ) graph under high working intensity of a multi-image comprehensive quality-preserving compression method based on entropy value analysis according to the present invention. Detailed Embodiments
[0043] The technical method of the present invention will be further described below with reference to the accompanying drawings and embodiments. It should be noted that: unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps described in these embodiments do not limit the scope of the present application.
[0044] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way a limitation on the present application, its application, or its use.
[0045] Technologies, systems, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the technologies, systems, and devices should be regarded as part of the specification.
[0046] In all the examples shown and discussed here, any specific values should be construed as merely exemplary and not as limitations. Thus, other examples of the exemplary embodiments may have different values.
[0047] Unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meaning as understood by those of ordinary skill in the field to which the present invention pertains.
[0048] The main object of the present invention is to solve some deficiencies existing in the existing image compression technology, especially when compressing images in machine vision applications, the relationship between the volume after image compression and the retention of effective visual data information of the image is not fully considered, and then the problem that the effective visual information is lost after the image is compressed according to a fixed compression ratio.
[0049] The present invention provides a multi-picture comprehensive quality-preserving compression method based on entropy value analysis. S1. Perform information entropy analysis and calculation on the pictures generated in machine vision applications to obtain the information entropy of the pictures.
[0050] Information entropy is a statistical measure of the amount of information. For pictures, information entropy can be regarded as a quantitative index of the complexity of the visual data content of the image, reflecting the diversity and uncertainty of the image information. The two-dimensional information entropy of an image particularly focuses on the spatial relationship between pixels in the image and provides a method for measuring the structural complexity of the image for the compression method. The present invention uses to represent an image processing task generated by the application . For a specific visual data , first calculate the two-dimensional entropy (amount of information) of the image. The entropy of
[0051] ;
[0052] Here represents the picture information entropy calculation function, and this function can be specifically expressed as:
[0053] ;
[0054] In this function, represents the probability of the pixel pair appearing in the picture, and its calculation formula is as follows:
[0055] ;
[0056] In the above formula, the pixel pair in represents the gray value of the central pixel within the sliding window, b is the average gray value of the pixels within the sliding window excluding the central pixel, is the pixel pair appearance frequency among all pixels of the entire picture, is the number of pixels, is a specific visual data processing task, equivalent to a picture, in represents the order of the picture itself. For example, if an application generates pictures from running to ending, then the order of this picture is i.
[0057] Through the above formula, the present invention obtains the information entropy of the picture.
[0058] Next, by combining with reinforcement learning technology, the optimal compression method for pictures in machine vision applications is realized. The present invention uses DQN the reinforcement learning algorithm. Through DQN continuous learning and updating, a deep reinforcement learning neural network that can make the best compression decision based on three pieces of information, namely the information entropy of the picture, the picture size, and the number n of pictures to be processed generated by the machine vision device at the same moment, is trained, that is, DQN the (Deep Reinforcement Learning) neural network.
[0059] S2. Combine the information entropy of the picture, the picture size, and the machine vision device with reinforcement learning technology. Through reinforcement learning training, the corresponding DQN network is obtained. The machine vision device uses DQN the network to process the generated pictures, obtain the target compression ratio, and obtain the target compression scheme for pictures in machine vision applications.
[0060] Preferably, the specific content in S2 includes the definition of the DQN state state , the definition of the DQN action action , and the definition of the DQN reward reward ;
[0061] Among them, the specific content of the definition of the DQN state state is as follows:
[0062] Define the statestate For , where e represents the information entropy of the image, reflecting the complexity and information volume of the image, s represents the original size of the image, and k represents the number of pictures that need to be processed generated by the machine vision device at the same moment;
[0063] Among them, for DQN the action of action the specific content defined is:
[0064] For an original picture, DQN the action performed on it is:
[0065] ;
[0066] Preferably, DQN the reward of reward the specific content is:
[0067] For each picture processing task , define the time from its generation, calculating the information entropy of the picture, compressing to being processed by a specific application YOLO until it is processed and completed as , and the accuracy rate of the processing result is ;
[0068] Define that after a specific picture processing task is processed and completed, the total number of pictures processed by the specific application is ;
[0069] Furthermore, define that after a specific picture processing task is processed and completed, DQN the reward obtained is:
[0070] .
[0071] Preferably, DQN the training steps of the network are:
[0072] Step 1, Initialize DQN the Q network and the experience replay pool D ;
[0073] Step 2, Start multiple rounds of training;
[0074] The specific content of starting multiple rounds of training in Step 2 is:
[0075] Step 2.1, Obtain the first state from the external environment, which is ( , , );
[0076] Step 2.2, DQN With probability ε, randomly select an action, or with probability 1 - ε, select the action with value argmax_a'Q ( state,a' ) and record the selected action as ;
[0077] Step 2.3, DQN Execute the action , obtain the reward , then obtain the new state next_state from the outside world, and then store the ( state , , , next_state ) tuple into the experience replay pool D , and set the new state next_ state obtained from the outside world as the current state . At this time, check the experience replay pool D ;
[0078] If the size of the experience replay pool is sufficient, randomly draw a small batch of transitions ( D ) from state , , , next_state ), calculate = + for each transition, and then use to update the network Q with the loss;
[0079] At this time, if it is not the last round of training, decrease the value of ε and repeat Steps 2.1 to 2.3.
[0080] Among them, DQN The specific coding for training the network is:
[0081] Initialize DQN the network Q ,
[0082] Initialize the experience replay pool D ,
[0083] for episode = 1 to M :
[0084] Initialize the state state = ( , , )
[0085] while not done:
[0086] Selection action :
[0087] Randomly select an action with probability ε
[0088] Or with probability 1-ε Select an action = argmax_a'Q ( state,a' )
[0089] Execute the action , and obtain a reward and a new state next_state
[0090] Store the transition (state, , , next_state) into the experience replay pool D
[0091] Let state = next_state
[0092] The experience replay pool is sufficient:
[0093] Randomly sample a mini - batch of transitions from D ( state, , , next_state )
[0094] Calculate for each transition:
[0095] = +
[0096] Use to update the network for the loss Q
[0097] Decrease ε (exploration rate).
[0098] After analyzing the information entropy, compression ratio, and accuracy of the compressed images in a large amount of image data, the following results can be obtained:
[0099] Figure 1Shows the relationship between image entropy and accuracy under different compression ratios. The results show that at the same compression ratio, images with high entropy have higher accuracy, indicating better image recognition performance. In addition, the results show that high-entropy images have better accuracy at λ = 0.9 than low-entropy images at a compression ratio of λ = 0.7. This shows that images with higher entropy can better retain information, have greater compression potential, and thus lower time consumption. These findings emphasize the key role of visual data entropy in determining the compression ratio of visual tasks.
[0100] DQN After the network is trained, it is tested on the test set to obtain the time consumption and accuracy after processing for each image processing task. The method of the present invention adjusts DQN of reward in the calculation formula value to adjust the emphasis on time consumption and accuracy after image processing. The results are as follows:
[0101] Figure 2 Shows the weighted gain of accuracy and time in when reward . The figure shows that at around 0.7, the weighted accuracy and weighted time are equal, marking a reward close to zero state. When is less than 0.7, reward is below zero, and when is greater than 0.7, reward exceeds zero. We set this balance point as = 0.7, indicating that DQN considers accuracy and time almost equally. In addition, we define = 0.2 to indicate that DQN mainly focuses on time consumption, and = 0.9 to indicate that DQN mainly focuses on output accuracy. It should be noted that there is no optimal compression ratio applicable to all cases. After sufficient training, DQN makes the best compression decision for the current state received according to the current state . During training, the appropriate value can be selected according to the trainer's emphasis on time consumption and accuracy to train the corresponding DQN network.
[0102] We also compared two other strategies, namely N-C : Do not compress the images generated by each machine vision application, Random : Take a random compression ratio for the images generated by each machine vision application. The results are as follows:
[0103] Figures 3 - 6 For the cumulative distribution function distributions of time and accuracy under low and high working intensity conditions. The present invention observes that in the high working intensity condition scenario, the maximum time can exceed 2.5 seconds, while in the low working intensity scenario, the maximum time is about 1.5 seconds, which is due to the increased competition for resources between different pictures under high working intensity conditions.
[0104] As Figure 4 and Figure 6 shown, except = 0.2 and = 0.7 ( represents the parameter that adjusts the weights of both the picture accuracy rate and the picture processing time in reward ), for all other algorithms, the accuracy rate for certain tasks is higher than 0.8 under both low and high working intensities. This is because = 0.2 and = 0.7 prioritize time efficiency, resulting in a lower accuracy rate.
[0105] Figure 3 and Figure 5 illustrate the time distributions under low and high working intensity conditions respectively. In low traffic, only the tasks of N - C exceed 1 second, while = 0.7, = 0.9 and Random are mostly between 0.9 and 1.0 seconds. Under high working intensity, due to the fierce resource competition, = 0.9, Random and N-C 's tasks exceed 2 seconds. While = 0.7 maintains a stable balance in both cases, with a stable cumulative distribution function curve and a maximum value significantly lower than N-C , demonstrating the advantage of = 0.7.
[0106] Figure 4 and Figure 6 show the accuracy rate distributions under two working intensities. = 0.2 hardly shows tasks with an accuracy rate higher than 0.2. In the range of [0.4 - 0.6] under high working intensity, Random and = 0.7 are both higher than = 0.9 and N-C , indicating that = 0.9 and N-C have more tasks and higher accuracy rates, Random performs the worst, N-C performs the best. When the accuracy rate reaches 0.7, = 0.9 and N-C the curves are almost the same, highlighting the effectiveness of = 0.9 in the accuracy priority ranking.
[0107] It proves the effectiveness of the present method in balancing time consumption and accuracy under various working intensities. At low working intensities, it maintains a low processing time and acceptable accuracy, while at high working intensities, it adapts to the increased working intensity, ensuring efficient processing time while improving accuracy.
[0108] The present invention dynamically adjusts the compression ratio based on the information entropy of the pictures generated by machine vision applications. Combining with the reinforcement learning DQN algorithm, using the state state as the basis for dynamically adjusting the compression level, where represents the information entropy of the image, reflecting the complexity and information volume of the image; s represents the original size of the image, indicating the visual data size of the visual data captured by the vehicle set; represents the number of pictures to be processed generated by the machine vision device at the same moment.
[0109] Based on this state information, DQN it can select the optimal compression ratio from the compression ratio set (such as 1.0、0.9、0.8 etc.), ensuring that the visual pictures automatically optimize the compression ratio according to the different real-time working intensities.
[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical method of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify the technical method of the present invention or make equivalent replacements, and these modifications or equivalent replacements cannot make the modified technical method deviate from the spirit and scope of the technical method of the present invention.
Claims
1. A comprehensive quality-preserving compression method for multiple images based on entropy analysis, characterized in that: The following steps are involved: S1. Perform information entropy analysis on the images generated in the machine vision application to obtain the information entropy of the images; S2, combining the information entropy, image size, and machine vision equipment with reinforcement learning technology, and obtaining the corresponding DQN Network, machine vision equipment DQN The network is used to process the generated images, obtain the target compression ratio, and obtain the target compression scheme of the images in the machine vision application; The specific contents of S2 include DQN Status definition, DQN Action Definition and DQN Rewards definition; right DQN Status The specific contents of the definition are: Defining states for ,in e The information entropy of the image reflects the complexity and amount of information in the image. s Represents the original size of the image, Represents the number of images that need to be processed by the machine vision device at the same time; right DQN Action The specific contents of the definition are: For an original image, DQN Made to it for: ; in, for action; DQN Rewards The specific content is: For image processing tasks , define the image processing task From generating, calculating image information entropy, compressing to being used for specific applications YOLO The processing time is The processing result accuracy is ; Defining image processing tasks After the processing is completed, the total number of images processed by the specific application is ; Defining image processing tasks After being processed, DQN What you get reward for: ; in, To adjust the weights of image accuracy and image processing time in reward, For the order of pictures, For the The total number of images processed by a specific application after the images are processed.
2. The method for comprehensive quality-preserving compression of multiple images based on entropy analysis according to claim 1, characterized in that: The information entropy of the pictures generated in machine vision applications is analyzed and calculated, and the expression of the information entropy of the pictures is: ; in, is the information entropy of the image, Represents pixel pairs The probability of appearing in the image.
3. The method for comprehensive quality-preserving compression of multiple images based on entropy analysis according to claim 2, characterized in that: in, The calculation formula is as follows: ; Among them, the pixel pair middle represents the gray value of the center pixel in the sliding window, is the average gray value of the pixels in the sliding window excluding the central pixel, Is a pixel pair The frequency of occurrence of all pixels in the entire image, is the number of pixels, For a specific visual data processing task, it is equivalent to a picture. The i in represents the order of the pictures themselves.
4. The method for comprehensive quality-preserving compression of multiple images based on entropy analysis according to claim 1, characterized in that: DQN The network training steps are: Step 1. Initialization DQN of Q Network and Experience Replay Pool D ; Step 2: Start multiple rounds of training.
5. The method for comprehensive quality-preserving compression of multiple images based on entropy analysis according to claim 1, characterized in that: The specific contents of starting multiple rounds of training in step 2 are: Step 2.1, get the first state from the external environment, which is ,in Represents visual data processing tasks The information entropy of the picture in Represents visual data processing tasks The original size of the image in Indicates the total number of images that need to be processed when the machine vision device generates the image; Step 2.2 DQN have Choose an action randomly with probability The probability of choosing a value is actions, where Express The specific action with the largest value is recorded as ; Step 2.3 DQN Execute an action , get rewarded , and then get the new state from the outside world , and then Tuples are stored in the experience replay pool D , the new state obtained from the outside world Set as Current , check the experience replay pool at this time D ; If the experience replay pool is large enough, D Randomly select a small batch of transformations from , calculate for each transformation ,in is a discount factor that adjusts the weight of future rewards in DQN, and then uses Update the network for the loss Q ; If it is not the last round of training, reduce ε Repeat steps 2.1 to 2.3 for the values of
Citation Information
Patent Citations
Picture dynamic adaptive compression method based on reinforcement learning
CN110210548A
Machine learning model compression method, device and equipment
CN112508187A
Motor adaptive control method and system based on deep reinforcement learning
CN118508817A