Data Compression Method, Device, Equipment and Storage Medium for Chip
The chip data compression method addresses inefficiencies in traditional methods by analyzing data type and using content-aware algorithms with Huffman coding, achieving high compression efficiency and quality in embedded systems and mobile devices.
Patent Information
- Application Number
- CN202510253193.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-03-05
AI Technical Summary
Traditional data compression methods are not efficient in embedded systems and mobile devices, and cannot meet the high requirements for data quality and real-time performance such as image processing and video streaming. The existing solution algorithm selection is not intelligent enough, the compression ratio is not high, and the decoding speed is slow.
By performing type analysis of compressed data, combining content-aware algorithms and Huffman encoding, selecting appropriate data compression algorithms, data type identification and key information retention, and optimizing encoding strategies to improve compression rate and data quality.
It realizes that while ensuring data quality, data compression efficiency is improved, data compatibility and portability are enhanced, data exchange and sharing are adapted to data exchange and sharing in different systems, and system performance and user experience are improved.
Smart Images

Figure CN119788088B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data compression, and particularly to a data compression method, apparatus, device and storage medium for a chip. Background Art
[0002] With the rapid development of information technology, the amount of data has increased explosively, which poses higher requirements for data storage and transmission efficiency. Especially in embedded systems and mobile devices, due to more obvious resource limitations, efficient data compression technology has become one of the key factors to improve system performance. However, traditional data compression methods often ignore the characteristics of the data itself and the application scenarios, resulting in low compression efficiency or a decrease in the quality of the decompressed data, and unable to meet the high requirements for data quality and real-time performance in specific fields such as image processing and video stream transmission.
[0003] In addition, in the process of chip design and application, different types of data (such as text, images, audio) have different structural characteristics, which requires the data compression method to be able to flexibly select the most suitable compression algorithm according to the data type. Although some existing data compression solutions provide a certain degree of flexibility, there are still problems such as insufficiently intelligent algorithm selection, low compression ratio, and slow decoding speed in actual applications. Especially for data with high content complexity, how to achieve efficient compression while ensuring data quality is a difficult point in current research.
[0004] To address these problems, this study proposes a new chip data compression method, aiming to improve the efficiency and quality of data compression by combining data type analysis, content-aware technology, and efficient coding strategies. This method first deeply analyzes the data to be compressed of the target chip to understand the specific type and characteristics of the data; then uses content-aware algorithms to further identify the key information in the data to ensure that this information is retained during the compression process; finally, uses an optimized data compression algorithm after selection for Huffman coding, which not only improves the compression ratio but also ensures the quality of the decompressed data. This method provides a new idea for solving the problems existing in the current data compression technology and helps to promote the development and application of related technologies. Summary of the Invention
[0005] The main object of the present invention is to provide a data compression method, apparatus, device and storage medium for a chip, which solves the technical problem of low traditional compression efficiency or a decrease in the quality of the decompressed data.
[0006] To achieve the above object, the present invention provides a data compression method for a chip, including the following steps:
[0007] Obtain the data to be compressed of the target chip, and perform type analysis on the data to be compressed to obtain the type of the data to be compressed;
[0008] Based on the type of the data to be compressed, select a data compression algorithm for the data to be compressed to obtain the selected data compression algorithm;
[0009] Perform content perception recognition on the data to be compressed through a preset content perception algorithm to obtain the perceived content recognition data;
[0010] Perform Huffman coding on the perceived content recognition data and the unperceived content recognition data in the data to be compressed through the selected data compression algorithm to obtain the compressed coding data;
[0011] Perform format conversion on the compressed coding data to obtain the standardized coding data, and store the standardized coding data in the target chip.
[0012] Further, the obtaining the data to be compressed of the target chip, and performing type analysis on the data to be compressed to obtain the type of the data to be compressed includes:
[0013] Perform a storage area scan on the target chip through a preset scan algorithm to obtain the data to be compressed;
[0014] Extract features from the data to be compressed to obtain the extracted scan features;
[0015] Use the Shannon entropy formula to calculate the information entropy of the extracted scan features respectively, and arrange all the information entropy into an information entropy vector;
[0016] Input the information entropy vector into a preset neural network model for classification to obtain a classification label;
[0017] Determine the type of the data to be compressed based on the classification label; wherein, the type of the data to be compressed includes message data, image data, and audio data.
[0018] Further, the using the Shannon entropy formula to calculate the information entropy of the extracted scan features respectively, and arranging all the information entropy into an information entropy vector includes:
[0019] Perform discretization processing on the extracted scan features to convert the continuous extracted scan features into discrete value features to obtain the discretized scan features;
[0020] Use a parallel computing architecture to calculate the probability of occurrence of each discrete value in the discretized scan features to obtain the feature probability distribution;
[0021] Use the Shannon entropy formula to calculate the information entropy of each discrete value based on the feature probability distribution;
[0022] The information entropy is obtained using the following formula;
[0023] ;
[0024] is the information entropy, is the total number of discrete values, is the discretized scan feature the probability of the th discrete value, is the logarithm to the base 2;
[0025] The index marking algorithm is used to assign a unique index to each of the information entropies to obtain the indexed information entropy values;
[0026] All the indexed information entropy values are arranged according to the index order in the index mapping table to obtain the information entropy vector.
[0027] Further, based on the data type to be compressed, a data compression algorithm is selected for the data to be compressed to obtain the selected data compression algorithm, including:
[0028] A preset algorithm library is called, and a list of candidate compression algorithms corresponding to the data type to be compressed is determined within the algorithm library based on the data type to be compressed;
[0029] The information entropy vector is dimensionally decomposed to obtain vector dimension information; wherein, the vector dimension information includes vector dimension correlation, vector dimension distribution, vector dimension weight, vector dimension sparsity, and vector dimension size;
[0030] The wavelet transform algorithm is used to perform frequency domain analysis on each of the vector dimension information to obtain frequency domain feature information;
[0031] Based on the frequency domain feature information, the information entropy vector is elementarily analyzed to obtain the analyzed vector element features;
[0032] Based on the analyzed vector element features, the performance of the candidate compression algorithms in the list of candidate compression algorithms is predicted to obtain a performance prediction probability distribution;
[0033] Based on the performance prediction probability distribution, the candidate compression algorithms are prioritized and selected to obtain the selected data compression algorithm.
[0034] Further, the content of the data to be compressed is sensed and identified through a preset content awareness algorithm to obtain the sensed content identification data, including:
[0035] Perceptually identify the data to be compressed through a preset content perception algorithm to obtain the identified perceptual content;
[0036] Mark the identified perceptual content to obtain the marked identified perceptual content;
[0037] Chunk the marked identified perceptual content to obtain content data chunks;
[0038] Perform hash transformation on the content data chunks to obtain transformed hash values;
[0039] Detect whether there are duplicate transformed hash values. If so, delete the identified perceptual content corresponding to the duplicate transformed hash values to obtain the deleted marked identified perceptual content;
[0040] Perform encryption processing on the deleted marked identified perceptual content to obtain encrypted perceptual content data, and use the encrypted perceptual content data as perceptual content identification data.
[0041] Further, the Huffman encoding of the perceptual content identification data and the unperceived content identification data in the to-be-compressed data by the selected data compression algorithm to obtain compressed encoded data includes:
[0042] Use the sliding window technique to analyze and record the occurrence frequency of each symbol in the unperceived content identification data to obtain a frequency statistical table; wherein, the frequency statistical table is used to reflect the distribution and frequency of each symbol;
[0043] Obtain the symbol frequency statistical result based on the frequency statistical table;
[0044] Construct a Huffman tree based on the frequency statistical result and the frequency statistical table;
[0045] Assign a unique binary codeword to each node in the constructed Huffman tree through a preset recursive encoding algorithm to obtain an encoding mapping table;
[0046] Perform Huffman encoding conversion on the symbols in the unperceived content identification data through the encoding mapping table to generate a preliminary compressed data stream;
[0047] Merge the perceptual content identification data with the preliminary compressed data stream to form compressed encoded data.
[0048] Further, the format conversion of the compressed encoded data to obtain standardized encoded data and store the standardized encoded data in the target chip includes:
[0049] Parse the compressed encoded data through a preset data parsing algorithm to obtain a compressed encoded data structure; wherein, the compressed encoded data structure includes a data header, a data body, and a check bit;
[0050] Perform format mapping on the compressed encoded data based on the compressed encoded data structure to obtain mapped format data, and use the mapped format data as standardized encoded data;
[0051] Perform encapsulation processing on the mapped format data through predefined data encapsulation rules to obtain an encapsulated data packet;
[0052] Perform encryption processing on the encapsulated data packet using a data encryption algorithm to obtain an encrypted data packet;
[0053] Perform storage path planning on the encrypted data packet based on a preset data storage protocol to obtain storage path information;
[0054] Write the encrypted data packet to a specified location of the target chip through a specific data writing interface.
[0055] The present invention also provides a data compression device for a chip, including:
[0056] An acquisition module, configured to acquire the data to be compressed of the target chip, and perform type analysis on the data to be compressed to obtain the type of the data to be compressed;
[0057] A selection module, configured to select a data compression algorithm for the data to be compressed based on the type of the data to be compressed to obtain a selected data compression algorithm;
[0058] An identification module, configured to perform content awareness identification on the data to be compressed through a preset content awareness algorithm to obtain perception content identification data;
[0059] An encoding module, configured to perform Huffman encoding on the perception content identification data and the unperceived content identification data in the data to be compressed through the selected data compression algorithm to obtain compressed encoded data;
[0060] A conversion module, configured to perform format conversion on the compressed encoded data to obtain standardized encoded data, and store the standardized encoded data in the target chip.
[0061] The present invention also provides a computer device, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the steps of the method described in any one of the above are implemented.
[0062] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.
[0063] The data compression method of the chip provided by the present invention includes the following steps: obtaining the data to be compressed of the target chip, and performing type analysis on the data to be compressed to obtain the type of the data to be compressed; based on the type of the data to be compressed, selecting a data compression algorithm for the data to be compressed to obtain the selected data compression algorithm; performing content awareness recognition on the data to be compressed through a preset content awareness algorithm to obtain the perceived content recognition data; performing Huffman coding on the perceived content recognition data and the unperceived content recognition data in the data to be compressed through the selected data compression algorithm to obtain the compressed coding data; performing format conversion on the compressed coding data to obtain the standardized coding data, and storing the standardized coding data in the target chip. Through the above technical means, the technical problem of low traditional compression efficiency or the decline of the data quality after decompression is solved, and the technical effect of generating the standardized coding data and storing it in the target chip by performing format conversion on the compressed coded data is realized. This process not only facilitates the unified management and rapid retrieval of data, but also facilitates the data exchange and sharing between different systems, and improves the compatibility and portability of data. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 is a schematic diagram of the steps of the data compression method of the chip in an embodiment of the present invention;
[0065] Figure 2 is a structural block diagram of the data compression device of the chip in an embodiment of the present invention;
[0066] Figure 3 is a schematic structural diagram of a computer device in an embodiment of the present invention.
[0067] The implementation, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0068] In order to make the object, technical solution and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0069] As Figure 1 shown, Figure 1 is a schematic diagram of the steps of a data compression method of a chip in an embodiment of the present invention;
[0070] In an embodiment of the present invention, a data compression method for a chip is provided, including the following steps:
[0071] Step S1: Obtain the data to be compressed of the target chip, and perform type analysis on the data to be compressed to obtain the type of the data to be compressed.
[0072] Specifically, the step of obtaining the data to be compressed of the target chip and performing type analysis on the data to be compressed to obtain the type of the data to be compressed means that before starting the data compression process, the data that needs to be compressed is first extracted from the target chip. Then, through a series of preset analysis tools or algorithms, in-depth type analysis is performed on these data to be compressed to determine their specific properties and categories. For example, in the application scenario of image processing, if the data to be compressed is a set of high-definition image files, the type analysis process includes identifying the format of these image files (such as JPEG, PNG), resolution size, color depth and other characteristics. Through such analysis, the system can understand whether these image files contain a large amount of color information, whether there are repeated texture patterns or a simple background, etc., thereby providing a basis for the selection of the subsequent data compression algorithm. Similarly, in the application of video stream transmission, the data to be compressed may be a continuous sequence of video frames, and the type analysis will focus on factors such as the video coding format (such as H.264, HEVC), frame rate, resolution, as well as the dynamic range and motion complexity of the video content, etc., so as to be able to more accurately select a compression method suitable for this type of video data in the future. This pre-analysis based on data type not only helps to improve the efficiency of subsequent compression processing, but also can guarantee the quality of the compressed data to a certain extent and meet the specific requirements in different application scenarios.
[0073] Step S2: Based on the type of the data to be compressed, select a data compression algorithm for the data to be compressed to obtain the selected data compression algorithm.
[0074] Specifically, based on the type of data to be compressed, a data compression algorithm is selected for the data to be compressed, and the selected data compression algorithm is obtained. This means that after completing the analysis of the data type, the most suitable compression algorithm is selected according to the analysis results. This process not only depends on the basic understanding of the data type, but also needs to consider the specific characteristics of the data and the requirements of the application scenario. For example, in the application scenario of image processing, if the data to be compressed mainly consists of static images with high resolution and rich color levels, then compression algorithms such as JPEG or WebP, which focus on maintaining image quality while reducing file size, can be selected. Both of these algorithms can effectively remove redundant information in the image, such as spatial redundancy and visual redundancy, thus achieving efficient compression without significantly affecting the image quality. On the contrary, if the data to be compressed are some icons or graphic files with a large number of solid color areas, it may be more appropriate to choose the lossless compression algorithm PNG at this time, because PNG can well preserve the details and transparency information of these images. Similarly, in the application scenario of video stream transmission, the type analysis of the data to be compressed may reveal the characteristics of the video content, such as whether it contains a large number of fast-action scenes, high-contrast picture changes, etc. Based on these analysis results, advanced video compression standards such as H.264 or HEVC can be selected. H.264 is widely used in network video streaming transmission and can provide good compression efficiency and compatibility; while HEVC is the next-generation standard of H.264 and provides a higher compression ratio, especially suitable for video content with 4K and higher resolutions. Both of these algorithms can effectively reduce the size of the video file while minimizing the impact on the video quality, which is very suitable for application scenarios that require efficient transmission of a large amount of video data. In short, by deeply analyzing the type of data to be compressed and selecting the most suitable data compression algorithm accordingly, it can be ensured that the compression process is both efficient and able to meet the quality requirements under specific application scenarios. This intelligent selection mechanism not only improves the overall performance of data compression, but also lays a good foundation for subsequent data processing and transmission.
[0075] Step S3, perform content-aware recognition on the data to be compressed through a preset content-aware algorithm to obtain the recognized content-aware data.
[0076] Specifically, the content of the data to be compressed is sensed and recognized through a preset content perception algorithm to obtain the perception content recognition data. This step is to further refine and optimize the data compression process. After completing the data type analysis and selecting the appropriate compression algorithm, the introduction of the content perception algorithm can help the system more accurately understand the information distribution and importance within the data, so as to take more targeted measures during the compression process. For example, in the application scenario of image processing, the content perception algorithm can analyze features such as edges, textures, and color distributions in the image to identify which parts are more easily noticed by the human eye, such as faces, text, or specific objects. In this way, during the subsequent compression process, the information in these key areas can be preferentially protected to avoid distortion or blurring caused by excessive compression, ensuring that the decompressed image still maintains a high visual quality. Similarly, in the application of video stream transmission, the content perception algorithm can identify elements such as the main objects, motion trajectories, and background changes in the video. Through these analyses, the system can distinguish which parts of the video are dynamic and frequently changing, and which parts are relatively static. For the dynamic parts, the algorithm can allocate more bitrates to ensure the smoothness and clarity of the motion; while for the static or less-changing parts, the bitrate can be appropriately reduced to reduce unnecessary data redundancy. In this way, not only can the efficiency of video compression be improved, but also the fluency and visual effects during video playback can be ensured. In addition, the application of the content perception algorithm can also help discover repeated patterns or redundant information in the data. In image processing, certain textures or patterns may repeat at different positions in the image. By identifying these repeated parts, more effective coding strategies can be adopted during the compression process to reduce the number of bits required to store the same information. In video processing, there may be a large amount of similar content between adjacent frames. The content perception algorithm can identify these similarities and use predictive coding techniques to reduce inter-frame redundancy, further improving the compression efficiency. In summary, by performing content perception and recognition on the data to be compressed through a preset content perception algorithm, not only can the system's understanding of the internal structure of the data be deepened, but also the subsequent data compression process can be guided to achieve a more refined and efficient compression effect. This content perception-based data processing method provides strong technical support for the efficient storage and transmission of multimedia data such as images and videos.
[0077] Step S4, perform Huffman coding on the perception content recognition data and the unperceived content recognition data in the to-be-compressed data through the selected data compression algorithm to obtain compressed coding data.
[0078] Specifically, the selected data compression algorithm performs Huffman encoding on the perceived content recognition data and the unperceived content recognition data in the data to be compressed, obtaining compressed encoded data. This process is a crucial step after content perception and recognition, aiming to further improve the efficiency and quality of data compression. After completing the content perception analysis, the system can already identify the key information and non-key information in the data. Next, an efficient encoding method needs to be adopted to achieve data compression. Here, Huffman encoding is selected. This is a lossless compression algorithm based on probability statistics, which can assign different encoding lengths to symbols according to the frequencies of their occurrences in the data, such that symbols with high occurrence frequencies are represented by shorter encodings, while symbols with low occurrence frequencies are represented by longer encodings, thus achieving the purpose of compression. For example, in the application scenario of image processing, assuming that the content perception algorithm has identified the face region in the image as key information, this part of the data will be marked as content that needs to be particularly protected. For this part of the perceived content recognition data, the system will use the selected data compression algorithm and combine it with Huffman encoding for processing. Since the face region usually contains more details and variations, Huffman encoding can assign the optimal encoding length according to the frequencies of each pixel value in these details, thereby achieving efficient compression while maintaining high quality. At the same time, for other regions in the image that are not perceived as key information, such as the background or unimportant textures, the system will also use Huffman encoding, but may allow a larger compression ratio because these regions have less impact on the overall visual effect. In the application scenario of video stream transmission, the content perception algorithm may have identified the main moving objects and background regions in the video. For this part of the perceived content recognition data, Huffman encoding can assign the optimal encoding according to the change frequency of the pixel values of the moving objects, ensuring that these objects remain clear and smooth in the compressed video. For the background region or other parts with little change and less impact on the viewing experience, Huffman encoding can adopt a higher compression ratio to reduce the storage space occupied by these parts. This differential encoding strategy can not only improve the overall compression efficiency of the video but also ensure the integrity of key information, providing a better viewing experience for users. In this way, the selected data compression algorithm combined with Huffman encoding processes the perceived content recognition data and the unperceived content recognition data in the data to be compressed, not only achieving efficient data compression but also ensuring the quality of the compressed data. The application of this technology provides strong support for the efficient storage and transmission of multimedia data such as images and videos, especially in resource-constrained embedded systems and mobile devices, which can significantly improve the performance of the system and the user experience.
[0079] Step S5: Convert the format of the compressed encoded data to obtain standardized encoded data, and store the standardized encoded data in the target chip.
[0080] Specifically, the compressed encoded data is format-converted to obtain standardized encoded data, and the standardized encoded data is stored in the target chip. This is the last important step in the data compression process, aiming to ensure that the compressed data can be stored in a standardized form for subsequent management and use. After Huffman encoding, although the compressed encoded data has achieved a high compression efficiency, these data usually exist in a specific encoding format and may not be directly applicable to all systems or applications. Therefore, through format conversion, these compressed encoded data need to be converted into a standardized encoding format that can be widely accepted and supported. For example, in the application scenario of image processing, the image data after Huffman encoding may exist in an internal proprietary binary format. To enable this data to be smoothly used on various image processing software and hardware platforms, it needs to be converted into a common standardized format, such as JPEG, PNG, or WebP. During this format conversion process, the system will make necessary adjustments and reorganizations to the compressed encoded data according to the requirements of the target format. For example, for the JPEG format, the system needs to add corresponding file header information, define parameters such as the width, height, and color space of the image, and organize data blocks according to the JPEG standard format. In this way, the converted standardized encoded data can be directly read and displayed by various devices and software that support the JPEG format. Similarly, in the application scenario of video stream transmission, the video data after Huffman encoding also needs to be format-converted to adapt to different players and transmission protocols. Assuming the target format is H.264, then the system needs to reorganize the compressed encoded data, add necessary encoding parameters and timestamp information to ensure that each video frame meets the requirements of the H.264 standard. In addition, it is necessary to generate file headers and index information that conform to the ISO Base Media File Format (such as the MP4 file format) for the smooth playback and random access of video files. Through this standardized format conversion, it can be ensured that the compressed video data can be seamlessly played on various devices and platforms, whether it is a smartphone, a tablet, or a smart TV. After obtaining the standardized encoded data through format conversion, the next step is to store this data in the target chip. The target chip can be the internal memory of an embedded system, an SD card, or other types of external storage devices. The storage process needs to ensure the integrity and security of the data to prevent errors or damage during transmission. For example, in an embedded system, the system can write the standardized encoded data into the internal flash memory or an external SD card through interfaces such as SPI, I2C, or SDIO. During the writing process, the system will perform checksums and error detection to ensure that each piece of data is correctly stored.Generally speaking, by converting the format of compressed encoded data to obtain standardized encoded data and storing these data in the target chip, it can not only ensure the compatibility and portability of the data, but also improve the convenience and reliability of data management. This standardized data processing method provides a solid foundation for the efficient storage and transmission of multimedia data such as images and videos. Especially in resource-constrained embedded systems and mobile devices, it can significantly improve the overall performance of the system and the user experience.
[0081] In a specific embodiment, the obtaining of the data to be compressed in the target chip and the type analysis of the data to be compressed to obtain the data type to be compressed include:
[0082] Scanning the storage area of the target chip through a preset scanning algorithm to obtain the data to be compressed;
[0083] Extracting features from the data to be compressed to obtain the extracted scanning features;
[0084] Calculating the information entropy of the extracted scanning features respectively using the Shannon entropy formula, and arranging all the information entropy into an information entropy vector;
[0085] Inputting the information entropy vector into a preset neural network model for classification to obtain a classification label;
[0086] Determining the data type to be compressed based on the classification label; wherein, the data type to be compressed includes message data, image data and audio data.
[0087] Specifically, first, a preset scanning algorithm is used to scan the storage area of the target chip to obtain the data to be compressed. This step is to extract the data that needs to be compressed from the target chip. The scanning algorithm will traverse the storage area of the chip to identify which data needs to be compressed. For example, in an embedded system, this data may be stored in internal flash memory or an external SD card. The scanning algorithm will read each data block in these storage areas one by one to ensure that all the data to be compressed is extracted completely. Next, feature extraction is performed on the data to be compressed to obtain the extracted scanning features. Feature extraction is a key step in data type analysis. By analyzing the data in detail, key information that can reflect the characteristics of the data is extracted. These features may include the data format, structure, frequency distribution, etc. For example, in image data, feature extraction may involve color histograms, edge detection, texture analysis, etc.; in audio data, it may involve spectral analysis, sound intensity distribution, etc. Through these feature extractions, the system can obtain detailed information about the data, laying a foundation for further analysis and classification. Subsequently, the information entropy of the extracted scanning features is calculated using the Shannon entropy formula respectively, and all the information entropies are arranged into an information entropy vector. Shannon entropy is an indicator that measures the uncertainty of data. By calculating the information entropy of each feature, the randomness and complexity of these features can be quantified. By calculating the information entropy of each extracted feature, the system can obtain a set of values that reflect the complexity and uncertainty of the data. Arranging all the information entropy values into a vector forms the information entropy vector, providing input data for subsequent classification. Then, the information entropy vector is input into a preset neural network model for classification to obtain a classification label. The neural network model is a powerful machine learning tool that can learn the internal laws and patterns of the data through training. In this step, the information entropy vector is fed into a pre-trained neural network model as input data. The neural network model extracts high-level features of the data step by step through the calculations of multiple layers of neurons and finally outputs a classification label. This classification label represents the type of the data to be compressed, such as packet data, image data, or audio data. Finally, the type of the data to be compressed is determined based on the classification label. The classification label is the result output by the neural network model. By parsing this label, the system can clearly know the specific type of the data to be compressed. For example, if the classification label indicates that the data type is image data, then the system will select a compression algorithm suitable for image data, such as JPEG or WebP, in the subsequent compression step; if the classification label indicates that the data type is audio data, then the system will select a compression algorithm suitable for audio data, such as MP3 or AAC. Through this series of steps, the system can accurately identify and classify the type of the data to be compressed, providing an important foundation for subsequent data compression. This method of data type analysis based on feature extraction and machine learning not only improves the efficiency and quality of data compression but also can adapt to the characteristics and requirements of different types of data.For example, in the application scenario of image processing, after the system identifies the image data through the above steps, it can select efficient image compression algorithms such as JPEG or WebP to ensure the quality of the compressed image; in the application scenario of video stream transmission, after the system identifies the video data, it can select advanced video compression standards such as H.264 or HEVC to ensure the efficient transmission and playback quality of the video. This intelligent data processing method provides strong technical support for the efficient storage and transmission of multimedia data.
[0088] In a specific embodiment, the method of calculating the information entropy of the extracted scan features respectively by using the Shannon entropy formula and arranging all the information entropy into an information entropy vector includes:
[0089] Perform discretization processing on the extracted scan features to convert the continuous extracted scan features into discrete value features, and obtain discretized scan features;
[0090] Use a parallel computing architecture to calculate the probability of occurrence of each discrete value in the discretized scan features, and obtain a feature probability distribution;
[0091] Use the Shannon entropy formula to calculate the information entropy of each discrete value based on the feature probability distribution;
[0092] The following formula is used to obtain the information entropy;
[0093] ;
[0094] is the information entropy, is the total number of discrete values, is the discretized scan feature of the - th discrete value probability, is the logarithm to the base 2;
[0095] Use an index marking algorithm to assign a unique index to each information entropy to obtain an information entropy value with an index;
[0096] Arrange all the information entropy values with indexes according to the index order in the index mapping table to obtain an information entropy vector.
[0097] Specifically, first, the extracted scan features are discretized to convert the continuous extracted scan features into discrete value features, obtaining discretized scan features. Discretization is to divide continuous data values into several intervals, and each interval corresponds to a discrete value. For example, in the application scenario of image processing, the extracted features may be the gray values of pixels, and these gray values are usually continuous integers between 0 and 255. To perform discretization, these gray values can be divided into several intervals, such as 0 - 63, 64 - 127, 128 - 191, 192 - 255, and each interval corresponds to a discrete value. Through this processing, the continuous gray values are converted into discrete interval values, which is convenient for subsequent calculation and analysis.
[0098] Next, using a parallel computing architecture, the probability of each discrete value in the discretized scan features is calculated to obtain a feature probability distribution. The parallel computing architecture can significantly improve the computing efficiency, especially when dealing with large-scale data. For example, a GPU or a multi-core CPU can be used for parallel computing. Each computing unit is responsible for calculating the occurrence probability of a part of the discrete values. By counting the number of occurrences of each discrete value in all data and dividing by the total number of samples, the probability of each discrete value can be obtained. For example, in image data, if a certain gray value interval appears 1000 times among all pixels and the total number of pixels is 10000, then the probability of this interval is 0.1.
[0099] Then, using the Shannon entropy formula , based on the feature probability distribution, the information entropy of each discrete value is calculated. The Shannon entropy formula is used to measure the uncertainty or randomness of data. For each discrete value , its information entropy is calculated . Specifically, the probability of each discrete value is substituted into the formula to calculate the corresponding information entropy. For example, if the probability of a certain discrete value is 0.1, then its information entropy is 0.332.
[0100] Next, an index marking algorithm is used to assign a unique index to each information entropy, obtaining indexed information entropy values. The index marking algorithm is used to assign a unique identifier to each information entropy value for subsequent sorting and processing. For example, an integer index, such as 1, 2, 3, etc., can be assigned to each information entropy value in sequence. In this way, each information entropy value has a corresponding index, forming indexed information entropy values.
[0101] Finally, arrange all the indexed information entropy values according to the index order in the index mapping table to obtain an information entropy vector. The index mapping table is a table that records the correspondence between indexes and information entropy values. Through this table, all the indexed information entropy values can be arranged according to the index order to form an ordered information entropy vector. For example, if the information entropy values and their indexes are (0.332, 1), (0.5, 2), (0.1, 3) respectively, the information entropy vector after arranging according to the index order is [0.332, 0.5, 0.1].
[0102] Through this series of steps, the system can convert the extracted scan features into information entropy vectors. These information entropy vectors contain the uncertainty information of the data and can be used as input data to be sent into the neural network model for classification. For example, in the application scenario of image processing, the information entropy vector can reflect the complexity and randomness of the image and help the system identify the image data; in the application scenario of video stream transmission, the information entropy vector can reflect the change of video frames and help the system identify the video data. This feature extraction and classification method based on information entropy not only improves the accuracy of data type recognition but also provides an important reference basis for subsequent data compression. In this way, the system can more intelligently select the appropriate compression algorithm and improve the efficiency and quality of data compression.
[0103] In a specific embodiment, the selection of the data compression algorithm for the data to be compressed based on the data type to be compressed to obtain the selected data compression algorithm includes:
[0104] Call a preset algorithm library, and determine a list of candidate compression algorithms corresponding to the data type to be compressed within the algorithm library based on the data type to be compressed;
[0105] Perform dimensional decomposition on the information entropy vector to obtain vector dimension information; wherein, the vector dimension information includes vector dimension correlation, vector dimension distribution, vector dimension weight, vector dimension sparsity, and vector dimension size;
[0106] Use the wavelet transform algorithm to perform frequency-domain analysis on each piece of vector dimension information to obtain frequency-domain feature information;
[0107] Perform element analysis on the information entropy vector based on the frequency-domain feature information to obtain the characteristics of the parsed vector elements;
[0108] Perform performance prediction on the candidate compression algorithms in the list of candidate compression algorithms based on the characteristics of the parsed vector elements to obtain a performance prediction probability distribution;
[0109] Rank and select the candidate compression algorithms based on the performance prediction probability distribution to obtain the selected data compression algorithm.
[0110] Specifically, first, call the preset algorithm library, and determine a list of candidate compression algorithms corresponding to the data type to be compressed within the algorithm library. The algorithm library is a collection containing various data compression algorithms, and these algorithms are applicable to different types of data. For example, for image data, the algorithm library may include compression algorithms such as JPEG, PNG, and WebP; for audio data, it may include MP3, AAC, etc.; for video data, it may include H.264, HEVC, etc. The system filters out applicable candidate compression algorithms from the algorithm library according to the data type to be compressed determined in the previous step to form a list of candidate compression algorithms. For example, if the data type to be compressed is image data, the system will select algorithms such as JPEG, PNG, and WebP from the algorithm library as candidates. Next, perform dimensional decomposition on the information entropy vector to obtain vector dimension information. The information entropy vector contains information in multiple dimensions, and these dimensions reflect the complexity and uncertainty of the data. The purpose of dimensional decomposition is to extract the characteristics of these dimensions for subsequent analysis and processing. The vector dimension information includes vector dimension correlation, vector dimension distribution, vector dimension weight, vector dimension sparsity, and vector dimension size. Vector dimension correlation reflects the degree of association between different dimensions; vector dimension distribution describes the distribution of each dimension value; vector dimension weight represents the importance of each dimension; vector dimension sparsity reflects the proportion of zero values in the vector; vector dimension size represents the length of the vector. For example, in image data, the dimensions of the information entropy vector may include color distribution, texture complexity, edge strength, etc., and these characteristics can be extracted through dimensional decomposition. Then, use the wavelet transform algorithm to perform frequency domain analysis on each piece of vector dimension information to obtain frequency domain feature information. The wavelet transform is a multi-resolution analysis method that can decompose a signal into components of different frequencies. By performing wavelet transform on the vector dimension information, the characteristics of each dimension at different frequencies can be extracted. The frequency domain feature information reflects the distribution of data at different frequencies and helps to identify periodic features and local changes in the data. For example, in image data, high-frequency details (such as edges and textures) and low-frequency backgrounds (such as smooth regions) in the image can be identified through wavelet transform. Next, perform element parsing on the information entropy vector based on the frequency domain feature information to obtain parsed vector element characteristics. The parsed vector element characteristics are the detailed analysis results of each element in the information entropy vector, and these characteristics describe the specific performance of the data in different dimensions. For example, the parsed vector element characteristics may include statistical information such as the maximum value, minimum value, average value, and variance of each dimension, as well as the correlation coefficients between dimensions. These characteristics provide basic data for subsequent performance prediction. Then, perform performance prediction on the candidate compression algorithms in the list of candidate compression algorithms based on the parsed vector element characteristics to obtain a performance prediction probability distribution.Performance prediction is to predict the performance of each candidate compression algorithm on the current data by analyzing the characteristics of the parsed vector elements. Performance prediction can be carried out using machine learning models (such as decision trees, support vector machines, neural networks, etc.) based on historical data and known algorithm performance. The performance prediction probability distribution represents the probability that each candidate compression algorithm will perform excellently on the current data. For example, if the characteristics of the parsed vector elements of a certain image data indicate that the image has complex textures and edges, then the JPEG algorithm may find a better balance between the compression ratio and image quality, so its performance prediction probability is higher. Finally, the candidate compression algorithms are prioritized and selected based on the performance prediction probability distribution to obtain the selected data compression algorithm. The prioritization is to arrange the candidate compression algorithms in descending order according to the performance prediction probability distribution. The selected data compression algorithm is to select one or more algorithms with the best performance from the sorted candidate compression algorithms. For example, if the JPEG algorithm has the highest performance prediction probability, then the system will select JPEG as the final data compression algorithm. If further optimization is needed, several algorithms with the highest performance prediction probability can be selected for actual testing, and finally the algorithm with the best performance is selected. Through this series of steps, the system can intelligently select the most suitable compression algorithm based on the type and characteristics of the data to be compressed. This algorithm selection method based on data characteristics and performance prediction not only improves the efficiency of data compression but also ensures the quality of the compressed data. For example, in the application scenario of image processing, the system selects the JPEG algorithm through the above steps, which can achieve efficient compression while maintaining the image quality; in the application scenario of video stream transmission, the system selects the H.264 algorithm, which can ensure the efficient transmission and smooth playback of the video. This intelligent data processing method provides strong technical support for the efficient storage and transmission of multimedia data.
[0111] In a specific embodiment, the content-aware recognition of the data to be compressed by a preset content-aware algorithm to obtain the perception content recognition data includes:
[0112] Performing content-aware recognition on the data to be compressed by a preset content-aware algorithm to obtain the recognized perception content;
[0113] Marking the recognized perception content to obtain the marked recognized perception content;
[0114] Performing data chunking on the marked recognized perception content to obtain content data chunks;
[0115] Performing hash transformation on the content data chunks to obtain the transformed hash values;
[0116] Detect whether there is a duplicate of the converted hash value. If so, delete the recognition and perception content corresponding to the duplicate converted hash value, and obtain the recognition and perception content with the deletion mark.
[0117] Perform encryption processing on the recognition and perception content with the deletion mark to obtain encrypted perception content data, and use the encrypted perception content data as perception content recognition data.
[0118] Specifically, first, content-aware recognition is performed on the data to be compressed through a preset content-aware algorithm to obtain recognized and perceived content. The content-aware algorithm is a technology that can understand and recognize the content of data and can be used to extract key information and features in the data. For example, in the application scenario of image processing, the content-aware algorithm can identify key information such as human faces, text, and objects in the image by analyzing features such as edges, textures, and color distributions in the image. In the application scenario of video stream transmission, the content-aware algorithm can identify the main objects, motion trajectories, and background changes in the video. Through these recognitions, the system can understand the important parts in the data and provide a basis for subsequent processing. Next, the recognized and perceived content is marked to obtain the marked recognized and perceived content. The marking process is to label the recognized key information for subsequent processing and management. For example, in image data, the system can add bounding boxes or other marks around key information such as the human face area and text area, so that the subsequent compression algorithm can preferentially protect these areas. In video data, the system can add labels to each main object to record its position and motion trajectory. These marks not only help the compression algorithm process the data better, but also can be used for subsequent data retrieval and analysis. Then, the marked recognized and perceived content is divided into data blocks to obtain content data blocks. Data block division is to divide the data into multiple small blocks, and each small block contains a part of the marked recognized and perceived content. This division helps to improve the parallelism and efficiency of processing. For example, in image data, the image can be divided into multiple small blocks, and each small block contains a part of the human face area or text area; in video data, the video frame can be divided into multiple areas, and each area contains a main object or a background part. Through data block division, multiple small blocks can be processed in parallel to improve the processing speed. Next, the content data blocks are subjected to hash transformation to obtain transformed hash values. Hash transformation is to convert each data block into a hash value with a fixed length, and this hash value can uniquely identify the content of the data block. Hash transformation helps to detect the repeatability and uniqueness of data blocks. For example, hash algorithms such as MD5 and SHA-1 can be used to perform hash transformation on each data block. Through hash transformation, the system can quickly compare the hash values of different data blocks to detect whether there are duplicate data blocks. Then, it is detected whether there are duplicate transformed hash values. If so, the recognized and perceived content corresponding to the duplicate transformed hash values is deleted to obtain the marked recognized and perceived content after deletion. The deletion of duplicate data blocks is to reduce data redundancy and improve compression efficiency. For example, if the hash values of multiple data blocks are the same, it means that the contents of these data blocks are exactly the same, and the system can only retain one data block and delete the remaining duplicate data blocks. In this way, the storage space of the data can be significantly reduced and the compression efficiency can be improved. At the same time, the hash value and position information of one data block are retained for subsequent decompression and recovery.Finally, encrypt the recognized and perceived content after the deletion of the tags to obtain encrypted perceived content data, and use the encrypted perceived content data as the perceived content recognition data. The encryption process is to protect the security and privacy of the data. For example, encryption algorithms such as AES and RSA can be used to encrypt the data blocks after removing duplicates. The encrypted data blocks can only be decrypted with the correct key, thus preventing unauthorized access and tampering. Through the encryption process, the system can ensure that the compressed data will not be illegally accessed during storage and transmission, improving the security of the data. Through this series of steps, the system can identify and process the key information in the data through the content-aware algorithm, reduce data redundancy, improve the compression efficiency, and ensure the security and privacy of the data. For example, in the application scenario of image processing, the system identifies the face and text areas in the image through the content-aware algorithm, marks these areas, performs data chunking and hash transformation, deletes duplicate data blocks, and finally encrypts the remaining data blocks to ensure that the compressed image data is both efficient and secure. In the application scenario of video stream transmission, the system identifies the main objects in the video through a similar method, marks and processes these objects, reduces redundant data, improves the compression efficiency and transmission speed of the video, and protects the security of the video content. This content-aware data processing method provides strong technical support for the efficient storage and transmission of multimedia data.
[0119] In a specific embodiment, the Huffman encoding of the perceived content recognition data and the unperceived content recognition data in the to-be-compressed data through the selected data compression algorithm to obtain compressed encoded data includes:
[0120] Using the sliding window technique, analyze and record the occurrence frequency of each symbol in the unperceived content recognition data to obtain a frequency statistical table; wherein, the frequency statistical table is used to reflect the distribution and frequency of each symbol;
[0121] Obtain the symbol frequency statistical result based on the frequency statistical table;
[0122] Construct a Huffman tree based on the frequency statistical result and the frequency statistical table;
[0123] Assign a unique binary codeword to each node in the constructed Huffman tree through a preset recursive encoding algorithm to obtain an encoding mapping table;
[0124] Perform Huffman encoding conversion on the symbols in the unperceived content recognition data through the encoding mapping table to generate a preliminary compressed data stream;
[0125] Merge the perceived content recognition data with the preliminary compressed data stream to form compressed encoded data.
[0126] Specifically, first, using the sliding window technique, analyze and record the occurrence frequency of each symbol in the unperceived content recognition data to obtain a frequency statistical table. The sliding window technique is a commonly used data analysis method that analyzes the symbol frequencies in data by sliding within a fixed-size window. For example, in the application scenario of image processing, the unperceived content recognition data may include the background area or unimportant textures in the image. The system can set a sliding window, and the window size can be adjusted according to actual needs. As the window slides through the data, the number of occurrences of each symbol (such as pixel values) is recorded. In this way, the system can count the occurrence frequency of each symbol in the unperceived content recognition data and generate a frequency statistical table. The frequency statistical table details the distribution and frequency of each symbol, providing basic data for subsequent encoding. Next, obtain the symbol frequency statistical result based on the frequency statistical table. The symbol frequency statistical result is to summarize and organize the data in the frequency statistical table to form a clear symbol frequency distribution. For example, the frequency statistical table may record the number of occurrences of each pixel value in the image, and the symbol frequency statistical result will summarize which pixel values have the highest occurrence frequency and which have the lowest. These statistical results are helpful for subsequent Huffman tree construction and encoding optimization. Then, construct a Huffman tree based on the frequency statistical result and the frequency statistical table. A Huffman tree is a binary tree structure used to generate optimal prefix codes. The process of constructing a Huffman tree is as follows: First, take each symbol in the frequency statistical table and its frequency as a node to create an initial node set. Then, select the two nodes with the smallest frequencies, merge them into a new node, and the frequency of the new node is the sum of the frequencies of these two nodes. Add the new node to the set and remove the original two nodes from the set. Repeat this process until there is only one node left in the set, which is the root node of the Huffman tree. In this way, the system can construct a Huffman tree, and each leaf node in the tree represents a symbol, and the path length and direction determine the encoding of the symbol. Next, assign a unique binary codeword to each node in the constructed Huffman tree through a preset recursive encoding algorithm to obtain an encoding mapping table. The recursive encoding algorithm starts from the root node of the Huffman tree and traverses down along the path of the tree to assign binary codewords to each node. Usually, the left branch is assigned 0 and the right branch is assigned 1. When reaching a leaf node, the symbol corresponding to this node is assigned a unique binary codeword. The binary codewords of all leaf nodes are combined together to form an encoding mapping table. The encoding mapping table records the mapping relationship between each symbol and its corresponding binary codeword, providing a basis for subsequent encoding conversion. Then, perform Huffman encoding conversion on the symbols in the unperceived content recognition data through the encoding mapping table to generate a preliminary compressed data stream. The encoding conversion process is to replace each symbol in the unperceived content recognition data with its corresponding binary codeword.For example, in image data, if the binary codeword of a certain pixel value is 010, then during the encoding conversion process, this pixel value will be replaced with 010. In this way, the system can convert each symbol in the unperceived content recognition data into a shorter binary codeword, generating a preliminary compressed data stream. The preliminary compressed data stream is the data after Huffman encoding, which occupies much less space than the original data. Finally, the perceived content recognition data is merged with the preliminary compressed data stream to form the compressed encoded data. The perceived content recognition data is the data recognized and processed by the content perception algorithm, and these data usually contain key information that needs to be protected preferentially during the compression process. The preliminary compressed data stream is the result of Huffman encoding of the unperceived content recognition data. Combining these two parts of data forms the final compressed encoded data. The process of data merging can be simple concatenation, or more complex interleaving or embedding methods, depending on the requirements of the application scenario. For example, in image processing, the perceived content recognition data (such as the face area) can be placed at the beginning of the compressed encoded data to ensure that these key information are preferentially restored during the decompression process. Through this series of steps, the system can efficiently compress data through Huffman encoding technology, ensuring that the compressed data is both efficient and reliable. For example, in the application scenario of image processing, the system analyzes the symbol frequency of the background area through the sliding window technology, constructs the Huffman tree and generates the encoding mapping table, performs Huffman encoding on the background area, and generates the preliminary compressed data stream. Then, the face area processed by the content perception algorithm is merged with the preliminary compressed data stream to form the final compressed encoded data. In this way, the compressed image data not only occupies less space, but also the key information is effectively protected. In the application scenario of video stream transmission, the system can also use a similar method to efficiently compress the background area in the video, ensuring the clarity and smoothness of the main object, and improving the efficiency and quality of video transmission. This data compression method based on Huffman encoding provides strong technical support for the efficient storage and transmission of multimedia data.
[0127] In a specific embodiment, the step of performing format conversion on the compressed encoded data to obtain standardized encoded data and storing the standardized encoded data in the target chip includes:
[0128] Parsing the compressed encoded data through a preset data parsing algorithm to obtain a compressed encoded data structure; wherein, the compressed encoded data structure includes a data header, a data body, and a check bit;
[0129] Performing format mapping on the compressed encoded data based on the compressed encoded data structure to obtain mapped format data, and using the mapped format data as the standardized encoded data;
[0130] Encapsulate the mapped format data according to predefined data encapsulation rules to obtain an encapsulated data packet;
[0131] Use a data encryption algorithm to encrypt the encapsulated data packet to obtain an encrypted data packet;
[0132] Based on a preset data storage protocol, plan a storage path for the encrypted data packet to obtain storage path information;
[0133] Write the encrypted data packet to a specified location of the target chip through a specific data writing interface.
[0134] Specifically, first, the compressed and encoded data is parsed through a preset data parsing algorithm to obtain a compressed and encoded data structure. The data parsing algorithm is used to analyze the internal structure of the compressed and encoded data and decompose it into different components. The compressed and encoded data structure usually includes a data header, a data body, and a check bit. The data header contains basic information about the data, such as data type, length, version number, etc.; the data body is the actual data content; the check bit is used to verify the integrity and correctness of the data. For example, in the application scenario of image processing, the compressed and encoded data may be a JPEG file. The data parsing algorithm will parse out information such as the identifier and data length in the file header. The data body is the image data after Huffman encoding, and the check bit is used to ensure that the data is not damaged during transmission. Next, based on the compressed and encoded data structure, format mapping is performed on the compressed and encoded data to obtain mapped format data, and the mapped format data is used as standardized encoded data. Format mapping is to convert the parsed data structure into a data format that conforms to a specific standard. For example, if the target format is JPEG, the system will reorganize the information in the data header according to the requirements of the JPEG standard format to ensure that each field conforms to the standard. The data body part will also be rearranged and encoded according to the standard format. In this way, the system can convert the compressed and encoded data into a standardized format so that it can be recognized and processed by various devices and software that support this format. Then, through predefined data encapsulation rules, encapsulation processing is performed on the mapped format data to obtain an encapsulated data packet. The data encapsulation rules define how to package the standardized encoded data into a complete data packet for easy transmission and storage. Encapsulation processing usually includes adding a file header, tail information, check codes, etc. For example, in image processing, the encapsulation rules may require adding a file identifier and length information at the beginning of the data packet and a check code at the end. Through encapsulation processing, the system can ensure the integrity and consistency of the data packet, facilitating subsequent processing and storage. Next, the encapsulated data packet is encrypted through a data encryption algorithm to obtain an encrypted data packet. Data encryption is to protect the security and privacy of data and prevent unauthorized access and tampering. Commonly used encryption algorithms include AES (Advanced Encryption Standard), RSA (Rivest-Shamir-Adleman), etc. The encryption processing converts the encapsulated data packet into ciphertext form, and only the receiving party with the correct key can decrypt and restore the original data. For example, in image processing, the system can use the AES algorithm to encrypt the encapsulated JPEG file to ensure that the image data is not illegally accessed during storage and transmission. Then, based on a preset data storage protocol, storage path planning is performed on the encrypted data packet to obtain storage path information. The data storage protocol defines the storage rules and paths of data in the storage device. Storage path planning is to ensure that the data can be correctly stored in the specified location of the target chip.For example, in an embedded system, the data storage protocol may specify the storage path of data in the internal flash memory or external SD card. The system will, according to these rules, assign a unique storage path to each encrypted data packet and generate storage path information. In this way, the system can ensure the orderly storage and quick retrieval of data. Finally, the encrypted data packet is written to the specified location of the target chip through a specific data writing interface. The data writing interface is a hardware interface for transferring data to the storage device. Common interfaces include SPI (Serial Peripheral Interface), I2C (Inter-Integrated Circuit Bus), SDIO (Secure Digital Input Output), etc. The system writes the encrypted data packet to the specified location of the target chip through these interfaces. For example, in an embedded system, the system can write the encrypted image data to a specified address in the internal flash memory through the SPI interface to ensure the secure storage of data. In this way, the system can efficiently store the compressed data in the target chip, improving the storage efficiency and security of the data. Through this series of steps, the system can convert the compressed and encoded data into a standardized format and securely store it in the target chip. This standardized and encrypted data processing method not only improves the compatibility and security of the data but also provides convenience for subsequent data management and use. For example, in the application scenario of image processing, the system stores the compressed image data in the internal flash memory of the embedded system in JPEG format through the above steps, ensuring the efficiency and security of the image data. In the application scenario of video stream transmission, the system can also, through a similar method, store the compressed video data in the target chip in H.264 format, improving the storage efficiency and transmission speed of the video data. This data processing method based on standardization and encryption provides strong technical support for the efficient storage and transmission of multimedia data.
[0135] The data compression method of the chip in the embodiment of the present invention has been described above. Next, the data compression system of the chip in the embodiment of the present invention will be described. Please refer to Figure 2 , an embodiment of the data compression system of the chip in the embodiment of the present invention includes:
[0136] An acquisition module 21, configured to acquire the data to be compressed of the target chip, perform type analysis on the data to be compressed, and obtain the type of the data to be compressed;
[0137] A selection module 22, configured to select a data compression algorithm for the data to be compressed based on the type of the data to be compressed, and obtain the selected data compression algorithm;
[0138] An identification module 23, configured to perform content awareness identification on the data to be compressed through a preset content awareness algorithm, and obtain the perceived content identification data;
[0139] The encoding module 24 is configured to perform Huffman encoding on the perceived content recognition data and the unperceived content recognition data in the data to be compressed by using a selected post-data compression algorithm, so as to obtain compressed encoded data;
[0140] The conversion module 25 is configured to perform format conversion on the compressed encoded data to obtain standardized encoded data, and store the standardized encoded data in the target chip.
[0141] In this embodiment, for the specific implementation of each unit in the above system embodiment, please refer to that described in the above method embodiment, and details are not described herein again.
[0142] Refer to Figure 3 , in an embodiment of the present invention, a computer device is further provided. The internal structure of the computer device may be as Figure 3 shown. The computer device includes a processor, a memory, a display screen, an input device, a network interface, and a database connected through a system bus. Among them, the processor of the computer design is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is configured to store the corresponding data in this embodiment. The network interface of the computer device is configured to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements the above method.
[0143] Those skilled in the art can understand that Figure 3 the structure shown in
[0144] is only a block diagram of a part of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied.
[0145] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided by the present invention and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.
[0146] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, apparatus, article, or method including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a process, apparatus, article, or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, apparatus, article, or method including the element.
[0147] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied to other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A data compression method for a chip, characterized in that, It includes the following steps: Obtain the data to be compressed of the target chip, and perform type analysis on the data to be compressed to obtain the type of the data to be compressed; Based on the type of the data to be compressed, select a data compression algorithm for the data to be compressed to obtain the selected data compression algorithm; Perform content awareness recognition on the data to be compressed through a preset content awareness algorithm to obtain the recognized content awareness data; Perform Huffman coding on the recognized content awareness data and the unrecognized content awareness data in the data to be compressed through the selected data compression algorithm to obtain compressed coding data; Perform format conversion on the compressed coding data to obtain standardized coding data, and store the standardized coding data in the target chip; The step of obtaining the data to be compressed of the target chip and performing type analysis on the data to be compressed to obtain the type of the data to be compressed includes: Scan the storage area of the target chip through a preset scanning algorithm to obtain the data to be compressed; Extract features from the data to be compressed to obtain the extracted scanning features; Use the Shannon entropy formula to calculate the information entropy of each of the extracted scanning features respectively, and arrange all the information entropy into an information entropy vector; Input the information entropy vector into a preset neural network model for classification to obtain a classification label; Determine the type of the data to be compressed based on the classification label; wherein, the type of the data to be compressed includes message data, image data, and audio data; The step of selecting a data compression algorithm for the data to be compressed based on the type of the data to be compressed to obtain the selected data compression algorithm includes: Call a preset algorithm library, and determine a list of candidate compression algorithms corresponding to the type of the data to be compressed within the algorithm library based on the type of the data to be compressed; Perform dimension decomposition on the information entropy vector to obtain vector dimension information; wherein, the vector dimension information includes vector dimension correlation, vector dimension distribution, vector dimension weight, vector dimension sparsity, and vector dimension size; Use the wavelet transform algorithm to perform frequency domain analysis on each of the vector dimension information to obtain frequency domain feature information; Perform element analysis on the information entropy vector based on the frequency domain feature information to obtain the analyzed vector element features; Perform performance prediction on the candidate compression algorithms in the list of candidate compression algorithms based on the analyzed vector element features to obtain a performance prediction probability distribution; Perform priority sorting and selection on the candidate compression algorithms based on the performance prediction probability distribution to obtain the selected data compression algorithm.
2. The data compression method of the chip according to claim 1, characterized in that, The step of using the Shannon entropy formula to calculate the information entropy of each of the extracted scanning features respectively and arranging all the information entropy into an information entropy vector includes: Perform discretization processing on the extracted scanning features to convert the continuous extracted scanning features into discrete value features to obtain discretized scanning features; Use a parallel computing architecture to calculate the probability of occurrence of each discrete value in the discretized scanning features to obtain a feature probability distribution; Use the Shannon entropy formula to calculate the information entropy of each of the discrete values based on the feature probability distribution; Obtain the information entropy using the following formula; ; is the information entropy, is the total number of discrete values, is the discretized scanning feature of the probability of the nth discrete value, where is the logarithm to the base 2; Use the index marking algorithm to assign a unique index to each of the information entropy values to obtain the indexed information entropy values; Arrange all the indexed information entropy values according to the index order in the index mapping table to obtain the information entropy vector.
3. The data compression method of the chip according to claim 1, wherein The content-aware recognition of the data to be compressed by the preset content-aware algorithm to obtain the perception content recognition data includes: Perform content-aware recognition on the data to be compressed by the preset content-aware algorithm to obtain the recognized perception content; Mark the recognized perception content to obtain the marked recognized perception content; Perform data chunking on the marked recognized perception content to obtain content data chunks; Perform hash transformation on the content data chunks to obtain the transformed hash values; Detect whether there are duplicate transformed hash values. If so, delete the recognized perception content corresponding to the duplicate transformed hash values to obtain the deleted marked recognized perception content; Perform encryption processing on the deleted marked recognized perception content to obtain the encrypted perception content data, and use the encrypted perception content data as the perception content recognition data.
4. The data compression method of the chip according to claim 1, characterized in that The Huffman encoding of the perception content recognition data and the unperceived content recognition data in the data to be compressed by the selected data compression algorithm to obtain the compressed encoded data includes: Use the sliding window technique to analyze and record the occurrence frequency of each symbol in the unperceived content recognition data to obtain a frequency statistics table; wherein, the frequency statistics table is used to reflect the distribution and frequency of each symbol; Obtain the symbol frequency statistical result based on the frequency statistics table; Construct a Huffman tree based on the frequency statistical result and the frequency statistics table; Assign a unique binary codeword to each node in the constructed Huffman tree through the preset recursive encoding algorithm to obtain the encoding mapping table; Perform Huffman encoding conversion on the symbols in the unperceived content recognition data through the encoding mapping table to generate a preliminary compressed data stream; Merge the perception content recognition data with the preliminary compressed data stream to form the compressed encoded data.
5. The data compression method of the chip according to claim 1, characterized in that, The format conversion of the compressed encoded data to obtain the standardized encoded data and store the standardized encoded data in the target chip includes: Parse the compressed encoded data through the preset data parsing algorithm to obtain the compressed encoded data structure; wherein, the compressed encoded data structure includes a data header, a data body, and a check bit; Perform format mapping on the compressed encoded data based on the compressed encoded data structure to obtain the mapped format data, and use the mapped format data as the standardized encoded data; Perform encapsulation processing on the mapped format data through the predefined data encapsulation rules to obtain the encapsulated data packet; Perform encryption processing on the encapsulated data packet using the data encryption algorithm to obtain the encrypted data packet; Perform storage path planning on the encrypted data packet based on the preset data storage protocol to obtain the storage path information; Write the encrypted data packet into the specified location of the target chip through a specific data writing interface.
6. A data compression device for a chip, characterized in that, Include: An acquisition module, configured to acquire data to be compressed of a target chip, and perform type analysis on the data to be compressed to obtain the type of the data to be compressed; A selection module, configured to select a data compression algorithm for the data to be compressed based on the type of the data to be compressed, to obtain a selected data compression algorithm; An identification module, configured to perform content awareness identification on the data to be compressed through a preset content awareness algorithm to obtain perception content identification data; An encoding module, configured to perform Huffman encoding on the perception content identification data and the data of the unperceived content identification in the data to be compressed through the selected data compression algorithm to obtain compressed encoded data; A conversion module, configured to perform format conversion on the compressed encoded data to obtain standardized encoded data, and store the standardized encoded data in the target chip; The acquisition of the data to be compressed of the target chip, and the type analysis of the data to be compressed to obtain the type of the data to be compressed, includes: Scanning the storage area of the target chip through a preset scanning algorithm to obtain data to be compressed; Performing feature extraction on the data to be compressed to obtain extracted scanning features; Calculating the information entropy of each of the extracted scanning features respectively by using the Shannon entropy formula, and arranging all the information entropy into an information entropy vector; Inputting the information entropy vector into a preset neural network model for classification to obtain a classification label; Determining the type of the data to be compressed based on the classification label; wherein, the type of the data to be compressed includes message data, image data, and audio data; The selection of a data compression algorithm for the data to be compressed based on the type of the data to be compressed to obtain a selected data compression algorithm, includes: Invoking a preset algorithm library, and determining a list of candidate compression algorithms corresponding to the type of the data to be compressed based on the type of the data to be compressed in the algorithm library; Performing dimension decomposition on the information entropy vector to obtain vector dimension information; wherein, the vector dimension information includes vector dimension correlation, vector dimension distribution, vector dimension weight, vector dimension sparsity, and vector dimension size; Performing frequency domain analysis on each of the vector dimension information by using a wavelet transform algorithm to obtain frequency domain feature information; Performing element analysis on the information entropy vector based on the frequency domain feature information to obtain analyzed vector element features; Performing performance prediction on the candidate compression algorithms in the list of candidate compression algorithms based on the analyzed vector element features to obtain a performance prediction probability distribution; Performing priority sorting and selection on the candidate compression algorithms based on the performance prediction probability distribution to obtain a selected data compression algorithm.
7. A computer device, comprising a memory and a processor, wherein a computer program is stored in the memory, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Entropy and support vector machine (SVM)-based detection method of advanced persistent threats (APTs)
CN109450860A
Data storage method based on lossless compression algorithm
CN119543957A