Content-based dynamic hybrid data compression

By configuring a processor in the information processing system to determine and generate symbol transformation matrices and probability matrices, and dynamically selecting data compression algorithms, the problem of inflexible data compression in existing technologies is solved, and data transmission and storage efficiency is improved.

CN115391298BActive Publication Date: 2026-03-27DELL PROD LP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-25
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing information processing systems lack flexibility and efficiency in data compression, failing to dynamically select the best compression algorithm based on content, resulting in resource waste and transmission delays.

Method used

By configuring the processor to process training data files, the optimal data compression algorithm is determined, and a symbol transformation matrix and probability matrix are generated based on probabilistic analysis. A suitable data compression algorithm is then dynamically selected for file compression.

Benefits of technology

It achieves dynamic hybrid data compression based on content, improving data transmission and storage efficiency while reducing network bandwidth requirements and resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115391298B_ABST
    Figure CN115391298B_ABST
Patent Text Reader

Abstract

An information handling system includes a processor configured to process a training data file to determine an optimal data compression algorithm. The processor can also perform a compression rate analysis including compressing the training data file using data compression algorithms, calculating a compression rate associated with each of the data compression algorithms, determining an optimal compression rate from the compression rates associated with each data compression algorithm, and determining a desired data compression algorithm associated with the training data file based on the optimal compression rate. The processor can also perform a probability analysis including generating a symbol transition matrix based on the desired data compression algorithm, extracting statistical feature data based on the symbol transition matrix, and generating a probability matrix based on the statistical feature data to determine the optimal data compression algorithm for each segment of a working data file.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates generally to information handling systems, and more particularly to dynamic hybrid data compression based on content. BACKGROUND

[0002] As the value and use of information continues to increase, individuals and businesses seek additional ways to process and store information. One option available to users is information handling systems. An information handling system generally processes, compiles, stores, or communicates information or data for business, personal, or other purposes. Because technology and information handling needs and requirements can vary significantly between different applications, information handling systems can also vary regarding the information handling resources they need to manipulate, the information pattern they process, store, or communicate, and the speeds with which they manipulate, store, or communicate that information. The variety of resources can span across personal, desktop, server, cluster, database, network, memory, and cloud computing environments, and can include one or more computers, graphics interface systems, data storage systems, networking systems, and mobile communication systems. Information handling systems can implement various virtualization architectures. Data and voice communications in information handling systems can occur via networks, which are wired, wireless, or some combination. SUMMARY

[0003] An information handling system includes a processor configured to process a training data file to determine an optimal data compression algorithm. The processor can also perform a compression rate analysis including compressing the training data file using data compression algorithms, calculating a compression rate associated with each of the data compression algorithms, determining an optimal compression rate from the compression rates associated with the each data compression algorithm, and determining a desired data compression algorithm associated with the training data file based on the optimal compression rate. The processor can also perform a probability analysis including generating a symbol conversion matrix based on the desired data compression algorithm, extracting statistical feature data based on the symbol conversion matrix, and generating a probability matrix based on the statistical feature data to determine the optimal data compression algorithm for each segment of a working data file. BRIEF DESCRIPTION OF DRAWINGS

[0004] It should be appreciated that for simplicity and clarity of illustration, elements shown in the figures have not necessarily been drawn to scale. For example, the dimensions of some of the elements can be exaggerated relative to other elements. Embodiments incorporating teachings of the present disclosure are illustrated and described herein in connection with the following drawings, in which:

[0005] It should be appreciated that for simplicity and clarity of illustration, elements shown in the figures have not necessarily been drawn to scale. For example, the dimensions of some of the elements can be exaggerated relative to other elements. Embodiments incorporating teachings of the present disclosure are illustrated and described herein in connection with the following drawings, in which:

[0005] Figure 1 is a block diagram illustrating an information handling system according to embodiments of the present disclosure;

[0006] Figure 2 is a block diagram illustrating an example of a system for content-based dynamic hybrid data compression according to embodiments of the present disclosure;

[0007] Figure 3 is a flow diagram illustrating an example of a process during a training mode for content-based dynamic hybrid data compression according to embodiments of the present disclosure;

[0008] Figure 4 illustrates an example of a symbol transition matrix according to embodiments of the present disclosure;

[0009] Figure 5A illustrates an example of an observation probability matrix according to embodiments of the present disclosure;

[0010] Figure 5B illustrates an example of an initial state distribution matrix according to embodiments of the present disclosure;

[0011] Figure 5C illustrates an example of a state transition probability matrix according to embodiments of the present disclosure;

[0012] Figure 6 is a flow diagram illustrating an example of a process during an operational mode for content-based dynamic hybrid data compression according to embodiments of the present disclosure;

[0013] Figure 7A , Figure 7B , Figure 7C , Figure 7D and Figure 7E illustrates an example of a state sequence generated during an operational mode for content-based dynamic hybrid compression according to embodiments of the present disclosure;

[0014] Figure 8A and Figure 8B is a flow diagram illustrating an example of a method in a training mode for content-based dynamic hybrid compression according to embodiments of the present disclosure;

[0015] Figure 9 is a flow diagram illustrating an example of a method in an operational mode for content-based dynamic hybrid compression according to embodiments of the present disclosure; and

[0016] Figure 10 is a flow diagram illustrating an example of a method for verifying an optimal data compression algorithm for content-based dynamic hybrid compression according to embodiments of the present disclosure.

[0017] The use of the same symbols in different drawings indicates similar or identical items. DETAILED DESCRIPTION

[0018] The following description in connection with the appended drawings is provided to assist in understanding the teachings disclosed herein. The description includes conceptual aspects, as well as

[0019] Figure 1 An embodiment of an information handling system 100 is shown that includes processors 102 and 104, a chipset 110, a memory 120, a graphics adapter 130 connected with a video display 134, a non-volatile RAM (NV-RAM) 140 that includes a basic input and output system / Extensible Firmware Interface (BIOS / EFI) module 142, a disk controller 150, a hard disk drive (HDD) 154, an optical disk drive 156, a disk emulator 160 connected with a solid state drive (SSD) 164, an input / output (I / O) interface 170 connected with expansion resources 174 and a trusted platform module (TPM) 176, a network interface 180, and a baseboard management controller (BMC) 190. Processor 102 is connected to chipset 110 via a processor interface 106, while processor 104 is connected to the chipset via a processor interface 108. In a particular embodiment, processors 102 and 104 are connected together via a high-capacity coherent fabric such as a HyperTransport link, a QuickPath interconnect, or the like. Chipset 110 represents an integrated circuit or a set of integrated circuits that manage data flow between processors 102 and 104 and other elements of information handling system 100. In a particular embodiment, chipset 110 represents a pair of integrated circuits such as a northbridge component and a southbridge component. In another embodiment, some or all of the functionality and features of chipset 110 are integrated with one or more of processors 102 and 104.

[0020] Memory 120 is connected to chipset 110 via a memory interface 122. An example of memory interface 122 includes a double data rate (DDR) memory channel, and memory 120 represents one or more DDR dual in-line memory modules (DIMMs). In a particular embodiment, memory interface 122 represents two or more DDR channels. In another embodiment, one or more of processors 102 and 104 include a memory interface that provides the processor with dedicated memory. The DDR channels and connected DDR DIMMs can conform to a particular DDR standard such as a DDR3 standard, a DDR4 standard, a DDR5 standard, or the like.

[0021] Memory 120 can further represent various combinations of memory types, such as dynamic random access memory (DRAM) DIMMs, static random access memory (SRAM) DIMMs, non-volatile DIMMs (NV-DIMMs), storage class memory devices, read-only memory (ROM) devices, and the like. Graphics adapter 130 connects to chipset 110 via graphics interface 132 and provides video display output 136 to video display 134. Examples of graphics interface 132 include a Peripheral Component Interconnect Express (PCIe) interface, and graphics adapter 130 can include a four-lane (x4) PCIe adapter, an eight-lane (x8) PCIe adapter, a 16-lane (xl6) PCIe adapter, or another configuration as needed or desired. In particular embodiments, graphics adapter 130 is set down on a system printed circuit board (PCB). Video display output 136 can include a digital video interface (DVI), a high-definition multimedia interface (HDMI), a DisplayPort interface, or the like, and video display 134 can include a monitor, a smart television, an embedded display such as a laptop computer display, or the like.

[0022] NV-RAM 140, disk controller 150, and I / O interface 170 connect to chipset 110 via I / O channel 112. An example of I / O channel 112 includes one or more point-to-point PCIe links between chipset 110 and each of NV-RAM 140, disk controller 150, and I / O interface 170. Chipset 110 can also include one or more other I / O interfaces, including a PCIe interface, an Industry Standard Architecture (ISA) interface, a Small Computer Serial Interface (SCSI) interface, an Inter-Integrated 2 C) interface, a system packet interface (SPI), a Universal Serial Bus (USB), another interface, or a combination thereof. NV-RAM 140 includes a BIOS / EFI module 142 that stores machine executable code (BIOS / EFI code) that operates to detect resources of information handling system 100, provide drivers for the resources, initialize the resources, and provide common access mechanisms for the resources. The functions and features of BIOS / EFI module 142 will be further described below.

[0023] The disk controller 150 includes a disk interface 152 that connects the disk controller to a hard disk drive (HDD) 154, an optical disk drive (ODD) 156, and a disk emulator 160. Examples of the disk interface 152 include an integrated drive electronics (IDE) interface, an advanced technology attachment (ATA) such as a parallel ATA (PATA) interface or a serial ATA (SATA) interface, a SCSI interface, a USB interface, a proprietary interface, or a combination thereof. The disk emulator 160 allows the SSD 164 to be connected to the information handling system 100 via an external interface 162. Examples of the external interface 162 include a USB interface, an Institute of Electrical and Electronics Engineers (IEEE) 1394 (FireWire) interface, a proprietary interface, or a combination thereof. Alternatively, the SSD 164 can be disposed within the information handling system 100.

[0024] The I / O interface 170 includes a peripheral interface 172 that connects the I / O interface to an expansion resource 174, a TPM 176, and a network interface 180. The peripheral interface 172 can be the same type of interface as the I / O channel 112, or can be a different type of interface. Thus, when the peripheral interface 172 and the I / O channel 112 are the same type, the I / O interface 170 extends the capabilities of the I / O channel, and when they are different types, the I / O interface translates information from a format suitable for the I / O channel to a format suitable for the peripheral interface 172. The expansion resource 174 can include a data storage system, an additional graphics interface, a network interface card (NIC), a sound / video processing card, another expansion resource, or a combination thereof. The expansion resource 174 can be located on the main circuit board, on a separate circuit board or expansion card disposed within the information handling system 100, on a device external to the information handling system, or a combination thereof.

[0025] The network interface 180 represents a network communication device disposed within the information handling system 100, located on the main circuit board of the information handling system, integrated into another component such as the chipset 110, located in another suitable location, or a combination thereof. The network interface 180 includes a network channel 182 that provides an interface for devices external to the information handling system 100. In a particular embodiment, the network channel 182 is a different type than the peripheral interface 172, and the network interface 180 translates information from a format suitable for the peripheral channel to a format suitable for the external devices.

[0026] In particular embodiments, network interface 180 includes a NIC or host bus adapter (HBA), and examples of network lanes 182 include InfiniBand lanes, Fibre Channel lanes, Gigabit Ethernet lanes, proprietary lane architectures, or combinations thereof. In another embodiment, network interface 180 includes a wireless communication interface, and network lanes 182 include Wi-Fi lanes, near-field communication (NFC) lanes, Bluetooth Low Energy (BLE) lanes, a cellular-based interface such as a Global System for Mobile (GSM) interface, a Code Division Multiple Access (CDMA) interface, a Universal Mobile Telecommunications System (UMTS) interface, a Long Term Evolution (LTE) interface, or another cellular-based interface, or combinations thereof. Network lanes 182 can connect to external network resources (not shown). Network resources can include another information handling system, a data storage system, another network, a grid management system, another suitable resource, or combinations thereof.

[0027] BMC 190 connects to the various elements of information handling system 100 via one or more management interfaces 192 to provide out-of-band monitoring, maintenance, and control of the elements of the information handling system. As such, BMC 190 represents a distinct processing device from processors 102 and 104 that provides various management functions for information handling system 100. For example, BMC 190 can be responsible for power management, cooling management, and the like. The term BMC is commonly used in the context of server systems, while in consumer-grade devices, the BMC can be referred to as an embedded controller (EC). A BMC included in a data storage system can be referred to as a storage sled processor. A BMC included in a chassis of a blade server can be referred to as a chassis management controller, while an embedded controller included in a blade of a blade server can be referred to as a blade management controller. The capabilities and functions provided by BMC 190 can vary greatly based on the type of information handling system. BMC 190 can operate according to the Intelligent Platform Management Interface (IPMI). Examples of BMC 190 include the Integrated Dell™ Remote Access Controller (iDRAC) from Dell®.

[0028] ​Management interface 192 represents one or more out-of-band communication interfaces between BMC 190 and elements of information handling system 100, and can include an Inter-Integrated Circuit (I2C) bus, a System Management Bus (SMBUS), a Power Management Bus (PMBUS), a Low Pin Count (LPC) interface, a serial bus such as a Universal Serial Bus (USB) or Serial Peripheral Interface (SPI), a network interface such as an Ethernet interface, a high-speed serial data link such as a PCIe interface, a Network Controller Sideband Interface (NC-SI), etc. As used herein, out-of-band access refers to operations performed on information handling system 100 separate from the BIOS / operating system execution environment (i.e., separate from execution of code by processors 102 and 104 and programs implemented on information handling system in response to the executed code).

[0029] BMC 190 operates to monitor and maintain system firmware, such as code stored in BIOS / EFI module 142, option ROM for graphics adapter 130, disk controller 150, expansion resources 174, network interface 180, or other elements of information handling system 100, as needed or desired. In particular, BMC 190 includes network interface 194, which can connect to a remote management system to receive firmware updates, as needed or desired. Here, BMC 190 receives a firmware update, stores the update to a data store associated with the BMC, transmits the firmware update to the NV-RAM of a device or system that is the subject of the firmware update, thereby replacing the currently operating firmware associated with the device or system, and reboots the information handling system, after which the device or system utilizes the updated firmware image.

[0030] BMC 190 utilizes various protocols and application programming interfaces (APIs) to direct and control processes for monitoring and maintaining system firmware. Examples of protocols or APIs for monitoring and maintaining system firmware include a graphical user interface (GUI) associated with BMC 190, interfaces defined by Distributed Management Task Force (DMTF) such as Web Services Management (WSMan) interface, Management Component Transport Protocol (MCTP), or interfaces (such as the Dell EMC Remote Access Controller Manager (RACADM) utility, Dell EMC OpenManage Enterprise, Dell EMC OpenManage Server Manager (OMSS) utility, Dell EMC OpenManage Storage Services (OMSS) utility, or Dell EMC OpenManage Deployment Toolkit (DTK) suite), a BIOS setup utility such as invoked by the “F2” boot option, or another protocol or API.

[0031] In particular embodiments, BMC 190 is included on a main circuit board (such as a motherboard, a baseboard, or any combination thereof) of information handling system 100, or integrated into another element of the information handling system (such as chipset 110) or another suitable element, as desired or as appropriate. As such, BMC 190 can be part of an integrated circuit or chipset within information handling system 100. Examples of BMC 190 include iDRAC and the like. BMC 190 can operate on a separate power layer from other resources in information handling system 100. Thus, upon power down of resources of information handling system 100, BMC 190 can communicate with a management system via network interface 194. Here, information can be sent from the management system to BMC 190, and the information can be stored in RAM or NV-RAM associated with the BMC. Information stored in RAM can be lost upon power down of the power layer of BMC 190, while information stored in NV-RAM can be preserved through power down / power up cycles of the power layer of the BMC.

[0032] Information handling system 100 can include additional components and additional buses not shown for the sake of clarity. For example, information handling system 100 can include multiple processor cores, audio devices, and the like. Although a particular arrangement of bus technology and interconnects is shown for purposes of example, one of skill in the art will appreciate that the technology disclosed herein is applicable to other system architectures. Information handling system 100 can include multiple central processing units (CPUs) and redundant bus controllers. One or more components can be integrated together. Information handling system 100 can include additional buses and bus protocols, such as I2C and the like. Additional components of information handling system 100 can include one or more storage devices that can store machine executable code, one or more communication ports for communicating with external devices, and various input and output (I / O) devices such as a keyboard, mouse, and video display.

[0033] For the purposes of this disclosure, the information handling system 100 can include any instrumentality or aggregate of instrumentalities operable to compute, classify, process, transmit, receive, retrieve, originate, switch, store, display, manifest, detect, record, reproduce, handle, or utilize any form of information, intelligence, or data for business, scientific, control, entertainment, or other purposes. For example, an information handling system 100 can be a personal computer, a laptop computer, a smartphone, a tablet device or other consumer electronic device, a network server, a network storage device, a switch, a router, or another network communication device, or any other suitable device, and can vary in size, shape, performance, function, and price, depending on intended use. Additionally, an information handling system 100 can include processing resources for executing machine executable code, such as a processor 102, a programmable logic array (PLA), an embedded device such as a system on a chip (SoC), or other control logic hardware. An information handling system 100 can also include one or more computer readable media for storing machine executable code, such as software or data.

[0034] Applications typically transfer data between a client and a server via uncompressed data files, such as hypertext markup language (HTML), extensible markup language (XML), telemetry, logs, manifest files, etc. Various data compression algorithms are used to reduce resources for storing and / or transferring data. These data compression algorithms can be classified as lossless or lossy. Lossless compression can be reversed to produce the original data, while lossy compression can introduce some errors in the reversal. Lossless compression is typically used for textual data, while lossy compression can be used for voice or images where some errors can be acceptable. Various metrics, such as compression ratio and speed, can be used to evaluate the performance of a data compression algorithm. An optimal data compression algorithm can help reduce network bandwidth when transferring files between a client and a server, and help reduce data storage requirements. The present disclosure provides a content-based dynamic hybrid data compression system and method with the ability to select an optimal data compression to quickly achieve optimal compression.

[0035] Figure 2 An example of an environment 200 for content-based dynamic hybrid data compression is shown. The environment 200 includes an information handling system 205 communicatively coupled with an information handling system 270 via a network 260. The information handling systems 205 and 270 are similar to the information handling system 100, and can include a processor 102, a memory 104, and a network interface 106. The information handling system 205 can include a data compression module 202, while the information handling system 270 can include a data compression module 272. Figure 1information handling system 100. Information handling system 205 includes data training suite 210, data storage 230, index table 240, compression rate analysis table 245, training data sets 250, and learning model 255. Data training suite 210 includes training module 215 and data analyzer 220. Information handling system 270 includes data management suite 275, target data file 290, and compressed data file 295. Data management suite 275 includes work module 280 and data compression features 285. The components of environment 200 can be implemented in hardware, software, firmware, or any combination thereof. The components shown are not drawn to scale, and environment 200 can include more or fewer components. Additional connections between components and / or sensors can be omitted for clarity of description.

[0036] The present disclosure can operate in two modes: a training mode and a work mode. The training mode includes compression rate analysis performed by training module 215 and probability analysis performed by data analyzer 220. In one embodiment, information handling system 205 can be a development system, where training module 215 reads a plurality of training data sets 250 to generate index table 240 and / or compression rate analysis table 245. Compression rate analysis table 245, also referred to as a mapping table, includes the best compression rate associated with each training data file. The association can be one of the factors used when selecting an appropriate compression algorithm for a target data file, also referred to as a work data file, or a segment thereof. Training module 215 can also generate data compression features 285, which can be transmitted to information handling system 270 during installation of data management suite 275. Data compression features 285 can include parameters with information for compression of target data file 290. For example, data compression features 285 can include index table 240 and compression rate analysis table 245. Data compression features 285 can also include matrix A, matrix B, and matrix p, as depicted by learned model 627 of Figure 6

[0037] ​The data analyzer 220 can be configured to perform a probabilistic analysis on the training dataset 250 and generate a learning model 255. The probabilistic analysis can include using a Hidden Markov Model to statistically model a system in which the output of the system, such as a sequence of symbols produced during data compression, is observable, but the specific state changes that produced the output, such as a data compression algorithm, are not observable. As a non-limiting example, the data analyzer 220 can generate a learning model 255 that includes a state transition probability matrix designated as matrix A, an observation probability matrix designated as matrix B, and an initial state distribution matrix designated as matrix p. The learning model 220 can be used to extract statistical features that represent a relationship between the content of the training data file and the data compression algorithm. After extracting the statistical features, the data analyzer 220 can utilize the Baum-Welch algorithm that uses an Expectation-Maximization algorithm to find the maximum likelihood estimate of the parameters of a Hidden Markov Model given a set of observed state sequences.

[0038] The work module 280 can be configured to divide the target data file 290 into one or more segments and predict the best data compression algorithm for each segment. The work module 280 can be configured to use a multi-threaded processing approach to process each segment in parallel. The work module 280 uses the statistical features extracted in the learning model 255 to predict the best data compression algorithm for each segment. The work module 280 can then use a uniform format to combine each of the compressed segments, thereby generating a compressed data file 295 that is similar to the output 620 of Figure 6

[0039] In one embodiment, the work module 280 can receive a request from the data management suite 275 to compress the target data file 290. In response to the request, the work module 280 can retrieve or receive the target data file 290 and use the methods described herein to generate the compressed data file 295. The target data file 290 can be divided into one or more segments of equal length. If the length of the last segment is not equal to the length of the other segments, a flag can be used to store the length of the last segment. The best data compression algorithm can be predicted for each segment based on the data compression features 285. Each segment can then be compressed using the best data compression algorithm. Each segment can be compressed using a different data compression algorithm than the other segments.

[0040] ​The network 260 can be implemented as, or be part of, a storage area network (SAN), a personal area network (PAN), a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a wireless local area network (WLAN), a virtual private network (VPN), an intranet, the Internet, or any other appropriate architecture or system that facilitates the transmission of signals, data, and / or messages or that can be a part of such a transmission. The network 260 can use any storage and / or communication protocol, including, but not limited to, fiber channel, frame relay, asynchronous transfer mode (ATM), Internet Protocol (IP), other packet-based protocols, Small Computer System Interface (SCSI), Internet SCSI (iSCSI), Serial Attached SCSI (SAS), or any other transport that is capable of supporting the communication of signals, data, and / or messages between the various elements of the system 200. The network 260 and its various components can be implemented using hardware, software, or any combination of these.

[0041] The data storage 230 can be a persistent data storage. The data storage 230 can include a solid state disk, a hard disk drive, a tape library, an optical disk drive, a magneto-optical disk drive, an optical disk drive, a disk array, a disk array controller, and / or any computer readable medium operable to store data. The data storage 230 can include a database or collection of data files associated with the processing of the training data set 250, which the data training suite 210 can store, retrieve, and utilize, such as the index table 240, the compression rate analysis table 245, the learning model 255, and the training data set 250.

[0042] Although the data training suite 210 and the data management suite 275 are shown as being deployed into two different information handling systems, the data training suite 210 and the data management suite 275 can be deployed in one information handling system or more than two information handling systems. For example, two information handling systems can each have a data training suite 210 operating in parallel. In another example, one information handling system can include the training module 215 while another information handling system can include the data analyzer 220.

[0043] Those of ordinary skill in the art will appreciate that, Figure 2The configuration, hardware, and / or software components of the depicted environment 200 can vary. For example, the illustrative components within environment 200 are not intended to be exhaustive, but rather are representative to highlight components that can be employed to implement aspects of the present disclosure. Other devices and / or components can be used in addition to or in place of those depicted. The depicted examples are not meant to convey or imply any architectural or other limitations with regard to the present disclosure. With regard to the current description and / or overall disclosure, the depicted examples do not imply or represent that the described arrangements are the only way to implement the aspects of the present disclosure. With regard to the discussion of the figures, reference can also be made to the components illustrated in the other figures for descriptive continuity.

[0044] Figure 3 A flowchart depicting a process 300 of compression analysis during a training mode of a system for content-based hybrid data compression is shown. The process 300 includes a training data set 305, training data files 310, data compression algorithms 315a-315n, a compression rate calculator 320, compressed data files 325a-325n, an index table 335, and a compression rate analysis table 340. Although the process 300 is described in terms of a training mode, the process 300 can be used in a working mode to compress data files. Figure 2 While the information handling system 205 is described in terms of implementing embodiments of the present disclosure, it should be recognized that other systems can be utilized to perform the described methods. Those skilled in the art will appreciate that the flowchart explains a typical example that can be extended to high-level applications or services in practice.

[0045] Figure 3 Annotate with a series of letters A to B1 / B2 to Bn to C. Each of these letters represents a phase of one or more operations. Although the phases are ordered for this example, the phases are shown as an example to aid in understanding the present disclosure and are not to be used to limit the claims. The subject matter falling within the scope of the claims can vary with respect to the order of operations.

[0046] One of the goals of the training mode is to generate a compression rate analysis table 340 that captures the best compression rate for each file, such as the training data files 310 of the training data set 305. The training data set 305 can be selected to represent the expected data set that will be compressed during the working mode. The training data set 305 can be collected from different files of different types or models of computer systems that can be encountered during the working mode. These different files are also referred to as training data files.

[0047] The desired data compression algorithm can also be interchangeably referred to herein as an optimal data compression algorithm for the training mode, or the desired data compression technique can be associated with each training data file of the training data set 305. To determine the desired data compression algorithm, an individual training data file, such as the training data file 310, can be selected from the training data set 305 at stage A. At stages Bl to Bn, each of a plurality of data compression algorithms or techniques, such as the data compression algorithms 315a to 315n, can be used to compress the individual training data file, resulting in the compressed data files 325a to 325n. The compressed data file 325a is generated by compressing the training data file 310 using the data compression algorithm 315a. The compressed data file 325b is generated by compressing the training data file 310 using the data compression algorithm 315b. The compressed data file 325n is generated by compressing the training data file 310 using the data compression algorithm 315n.

[0048] In one example, the data compression algorithm 315a can be a run-length encoding (RLE) data compression algorithm, which is a form of lossless encoding in which a sequence of repeated symbols in the uncompressed data set is replaced by a single control symbol and a single uncompressed symbol in the compressed data set. The data compression algorithm 315b can be a differential pulse code modulation (DPCM) data compression algorithm, which is a form of lossless encoding in which each subsequent symbol in the uncompressed data set is compared to a reference symbol, and the distance between their code points is encoded into the uncompressed data set, provided that the distance is below a distance threshold. DPCM takes advantage of the fact that symbols in the uncompressed data set can cluster within a local region of the data space, and therefore the distance between the reference symbol and the individual uncompressed symbol can be represented using fewer bits than the bits used to represent the individual uncompressed symbol.

[0049] The compression algorithm 315n can be a GZIP data compression algorithm, which refers to one of a variety of implementations of file compression and decompression based on Lempel, Ziv, Welch (LZW) and Huffman codes. Like LZW, GZIP identifies sequences of arbitrary length that have occurred previously and encodes one or more uncompressed symbols as a single control symbol that references the previously observed sequence. LZW is a lossless data compression algorithm that builds a dictionary of tracking symbol sequences. As symbols are read from an uncompressed data file, any identical sequence of symbols that already exists in the dictionary is sought until the dictionary pattern and the input pattern diverge. At this point, a code representing the matching portion of the pattern is passed to the compressed data file and the divergent symbol is added to the dictionary as an extension of the pattern preceding it. LZW can be implemented using variable length codes to allow the dictionary to grow until a separate control symbol is placed into the compressed data set that resets the dictionary and begins anew. According to LZW, the dictionary built by the decoder is identical to the dictionary built by the encoder in producing the compressed data set and thus is able to interpret the symbols in the compressed data set representing sequences.

[0050] At stage C, the best compression ratio is determined from the two or more compression ratios, such as CR1 through CRn. The compression ratio is the ratio between the uncompressed size and the compressed size of a data set or data file, such as:

[0051]

[0052] The best compression ratio is the compression ratio that achieves the greatest compression ratio of the data compression algorithm compared to other data compression algorithms. The association between the training data file and the best compression ratio can be stored in the compression ratio analysis table 340. In this example, the best compression ratio associated with the training data file 310 is CR1, as shown in the index table 335, which is mapped to the data compression algorithm 315a. Thus, the data compression algorithm 315a is the desired data compression algorithm for the training data file 310.

[0053] Figure 4An exemplary matrix 400 is shown for counting occurrences of symbol transitions in a training data set in association with a data compression algorithm. The dimensions of the matrix 400 can be used in the following manner: the symbols appearing vertically on the left and right of the matrix 400 correspond to the starting symbol of a symbol transition in the training data file, while the symbols appearing on the top of the matrix correspond to the ending symbol of a symbol transition in the training data file. The cells corresponding to the starting and ending symbols show the location where the symbol transition was seen. The cells are incremented for each transition and are associated with the desired data compression algorithm associated with the training data file being processed. Those skilled in the art will appreciate that the matrix 400, depicted as a 14x14 matrix, is a simplified example and that in real world conditions, the matrix 400 will be larger than the depicted matrix.

[0054] The matrix 400 is a symbol transition tracking matrix and it tabulates the number of particular symbol transitions that occur for a training data set compressed by a particular data compression algorithm. Shown at the bottom of the matrix 400 are a first training data file 405a, a second training data file 405b, and a third training data file 405n. For simplicity of the non-limiting example, the size of each training data file and the number of data sets is kept small. In this example, the RLE data compression algorithm is determined to be the desired data compression algorithm for the first training data file 405a. The DPCM data compression algorithm is determined to be the desired data compression algorithm for the second training data file 405b. The GZIP data compression algorithm is determined to be the desired data compression algorithm for the third training data file 405n. There can be more or less training data files shown. Additionally, the training data files can have different data compression algorithms that have been found to be the data compression algorithm for a particular training data file and / or training data set.

[0055] The matrix 400 can be initialized to zero in each cell. Each symbol transition in each of the training data sets or training data files can be checked and tabulated in the matrix 400. The matrix 400 can be updated as each training data file is processed. In this example, the first two symbols in the first training data file 405a are “AA”, so the cell 410 is incremented. The letters of the symbols appear on the left and top of the matrix 400 so that these symbol transitions can be indexed. The next symbol transition in the first training data file 405b is “AB”, so the cell 415 is incremented. The symbol transition “BB” occurs four times in a row, so the cell 417 is incremented four times. The rest of the first training data file 405a is checked and tabulated in the same manner. The increments associated with the first training data file 405a are shown underlined, such as “ 1 ”.

[0056] The first two symbols in the second training data file 405b are "AB", so the cell 415 is incremented. Note that the cell 415 is incremented higher than when the first training data file 405a was analyzed. The increment associated with the second training data file 405b is shown in italics, such as "1". The first two symbols in the third training data file 405n are "CG", so the cell 425 is incremented. The increment associated with the third training data file 405n is shown in bold, such as "1". Note that the cell 430 is incremented twice when the symbol transition "DE" occurs on both the second training data file 405b and the third training data file 405n.

[0057] Figure 5A The table 505 (also referred to as matrix B) is shown to be a probability of observation table, and it tabulates the probability of using a data compression algorithm after a particular symbol occurs in the data according to the training data set. The value in the cell 520 of the row 525 indicates that there will be a sixty-six percent probability of using the RLE data compression algorithm after the symbol "A" occurs in the training data file according to the training data set. The calculation of the foregoing value and the values in the row 525 includes the following: as Figure 4 As shown by the matrix 400, there are two symbol transitions from the first training data file 405a that begin with the symbol "A". As shown by the matrix 400, there is one symbol transition from the second training data file 405b that begins with the symbol "A". As shown by the matrix 400, there are no symbol transitions from the third training data file 405c that begin with the symbol "A". Thus, there are a total of three symbol transitions that begin with the symbol "A". Out of the three symbol transitions that begin with the symbol "A", two-thirds or sixty-seven percent are from the first training data file 405a that is associated with the RLE data compression algorithm, one-third or thirty-three percent are from the second training data file 405b that is associated with the DPCM data compression algorithm, and zero-thirds or zero percent are from the third training data file 405c that is associated with the GZIP data compression algorithm. These values populate the row 525. Other rows of the table 505 can be calculated in the same manner. Figure 4 Figure 4

[0058] Figure 5B ​​Table 510 (also referred to as matrix p) is shown to be the initial state distribution matrix, and it tabulates the probabilities of initially using a particular data compression algorithm. The values in matrix p are calculated at the end of the training mode by calculating how many percent of the training data set each data compression algorithm compressed most effectively. In this simple non-limiting example, only three training data files were used, and each training data file was compressed most effectively under a different data compression algorithm. The value in cell 530 of row 535 indicates that one or thirty-three percent of the training data files, out of a total of three training data files, were associated with the RLE data compression algorithm. The same calculations can be performed for the DPCM data compression algorithm and the GZIP data compression algorithm.

[0059] Figure 5C Table 515 (also referred to as matrix A) is shown to be the state transition probability matrix, and it tabulates the probabilities of transitioning from one data compression algorithm to another data compression algorithm based on the training data set. In this simple non-limiting example, the values in matrix A are calculated based on the values in matrix 400, which have been tabulated using the training data files in the training data set. Figure 4 The values in matrix A indicate the probability that the compression algorithm will change state. The values in matrix A can be calculated after matrix 400 has been tabulated using the training data files in the training data set.

[0060] The value in cell 540 of row 545 indicates that if the RLE compression algorithm is used for a symbol transition, then the RLE compression algorithm has a ninety percent probability of being used for the next symbol transition. The calculation of the foregoing value and other values in row 545 includes the following: there are nine symbol transitions associated with the first training data file 405a, as shown in matrix 400. In addition, there is one symbol transition that overlaps with the symbol transitions of the first training data file 405a. Specifically, the symbol transition "AB" is incremented twice in cell 415. The second increment is from the second training data file 405b, which is associated with the DPCM data compression algorithm. Thus, there are a total of ten symbol transitions. Because the first training data file 405a, which is associated with the RLE data compression algorithm, has nine symbol transitions, the RLE to RLE data compression algorithm transition probability is ninety percent out of the total of ten symbol transitions. Because the second training data file 405b, which is associated with the DPCM data compression algorithm, has one overlapping symbol transition, the RLE to DPCM data compression algorithm transition probability is ten percent out of the total of ten symbol transitions. Because the third training data file 405c, which is associated with the GZIP data compression algorithm, has no overlapping symbol transitions, the RLE to GZIP data compression algorithm transition probability is zero percent. Other rows of matrix A can be calculated in the same manner. Figure 4

[0061] Figure 6 ​A flowchart showing an example of a process 600 depicting stages of data compression during a working mode for content-based hybrid data compression is shown. Data compression can be performed on a working data set, a working data file, or a segment thereof. The process 600 includes a target file 605, segments 610a-610n, compressed segments 615a-615n, an output 620, a compression rate analysis table 625, a data structure 650, and a learning model 627 including matrices 630a-630c. Although the process 600 is described in terms of the environment 200 of FIG. 1, it is recognized that other systems can be utilized to perform the described methods. Those skilled in the art will appreciate that the flowchart explains a typical example, which can be extended to high-level applications or services in practice. Figure 2 Although the environment 200 of FIG. 1 is described in terms of the embodiments of the present disclosure, it is recognized that other systems can be utilized to perform the described methods. Those skilled in the art will appreciate that the flowchart explains a typical example, which can be extended to high-level applications or services in practice.

[0062] Figure 6 Annotated with a series of letters A to C1 / C2 to Cn to D. Each of these letters represents a stage of one or more operations. Although the stages are ordered for this example, the stages show one example to help understand the present disclosure and are not to be used to limit the claims. The subject matter falling within the scope of the claims can vary with respect to the order of operations.

[0063] At stage A, the target file 605 is divided into one or more segments. The target data file can also be referred to as a working data file. The segments can be of equal length, except that the last segment can be shorter or longer than the other segments. Here, the target file 605 is divided into segments 610a-610n of a particular length. At stage B, the optimal data compression algorithm associated with each segment can be predicted based on analysis during a training mode. At stages C1-Cn, each segment 610a-610n can be compressed using the best data compression algorithm predicted at stage B for each segment. For example, segment 615a can be compressed using an RLE data compression algorithm, while segment 610b can be compressed using a DPCM data compression algorithm, and segment 610n can be compressed using a GZIP data compression algorithm.

[0064] At stage D, the compressed segments 615a-n are combined using the uniform format based on the data structure 650, resulting in output 620. The output 620 includes delimiters between its elements, such as between compressed content and between compressed indices and compressed content. In the example shown, although commas and parentheses are used as delimiters: 8(1,1C2A5F)(2,A00011233)...(1,4P1A1B,6), other characters or markers can be used as delimiters. The data structure 650 includes a file header 640 and compressed data blocks 645a-n. The file header 640 can include metadata associated with the target file, segments, and / or compressed segments. For example, the file header 640 can include the size and length of each segment, file name, date and time stamp, size of the output or file size, checksum, etc. In some embodiments, the file header 640 can include the number and / or size of the following compressed data blocks, such as compressed data blocks 645a-n. The end of the file header can be marked with a delimiter. Each compressed data block can include an encoder tag 655 and compressed data payload that is a compressed segment, such as compressed segment 615a. The encoder tag 655 indicates which data compression algorithm was used to compress the compressed data segment. If the length of the last segment is different from the other segments, the last compressed data block, such as compressed data block 615n, can also include a length tag 665. If the last compressed data block does not include a length tag 665, the size of the last segment is the same as the other segments.

[0065] Figures 7A to 7E Sequences 700, 730, 740, 750, and 760 are shown, which are sequences of decisions made during the working mode. The sequences of decisions can be based on the set of statistical features extracted and / or computed from the training mode, and include strings of symbols observed in the data blocks of the target data file, history of previous state transitions, or a combination thereof. The sequences of decisions can be used to determine the best data compression algorithm based on symbol-to-symbol analysis. The characters or symbols in segment 705 of the target data file are used as a non-limiting example shown here.

[0066] Figure 7A Sequence 700 shows the first observed symbol at node 715, reading "C" from the data block, and in conjunction with the set of statistical features at decision A, predicting that DPCM is the data compression algorithm used. The prediction is based on the matrix p cell 532, where the value of DPCM data compression used as the initial data compression algorithm with the highest probability is thirty-five percent.

[0067] At decision B, which symbol is predicted to follow symbol C. Based on Figure 4of the three data compression algorithms, such as max(DPCM=0.82, GZIP=0.09, and RLE=0.09). Here, at decision C, there is an eighty-two percent probability that the state starting with the DPCM data compression algorithm will have an ending state of the DPCM data compression algorithm based on the DPCM data compression algorithm at decision C. Second, it is determined whether the symbol D or the symbol G is after the symbol C based on the foregoing decisions. At decision point D, rows 528 and 529 of table 505 (matrix B) are referenced based on their association with the symbols D and G, respectively. For the symbol D, it has a fifty percent probability that the cell 523 is after the symbol C. For the symbol G, it has a twenty-five percent probability that the cell 524 is after the symbol C. The maximum between the two probabilities is selected, such as max{D(DPCM)=0.5, G(DPCM)=0.25)}. Thus, as depicted in node 720, the symbol D is predicted as the next symbol.

[0068] In Figure 7B the system continues to retrieve the next symbol after the symbol in segment 705, which is the symbol B. Since the predicted symbol is different, the node 720 is overwritten with the retrieved symbol B. At decision point E, the next symbol after B is determined to be one of the symbols B, C, or H with a probability based on row 422 of matrix 400. Since there are three possible symbols after the symbol B, it can be determined which of the three symbols is after B. Based on cell 417, the symbol B is predicted to be after the earlier symbol B with four increments compared to each of the cells C or D being incremented by one. At decision point F, the maximum data value between the three data compression algorithms is selected based on row 547 of table 515 (matrix A), such as max(DPCM=0.82, GZIP=0.09, and RLE=0.09) can be determined. Here, if there is an eighty-two percent probability based on the DPCM data compression algorithm that the starting state is the DPCM data compression algorithm, then the DPCM data compression algorithm is predicted to be the ending state. However, based on row 526 of table 505 (matrix B), max(RLE=0.83, DPCM=0.17, GZIP=0), if the starting symbol is B, then there is an eighty-three percent probability to use the RLE data compression algorithm at node 725, which is greater than the eighty-two percent probability of the DPCM data compression algorithm.

[0069] Figure 7CThe sequence 740 is shown after several iterations of the state sequence, as Figure 7B The sequence 730 is shown with similar decisions. Since the symbol is B, the RLE data compression algorithm is also predicted in the next state sequence. Figure 7D The sequence 750 is shown, which depicts the state sequence when predicting the expected data compression algorithm for the next symbol A in the segment 705. The system retrieves the symbol A, and based on the row 421 of the matrix 400, continues to determine that the next symbol at stage G can be either symbol A or B. Based on the row 545 of the table 515 (matrix A), the RLE data compression algorithm is predicted to be the expected data compression algorithm at decision point H. The prediction is based on the maximum between the three data compression algorithms, such as max(RLE=0.9, DPCM=0.1, GZIP=0). Referring to the row 525 of the table 505 (matrix B), there is a sixty-seven percent probability associated with the starting symbol A according to the RLE data compression algorithm. In contrast, in the row 526, there is an eighty-three percent probability associated with the starting symbol B according to the RLE data compression algorithm. The symbol after symbol A is max(A=0.67, B=0.83). Therefore, symbol B is predicted to be the next symbol at decision point I.

[0070] Figure 7E The sequence 760 is shown, which includes the state sequence for the segment 705. Here, the DPCM data compression algorithm is predicted two times out of ten, while the RLE data compression algorithm is predicted eight times out of ten. Therefore, the best data compression algorithm is predicted to be max(RLE=0.8, DPCM=0.2), which is the RLE data compression algorithm with an eighty percent probability. To validate the prediction, the segment 705 can be compressed using the RLE data compression algorithm and the DPCM data compression algorithm, and the segment determines the best compression rate. Using the RLE data compression algorithm generates the compressed text 1C5B3A1C with a compression rate of 10 / 8 or 1.25. Using the DPCM data compression algorithm generates the compressed text C2111110002 with a compression rate of 10 / 11 or.909. The best data compression algorithm is max(RLE=1.25, DPCM=.909), which is the RLE data compression algorithm that confirms the prediction.

[0071] Figure 8A and Figure 8B A flowchart of a method 800 for content-based hybrid data compression is shown. The method 800 can be performed by the environment 200 and / or Figure 3One or more components of process 300 to perform. However, while embodiments of the present disclosure are described in terms of environment 200 and / or process 300, it should be recognized that other systems can be utilized to perform the described methods. Those skilled in the art will understand that the flowchart illustrates a typical example that can be extended in practice to high-level applications or services. Method 800 can be performed at a development system, such as on a server side of a client / server system.

[0072] Method 800 generally begins at block 805 when a plurality of training data sets (also referred to as data corpuses) are selected for use during a training mode of content-based hybrid data compression. The training data sets can be selected from a plurality of training data sets and can be based on the work data files to be processed. For example, if a client is identified as having audio and / or video files relative to text data files, the training data sets can be selected accordingly, such as audio and / or video training data files can be selected for the client processing audio and / or video files.

[0073] At block 810, the method identifies one or more data compression algorithms to use during the training mode. The data compression algorithms can be based on the type of training data files to be processed. At block 815, the method processes each training data file in the training data set. The training data file being processed can be referred to as the current training data file. At block 820, the method processes each training data file using each of the data compression algorithms. At block 825, the method runs the data compression algorithm on the current training data file being processed and calculates a compression rate at block 830. At block 835, the method determines whether additional data compression is to be run on the current training data file. If there is additional data compression, the "yes" branch is taken and the method proceeds to block 820 and processes the current training data file using the next data compression algorithm. If there is no additional data compression algorithm, the "no" branch is taken and the method proceeds to block 840 when the method determines the best compression rate based on the one or more compression rates calculated at block 830.

[0074] At block 845, the method determines whether there are additional training data files to process. If there are additional training data files to process, the method proceeds to block 815 where the method retrieves the next training data file in the training data set for processing. If there are no additional training data files to process, the method proceeds to block 850 where the method updates the compression rate analysis table with the best compression rate for each training data file.

[0075] At box 855, the method initializes and loads the symbolic transformation matrix and learning model into memory, the symbolic transformation matrix and learning model including matrix A, matrix B, and matrix π. The symbolic transformation matrix can be similar to... Figure 4 Matrix 400. At box 865, the method initializes and loads the learning parameters (also referred to herein as the learning model) into memory. The learning model includes models that can be associated with a Hidden Markov Model, such as a first matrix A, a second matrix B, and a third matrix π. The matrices can be defined using arrays that can be used during computation, and the matrices can be loaded into memory and then initialized with zero values ​​in each cell.

[0076] At box 860, the method processes each training data file in the training dataset. At box 865, the method updates the sign transformation matrix for each training data file. For example, as... Figure 4 As shown, the method can update the symbolic transformation matrix, such as matrix 400, for each of the training data files 405a to 405n. The symbolic sequence in each training data file associated with the data compression algorithm mapped to the optimal compression ratio can be used to represent the observed sequence in the symbolic transformation matrix. This can be based on a compression ratio analysis table and an index table (such as...). Figure 2 Compression ratio analysis table 245, index table 240, Figure 3 The optimal data compression algorithm is determined using index table 335 and compression ratio analysis table 340. A pair of symbols is sequentially read from each training data file and used to reflect the data compression ratio of that pair of symbols. Figure 4 The intersecting units in the depicted symbolic transformation matrix are increasing.

[0077] At box 870, the method determines whether there is an additional training data file to process. If an additional training data file exists, the "Yes" branch is taken, and the method proceeds to box 860 and processes the current training data file. If no additional training data file exists, the "No" branch is taken, and the method proceeds to box 875.

[0078] At box 875, the method can extract a set of statistical features from the sign transformation matrix. The method can perform one or more computations to determine the set of statistical features. For example, for matrix π, the method can apply the Baum-Welch algorithm, applying the Baum-Welch forward procedure to matrix A and the Baum-Welch backward procedure to matrix B. The Baum-Welch algorithm uses an expectation-maximization algorithm to find the maximum likelihood estimate of the hidden Markov model parameters given a set of observed feature vectors, such as the set of feature vectors in the sign transformation matrix.

[0079] At block 880, the method can update the learning model, such as matrix A, matrix B, and matrix p, based on the set of statistical features extracted in block 875. After the update, the information in the learning model can be packaged and become part of an installation package to be distributed on the customer’s site.

[0080] Figure 9 A flowchart of a method 900 for content-based hybrid compression is shown. The method 900 can be performed by one or more components of the environment 200 of Figure 2 the process 600 of Figure 6 However, while embodiments of the disclosure are described in terms of the environment 200 and / or the process 600, it should be recognized that other systems can be utilized to perform the described methods. Those skilled in the art will understand that the flowchart illustrates a typical example, which can be extended to high-level applications or services in practice. The method 900 can be performed by a client in a client / server system. Prior to compressing a data file using the method 900, an installation package including information from the learning model can be first installed.

[0081] The method 900 generally begins at block 905, in which a target data file is divided into several segments of equal length. The target data file can be divided based on the number of characters or symbols in the target data file. If the number of characters or symbols is not uniform, the last segment can have more or fewer characters or symbols than the other segments. For non-Latin-based text target data files, the text can be encoded using Unicode prior to dividing the target data file into segments. The same process can be performed for non-text target data files, such as audio or video files. For example, various encoding schemes, such as base64 encoding, can be used to encode the audio or video files prior to processing such audio or video files. Each segment is processed in parallel. Thus, the number of threads used when processing each segment concurrently will be based on the number of segments. For example, the method 900 can use n threads to process n segments. While the example shown below depicts blocks 910a-930a, similar processes can be used to process up to n segments.

[0082] At block 910a, the method stores a segment of the target data file in memory. The segment being processed can be referred to as the current segment. The method proceeds to block 915a, in which the method determines the best data compression algorithm for the current segment. The best data compression algorithm can be determined based on the knowledge gained during the training mode, such as in the process 600 of Figure 6The process described in 600. Specifically, given a sequence of symbols or text (such as fragments or data files), the optimal data compression algorithm can be determined based on statistical characteristics associated with the learned model (such as matrix A, matrix B, and matrix π). Figures 7A to 7E The described method generates an optimal state sequence for the fragment. The state (i.e., the data compression algorithm that appears most frequently in the state sequence) is the optimal data compression algorithm. At box 920a, the method uses the optimal data compression algorithm from box 915a to compress the fragment. The method can determine a first compression ratio for the compressed fragment. At box 925a, the method can verify the optimal data compression algorithm determined at box 915a.

[0083] At block 930a, the method determines to store the compressed fragments in memory. At block 935, the compressed fragments are combined to produce a result similar to... Figure 6 The output follows a uniform format similar to data structure 650. The output may include the size or length of each segment and compressed data block. The size or length can be the number of characters in each segment. Each compressed data block includes a prefix and a compressed segment. The prefix may be an index of the optimal data compression algorithm used for the compressed segment. Because the size of the last compressed segment may differ from the size of another compressed segment, the compressed data block associated with the last compressed segment may also include a suffix, where the suffix is ​​the size or length of the last compressed segment. If the size or length of the last compressed segment is the same as other segments, the suffix may not be included. After verification, the method ends.

[0084] Figure 10 A flowchart of a method 1000 for content-based hybrid data compression is shown. Specifically, method 1000 is... Figure 9 A detailed illustration of box 925 is provided. Method 1000 typically begins at box 1005, where the method determines from the training dataset whether the current segment has a matching data file. For example, a comparison of symbols and / or symbol transformations in the data file and the segment can be made based on one or more factors (such as the type of data file, such as whether the data file is audio, video, or text). At decision box 1010, if a matching training data file exists, the "Yes" branch is taken, and the method proceeds to box 1015. If no matching data file exists, the "No" branch is taken, and the method proceeds to box 1025.

[0085] At box 1015, based on the knowledge gained during training mode, a data compression algorithm associated with the matched training data file is selected. This can be achieved using... Figure 2the compression rate analysis table 340 to select a data compression algorithm. At block 1020, the method can validate the best data compression algorithm. The method can compress the segment using the data compression algorithm based on the compression rate analysis table and determine a second compression rate. The method can perform a comparison of the first compression rate and the second compression rate. If the first compression rate is greater than or equal to the second compression rate, the best compression rate is validated. At block 1025, the current segment is added to the training data set. The method proceeds to block 1030, in which the method proceeds to the training mode in which the target data file is trained. For example, blocks 910a-930a can be performed in parallel with blocks 910b-930b until Figure 9 blocks 910n-930n of the method 900.

[0086] Although in some implementations, Figure 8A 、 Figure 8B 、 Figure 9 and Figure 10 exemplary blocks of the method 800, the method 900, and the method 1000, the method 800, the method 900, and the method 1000 can include additional blocks, fewer blocks, different blocks, or differently arranged blocks than those depicted in Figure 8A 、 Figure 8B 、 Figure 9 and Figure 10 depicted in FIGS. 8-10. Additionally or alternatively, two or more of the blocks of the method 800, the method 900, and the method 1000 can be performed in parallel. For example, blocks 805 and 810 can be performed in parallel.

[0087] In accordance with various embodiments of the present disclosure, the methods described herein can be implemented by software programs that are executable by computer systems. Additionally, in an exemplary, non-limited embodiment, implementations can include distributed processing, component / object distributed processing, and parallel processing. Alternatively, virtual computer system processing can be constructed to implement one or more of the methods or functionality as described herein.

[0088] The present disclosure contemplates a computer-readable medium that includes instructions or receives and executes instructions responsive to a propagated signal; so that a device connected to a network can communicate voice, video or data over the network. Further, the instructions can be transmitted or received over the network via the network interface device.

[0089] Although the computer-readable medium is illustrated as a single medium, the term "computer-readable medium" includes a single medium or multiple media, such as a centralized or distributed database, and / or associated caches and servers that store one or more sets of instructions. The term "computer-readable medium" shall also include any medium that is capable of storing, encoding or carrying a set of instructions for execution by a processor or that cause a computer system to perform any one or more of the methods or operations disclosed herein.

[0090] In particular non-limiting example embodiments, the computer-readable medium can include a solid-state memory such as a memory card or other package that houses one or more non-volatile read-only memories. Further, the computer-readable medium can be a random access memory or other volatile re-writable memory. Additionally, the computer-readable medium can include a magnetic or optical medium, such as a disk or tape, or another storage device that stores information in the form of magnetic or optical signals. A digital file attachment to an e-mail or other self-contained information archive or set of archives is considered a distribution medium equivalent to a tangible storage medium. Accordingly, the disclosure is considered to include a computer-readable medium or a distribution medium that includes, without limitation, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, or a propagation medium that includes, without limitation, a carrier wave, signals, or other flow of information in coherent or other forms, in compliance with the general purpose of the disclosure. Accordingly, the disclosure is considered to include any one or more of a computer-readable medium or a distribution medium and other equivalents and successor media, in which data or instructions can be stored.

[0091] Although only a few exemplary embodiments have been described in detail above, those skilled in the art will readily comprehend numerous modifications. Many modifications can be made to the exemplary embodiments and the application described herein without departing from the novel teachings and advantages of the disclosure. Therefore, no limitation is placed on the scope of the disclosure by the detailed description and examples. All such modifications are to be included within the scope of the embodiments of the present disclosure as defined by the appended claims. In the claims, any means-plus-function clause is intended to cover the structures described herein as performing the recited function and also cover structures not expressly shown or described but which are not beyond the scope of the disclosure.

Claims

1. A content-based dynamic hybrid data compression method, comprising: Processing training data files to predict the optimal data compression algorithm, wherein the processing of the training data files includes: Perform compression ratio analysis, the compression ratio analysis including: The training data file is compressed using a variety of data compression algorithms; Calculate the compression ratio associated with each of the data compression algorithms in the data compression algorithms; The optimal compression ratio is determined based on the compression ratio associated with each of the data compression algorithms; and Identify the desired data compression algorithm associated with the optimal compression ratio of the training data file; and Perform probability analysis, the probability analysis including: Generate a symbol transformation matrix based on the desired data compression algorithm; Statistical feature data are extracted based on the symbol transformation matrix; and A probability matrix is ​​generated based on the statistical feature data to predict the optimal data compression algorithm for each segment of the working data file.

2. The method according to claim 1, wherein a compression ratio analysis table is generated, the compression ratio analysis table including the association between the training data file and the desired data compression algorithm.

3. The method of claim 1, wherein the optimal compression ratio has a maximum value compared to at least one other compression ratio associated with the training data file.

4. The method of claim 1, further comprising processing the working data file, wherein the processing of the working data file includes dividing the working data file into one or more segments and compressing each segment based on the optimal data compression algorithm to generate one or more compressed segments.

5. The method of claim 4, wherein each segment of the working data file has an equal length.

6. The method of claim 4, further comprising combining each compressed segment into a compressed data file using a uniform format.

7. The method of claim 6, wherein the first data compression algorithm is used to compress the first segment, and the second data compression algorithm is used to compress the second segment of the compressed data file.

8. The method according to claim 1, wherein the probability matrix includes an initial state distribution matrix, a state transition probability matrix, and an observation probability matrix.

9. An information processing system, comprising: A memory for storing a probability matrix for determining the optimal data compression algorithm; as well as A processor coupled to the memory, the processor being configured to process training data files to determine the optimal data compression algorithm, wherein the processor is further configured to: Perform compression ratio analysis, the compression ratio analysis including: The training data file is compressed using a variety of data compression algorithms; Calculate the compression ratio associated with each of the data compression algorithms in the data compression algorithms; The optimal compression ratio is determined from the compression ratio associated with each of the data compression algorithms; and Based on the optimal compression ratio, determine the desired data compression algorithm associated with the training data file; and Perform probability analysis, the probability analysis including: Generate a symbol transformation matrix based on the desired data compression algorithm; Statistical feature data are extracted based on the symbol transformation matrix; and A probability matrix is ​​generated based on the statistical feature data to determine the optimal data compression algorithm for each segment of the working data file.

10. The information processing system of claim 9, wherein the processor is further configured to generate a compression ratio analysis table, the compression ratio analysis table including the association between the training data file and the desired data compression algorithm.

11. The information processing system of claim 9, wherein the optimal compression ratio has a maximum value compared to at least one other compression ratio associated with the training data file.

12. The information processing system of claim 9, wherein the processor is further configured to process the working data file, the processing comprising dividing the working data file into one or more segments and compressing each segment based on the optimal data compression algorithm to generate one or more compressed segments.

13. The information processing system of claim 12, wherein the processor is further configured to combine each compressed segment into a compressed data file using a uniform format.

14. The information processing system according to claim 12, wherein each segment of the working data file has an equal length.

15. A non-transitory computer-readable medium comprising code, said code performing a method when executed, said method comprising: Processing training data files to determine the optimal data compression algorithm, wherein the processing of the training data files includes: Perform compression ratio analysis, the compression ratio analysis including: The training data file is compressed using a variety of data compression algorithms; Calculate the compression ratio associated with each of the data compression algorithms in the data compression algorithms; The optimal compression ratio is determined from the compression ratio associated with each of the data compression algorithms; and Based on the optimal compression ratio, determine the desired data compression algorithm associated with the training data file; and Perform probability analysis, the probability analysis including: Generate a symbol transformation matrix based on the desired data compression algorithm; Statistical feature data are extracted based on the symbol transformation matrix; and A probability matrix is ​​generated based on the statistical feature data to determine the optimal data compression algorithm for each segment of the working data file.

16. The method of claim 15, further comprising processing the working data file, wherein the processing of the working data file includes dividing the working data file into one or more segments and compressing each segment based on the optimal data compression algorithm to generate one or more compressed segments.

17. The method of claim 16, wherein each segment of the working data file has an equal length.

18. The method of claim 16, further comprising combining each compressed segment into a compressed data file using a uniform format.

19. The method of claim 18, wherein a first data compression algorithm is used to compress a first segment, and a second data compression algorithm is used to compress a second compressed segment of the compressed data file.

20. The method of claim 15, wherein the probability matrix comprises an initial state distribution matrix, a state transition probability matrix, and an observation probability matrix.

Citation Information

Patent Citations

  • Self-adaptation data prediction coding algorithm based on information entropy optimization

    CN103888144A

  • Data compression method and equipment and computer readable storage medium

    CN108197168A