An Internet data compression and storage method and system

By blocking the Internet data resources and dynamically selecting compressed flow characteristics, the inefficiency problem in traditional Internet data compression methods is solved, and more efficient and reliable data compression and storage are achieved.

CN119960684BActive Publication Date: 2025-07-25BEIJING SCI & TECH PATENT OFFICE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510032183.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-07-25
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

Traditional Internet data compression methods have problems such as inefficient compression efficiency, poor compression effect, and inability to adapt to dynamically changing compression requirements, resulting in a large space for compressed data.

Method used

By blocking the Internet data resources, the compression progress of the candidate resource blocks is obtained, and the compression flow characteristics are dynamically selected based on the compression progress, and applied to the compression unit for compression and storage.

Benefits of technology

It significantly improves the efficiency and accuracy of data compression, reduces the space required for data storage, and improves the speed of data compression and storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119960684B_ABST
    Figure CN119960684B_ABST
Patent Text Reader

Abstract

The present invention provides an Internet data compression and storage method and system, which relates to the field of computer technology. By performing block processing on the resource data to be compressed and dynamically selecting candidate resource blocks according to the compression progress to determine the compression stream characteristics, the efficiency and accuracy of data compression are significantly improved. Applying the compression stream characteristics to the compression unit further optimizes the compression process, making the compression and storage of Internet data resources on the cloud storage medium more efficient and reliable. This method not only reduces the space required for data storage, but also improves the speed of data compression and storage, thus contributing to the efficient management of large-scale Internet data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to an Internet data compression and storage method and system. Background Art

[0002] With the rapid development of Internet technology, Internet data resources have shown an explosive growth trend. These data resources are not only huge in quantity but also diverse in type, including various forms such as text, images, videos, and audios. In order to effectively store and manage these massive data, data compression technology has become an indispensable and important means. However, traditional data compression methods often have problems such as low compression efficiency, poor compression effect, and inability to adapt to dynamically changing compression requirements.

[0003] In the process of Internet data compression, how to efficiently divide data blocks, how to select a suitable compression algorithm according to data characteristics, and how to dynamically adjust the compression strategy during compression are all key factors affecting the compression effect. Traditional compression methods often adopt fixed block division methods and compression algorithms, and cannot be flexibly adjusted according to the actual characteristics of the data, resulting in low compression efficiency and relatively large space occupancy of the compressed data. Summary of the Invention

[0004] In view of the above-mentioned problems, in combination with the first aspect of the present invention, embodiments of the present invention provide an Internet data compression and storage method, and the method includes:

[0005] Obtain the resource data to be compressed corresponding to the Internet data resource, and determine a plurality of resource blocks according to the resource data to be compressed, and each of the resource blocks includes at least data span information;

[0006] When compressing the Internet data resource, obtain candidate resource blocks, and obtain the current compression progress of the Internet data resource;

[0007] When the compression progress is within the interval of the data span information included in the candidate resource block, determine the compression stream characteristics according to the candidate resource block;

[0008] Apply the compression stream characteristics to the compression unit, and load the compression unit onto the cloud storage medium, and the cloud storage medium is used to compress and store the Internet data resource.

[0009] In another aspect, embodiments of the present invention further provide an Internet data compression and storage system, including a processor and a machine-readable storage medium, the machine-readable storage medium is connected to the processor, the machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.

[0010] Based on the above aspects, in the embodiments of the present application, by performing block processing on the resource data to be compressed and dynamically selecting candidate resource blocks according to the compression progress to determine the compression stream characteristics, the efficiency and accuracy of data compression are significantly improved. Applying the compression stream characteristics to the compression unit further optimizes the compression process, making the compression and storage of Internet data resources on the cloud storage medium more efficient and reliable. This method not only reduces the space required for data storage, but also improves the speed of data compression and storage, thus contributing to the efficient management of large-scale Internet data. Brief Description of the Drawings

[0011] Figure 1 It is a schematic flowchart of the execution process of the Internet data compression and storage method provided by the embodiments of the present invention.

[0012] Figure 2 It is a schematic diagram of the hardware architecture of the Internet data compression and storage system provided by the embodiments of the present invention. Detailed Embodiments

[0013] The present invention will be specifically described below with reference to the accompanying drawings of the specification. Figure 1 It is a schematic flowchart of the Internet data compression and storage method provided by an embodiment of the present invention. The Internet data compression and storage method will be introduced in detail below.

[0014] Step S110: Obtain the resource data to be compressed corresponding to the Internet data resource, and determine a plurality of resource blocks according to the resource data to be compressed. Each resource block includes at least data span information.

[0015] In this embodiment, the Internet data resource is, for example, a data set in a large enterprise document management system. This data set contains various types of files, such as documents, images, audio, etc., with a total size reaching several TB.

[0016] First, for the acquisition of the resource data to be compressed, in the case of an integrated coding document. For example, in this enterprise document management system, there is a large project document, which is an integrated coding document. First, obtain the resource file corresponding to this Internet data resource. This file may contain information on various aspects of the project, such as project planning, execution status, relevant personnel information, etc. These information are integrated together in a specific coding manner. Then, disassemble this resource file. During the disassembly process, structured attribute data of the Internet data resource will be generated, such as the chapter structure and title hierarchy of the document; unstructured attribute data, such as some comments and annotations; and the resource data to be compressed. The resource data to be compressed here contains multiple compressed storage coding blocks, and these coding blocks may be divided according to different parts or functional modules of the document. Finally, extract these multiple compressed storage coding blocks to generate multiple resource chunks. Each resource chunk includes at least data span information, and this data span information can be understood as the position range of this resource chunk in the entire document. For example, the data content from line 1000 to line 2000.

[0017] In the case of a stand-alone coding document, assume there is a separate high-definition image file as the resource data to be compressed. Disassemble this resource data to be compressed. This image file may be encoded according to certain coding rules during storage. During the disassembly process, multiple compressed storage coding blocks will be generated. For example, this image may be divided into multiple small blocks for encoding and storage, and these small blocks are the compressed storage coding blocks. Then extract these compressed storage coding blocks to generate multiple resource chunks. Each resource chunk also includes data span information, and this data span information can be represented as a certain area in the image. For example, the upper left quarter area of the image corresponds to a resource chunk, and its data span information is the coordinate range of this area in the entire image.

[0018] Suppose the resource data to be compressed includes the first coding document and the second coding document. For example, the first coding document is a company's financial statement document, and the second coding document is a financial analysis report document related to it. First, extract the first coding document to determine multiple first resource chunks. For the financial statement document, multiple first resource chunks may be divided according to different financial indicators or different parts of the statement (such as the balance sheet part, the income statement part, etc.). Then extract the second coding document to determine multiple second resource chunks. For the financial analysis report document, multiple second resource chunks may be divided according to different aspects of the analysis (such as solvency analysis, profitability analysis, etc.).

[0019] For each first resource block, the first resource block and the corresponding second resource block are to be fused to generate a resource block. For example, the first resource block is the balance sheet part, and the corresponding second resource block is the solvency analysis part. If the first data span information in the first resource block matches the second data span information in the second resource block, then the first resource block and the second resource block are directly fused to generate a resource block. Suppose the data span information of the first resource block represents the liability data part in the balance sheet, and the data span information of the second resource block is also exactly the solvency analysis part for the liability data, then these two blocks can be directly fused.

[0020] If the first data span information in the first resource block is different from the second data span information in the second resource block and they have an overlapping part, for example, the data span information of the first resource block represents the current assets part in the balance sheet, and the data span information of the second resource block is the solvency analysis part for the current ratio (involving current assets and current liabilities), and there is an overlapping part here. When the first front-end position information precedes the second front-end position information, and the first end position information precedes the second end position information, according to the first front-end position information and the second end position information, a first data resource sub-block is decomposed from the first resource block, and a second data resource sub-block is decomposed from the second resource block. For example, if the front end of the current assets part in the first resource block is cash assets and the back end is inventory assets, and the front end of the current ratio analysis part in the second resource block is the overall concept of current assets and the back end is the part related to current liabilities, then a first data resource sub-block containing only cash assets can be decomposed from the first resource block, and a second data resource sub-block not containing the overall concept of current assets can be decomposed from the second resource block. According to the first end position information and the second end position information, a third data resource sub-block is decomposed from the first resource block, which can be a sub-block of the inventory assets part decomposed from the first resource block. According to the first front-end position information and the second front-end position information, a fourth data resource sub-block is decomposed from the second resource block, which can be a sub-block containing the overall concept of current assets decomposed from the second resource block. Then, two data resource sub-blocks with matching data span information among these data resource sub-blocks are fused to generate a fused data resource sub-block. Finally, the data resource sub-blocks with different data span information among the multiple data resource sub-blocks, and the fused data resource sub-block are fused according to the resource transfer order to generate a resource block.

[0021] Alternatively, when the first front-end position information precedes the second front-end position information and the first end position information follows the second end position information. For example, the first resource block is the entire balance sheet, and the second resource block is part of the solvency analysis. According to the first front-end position information and the second front-end position information, the fifth data resource sub-block is obtained by decomposing the second resource block, which may be the sub-block related to the front-end part of the balance sheet in the solvency analysis. According to the first front-end position information and the first end position information, the sixth data resource sub-block is obtained by decomposing the second resource block, which may be the sub-block related to the entire balance sheet in the solvency analysis. According to the first end position information and the second end position information, the seventh data resource sub-block is obtained by decomposing the second resource block, which may be the sub-block related to the back-end part of the balance sheet in the solvency analysis. The first resource block is output as the eighth data resource sub-block, and then fused according to the fusion rules mentioned above to generate a resource block.

[0022] Or, when the first front-end position information precedes the second front-end position information and the first end position information is the same as the second end position information, or when the first front-end position information is the same as the second front-end position information and the first end position information follows the second end position information. For example, the first resource block is a part of the balance sheet, and the second resource block is the related analysis part. According to the first front-end position information and the first end position information, the second resource block is decomposed to generate the ninth data resource sub-block and the tenth data resource sub-block. The first resource block is output as the eleventh data resource sub-block, and then fused to generate a resource block.

[0023] Step S120, when compressing the Internet data resources, obtain candidate resource blocks and obtain the current compression progress of the Internet data resources.

[0024] Continuing with the above enterprise document management system as an example, when starting to compress this large data set, the compression program will obtain candidate resource blocks from multiple previously determined resource blocks in a certain order. Assume that the compression program uses a block-by-block compression method and obtains candidate resource blocks according to the storage order of the resource blocks in the data set. For example, start obtaining from the resource block corresponding to the beginning of the project document.

[0025] Meanwhile, it is necessary to obtain the current compression progress of Internet data resources. This compression progress can be represented by the ratio of the amount of data that has been compressed to the total amount of data. For example, within a certain period after the start of compression, 100 MB of data has been compressed, and the total size of the entire Internet data resource is 1000 MB, then the current compression progress is 10%. The purpose of obtaining this compression progress is to determine how to process the candidate resource chunks based on its relationship with the data span information of the candidate resource chunks in the subsequent stage.

[0026] Step S130, when the compression progress is within the interval of the data span information included in the candidate resource chunk, determine the compression stream characteristics based on the candidate resource chunk.

[0027] For example, assume that when compressing project documents in an enterprise document management system, the current compression progress is 30%. The data span information of a candidate resource chunk indicates that this chunk contains the data content from the 30% to 40% positions in the document. Since the current compression progress is within the interval of the data span information of this candidate resource chunk, the compression stream characteristics are determined based on this candidate resource chunk.

[0028] Suppose this candidate resource chunk also includes resource node information, multiple resource content data, and the resource attributes of each resource content data. First, determine the compression storage mapping partition based on the resource node information in the candidate resource chunk. For example, if the resource node information indicates that this chunk is about the content of a specific stage in the project execution process, then a dedicated compression storage mapping partition is determined for the content of this stage.

[0029] Then, each resource content data in the candidate resource chunk is compressed into the compression storage mapping partition according to the resource attributes of the resource content data to generate the compression stream characteristics. For example, the resource content data in this chunk includes text descriptions, related images, and some audio annotations. For the resource content data with text attributes such as text descriptions, the Huffman coding algorithm in lossless compression algorithms is selected for compression. For the resource content data with image attributes such as images, the JPEG compression algorithm is selected for compression. For the resource content data with audio attributes such as audio annotations, the MP3 compression algorithm is selected for compression.

[0030] During the compression process, all resource content data in the candidate resource chunks are traversed. According to the resource attributes of each resource content data, the resource content data with the same or similar resource attributes are grouped together to generate resource content data groups classified by resource attributes. For each resource content data group, analyze the resource attribute characteristics corresponding to the resource content data group to determine the corresponding compression algorithm as described above. For each resource content data group, perform preprocessing operations according to the selected compression algorithm to generate preprocessed resource content data groups and a list of compression algorithms matching each resource content data group. For each preprocessed resource content data group, use the corresponding compression algorithm for compression operations to generate compressed resource content data blocks. Add marking information to each compressed resource content data block. The marking information includes the resource attribute identifier of the original resource content data and the sequence identifier in the original candidate resource chunks. Analyze the structural characteristics of the compressed storage mapping partition. According to the resource attribute identifier in the marking information of each compressed resource content data block, determine the reference storage area of each compressed resource content data block in the compressed storage mapping partition. In each reference storage area, determine the specific target storage location according to the sequence identifier or preset rules, so that each compressed resource content data block has a unique storage location in the implemented compressed storage mapping partition. Finally, configure the storage of each compressed resource content data block according to its target storage location in the compressed storage mapping partition to generate the compressed stream characteristics.

[0031] Again, assume the case where the candidate resource chunk contains the first resource chunk and the second resource chunk. The first resource chunk corresponds to the first resource type, and the second resource chunk corresponds to the second resource type. Based on the resource type determination instruction, obtain the reference resource type of the resource data to be compressed in the Internet data resource. If the reference resource type is the first resource type, for example, it is a text type, then determine the compressed stream characteristics based on the first resource chunk in the candidate resource chunk and operate according to the compression method for text resources mentioned above. If the reference resource type is the second resource type, for example, it is an image type, then determine the compressed stream characteristics based on the second resource chunk in the candidate resource chunk and operate according to the compression method for image resources. If the reference resource type is the first resource type and the second resource type, then determine the compressed stream characteristics based on the entire candidate resource chunk and comprehensively consider the compression processing of text and image resources.

[0032] Step S140, apply the compressed stream characteristics to the compression unit and load the compression unit onto the cloud storage medium, where the cloud storage medium is used to compress and store the Internet data resource.

[0033] In this embodiment, for the cloud compression service, the compression unit therein has specific type and functional requirements. First, according to the type of the compression unit, determine the compatibility requirements of the compression unit for the compression stream characteristics, and generate an adaptation analysis result. For example, the compression unit may be optimized for a certain format of compression algorithm. If the compression algorithm in the compression stream characteristics does not exactly match the requirements of the compression unit, then an incompatibility situation will be shown.

[0034] If the adaptation analysis result shows an incompatibility situation, convert the compression stream characteristics according to the compression requirements of the compression unit to generate the converted compression stream characteristics. For example, if the compression unit requires a compression algorithm in a specific coding format while the algorithm in the compression stream characteristics is in another format, conversion is needed.

[0035] Then, perform an initialization operation on the compression unit and load the compression stream characteristics into the compression unit after the initialization operation, so that the compression unit, according to the internal defined processing logic, compresses each resource content data block according to the storage configuration information and compression algorithm information in the implementation compression stream characteristics, and generates the compressed data after being processed by the compression unit.

[0036] Next, query the status information of the cloud storage medium. This status information includes the current available capacity of the cloud storage medium, the distribution of storage areas, and the performance parameters of the storage medium. For example, the cloud storage medium may be a cluster composed of multiple storage servers, with a current available capacity of 500TB, the storage areas are distributed in different data centers, and the performance parameters include read and write speeds, etc.

[0037] According to the type of the cloud storage medium, determine the storage strategy for storing the compressed data, and preprocess the compressed data according to the storage strategy. For example, if the cloud storage medium is based on object storage, the storage strategy may require splitting the compressed data into chunks, each chunk being 1MB in size, adding metadata to the compressed data, such as data source, compression time, etc., and encrypting the compressed data according to the storage strategy, using the AES encryption algorithm.

[0038] According to the status information of the cloud storage medium, determine the storage location of the compressed data in the cloud storage medium. If the storage area of a certain data center has a large available capacity and a fast read and write speed, store the preprocessed compressed data in the storage area of this data center. Finally, store the preprocessed compressed data in the storage location of the cloud storage medium.

[0039] During the process of loading the compression unit onto the cloud storage medium, obtain the first address information and the first capacity information of the cloud storage medium. Assume that the first address information of the cloud storage medium is a specific network address, such as 192.168.1.100, and the first capacity information is the aforementioned 500TB. Determine the second address information of the compression unit based on the first address information, the first capacity information, and the resource node information of the compression stream characteristics. For example, according to the resource node information, determine that this compression unit should be loaded into a specific storage partition in the cloud storage medium, and the address corresponding to this partition is the second address information of the compression unit, such as 192.168.1.100:8080 / partition11. Then load the cloud storage medium based on the first address information and load the compression unit based on the second address information.

[0040] In addition, if the compression progress mentioned in step S130 is not within the interval of the data span information, determine the interval parameter between the front-end position information and the compression progress in the data span information. For example, if the current compression progress is 50%, and the data span information of the candidate resource block is the content from 60% to 70%, then calculate the interval parameter as 10%. After this interval parameter, determine the compression stream characteristics based on the candidate resource block. Then operate according to the steps described above of applying the compression stream characteristics to the compression unit and loading it onto the cloud storage medium.

[0041] Based on the above steps, the embodiments of the present application significantly improve the efficiency and accuracy of data compression by performing block processing on the resource data to be compressed and dynamically selecting candidate resource blocks according to the compression progress to determine the compression stream characteristics. Applying the compression stream characteristics to the compression unit further optimizes the compression process, making the compression and storage of Internet data resources on the cloud storage medium more efficient and reliable. This method not only reduces the space required for data storage but also improves the speed of data compression and storage, thus contributing to the efficient management of large-scale Internet data.

[0042] In a possible implementation manner, when the resource data to be compressed is an integrated coding document, step S110 may include:

[0043] Obtain the resource file corresponding to the Internet data resource.

[0044] Decompose the resource file to generate the structured attribute data, unstructured attribute data, and resource data to be compressed of the Internet data resource, and the resource data to be compressed includes multiple compression storage coding blocks.

[0045] Extract the multiple compression storage coding blocks to generate multiple resource blocks.

[0046] In a possible implementation, when the resource data to be compressed is an independently encoded document, step S110 may further include:

[0047] Disassemble the resource data to be compressed to generate a plurality of compressed storage coding blocks.

[0048] Extract the plurality of compressed storage coding blocks to generate a plurality of resource sub-blocks.

[0049] In this embodiment, consider a knowledge management system within an enterprise that stores a large amount of Internet data resources. These Internet data resources exist in various forms, including integrated coding documents and independently encoded documents, etc. The following elaborates in detail the process of obtaining the resource data to be compressed and determining the resource sub-blocks for different types of documents.

[0050] For the case where the resource data to be compressed is an integrated coding document, take a large project document within an enterprise as an example. This document is an integrated coding document that covers all process information of the project from planning, execution to completion. First, obtain the resource file corresponding to this Internet data resource. This process involves locating and reading the storage file of this project document from the repository of the knowledge management system. This resource file may adopt a custom coding format to integrate various information in the project for easy storage and management.

[0051] Next, perform a disassembling operation on this resource file. During the disassembling process, structured attribute data of the Internet data resource will be generated. For example, the chapter title structure in the project document, such as under the project planning section, there are sub-titles like goal setting and resource budget. These title structures reflect the organizational structure of the document content and belong to structured attribute data; unstructured attribute data may include comments and annotations added by employees during the document editing process. These information have no fixed structural pattern; while the resource data to be compressed contains a plurality of compressed storage coding blocks. These compressed storage coding blocks are logical units formed during the encoding process of the resource file. For example, they are divided according to different stages or different functional modules of the project, such as the coding block corresponding to the task assignment module in the project execution stage, the coding block corresponding to the result evaluation module in the project acceptance stage, etc.

[0052] Finally, extract the plurality of compressed storage coding blocks to generate a plurality of resource sub-blocks. Each resource sub-block has a clear scope definition, identified by data span information. For example, the data span information of a certain resource sub-block indicates that it contains the content from line 500 to line 1000 of the project document. This data span information helps to determine the processing order and storage location, etc. of the resource sub-blocks in subsequent compression operations.

[0053] For the case where the resource data to be compressed is an independently encoded document, taking an image file of a product promotion poster within an enterprise as an example, this image file is an independently encoded document. First, perform a disassembling operation on this resource data to be compressed. This image file adopts a specific image encoding method during storage, such as JPEG encoding. During the disassembling process, multiple compressed storage encoding blocks will be generated according to the internal structure of the encoding. These encoding blocks may be related to the storage method of the pixel area or color information of the image. For example, the image is divided according to a certain matrix, and the encoding information corresponding to each small matrix area is a compressed storage encoding block.

[0054] Then, extract these multiple compressed storage encoding blocks to generate multiple resource chunks. Each resource chunk also contains data span information. For an image file, this data span information can be represented as a specific area in the image. For example, the data span information of a certain resource chunk may indicate that it corresponds to the pixel area in the image from the upper left coordinate (0, 0) to the lower right coordinate (100, 100). This way of dividing resource chunks helps to adopt different compression strategies according to the image characteristics of different regions during the compression process, and also facilitates the management and operation of the image during storage and transmission.

[0055] Through the above processing methods for the resource data to be compressed when it is an integrated encoding document and an independent encoding document, the Internet data resources can be effectively and reasonably chunked, providing a basis for subsequent compression operations, and ensuring that the characteristics of resource chunks can be fully utilized during the compression process of different types of documents, improving the compression efficiency and the convenience of resource management.

[0056] In a possible implementation manner, when the resource data to be compressed includes a first encoding document and a second encoding document, step S110 may further include:

[0057] Step S111, extract the first encoding document to determine multiple first resource chunks.

[0058] Step S112, extract the second encoding document to determine multiple second resource chunks.

[0059] Step S113, for each first resource chunk, fuse the first resource chunk and the second resource chunk corresponding to the first resource chunk to generate the resource chunk.

[0060] In a possible implementation manner, step S113 includes:

[0061] Step S1131, when the first data span information in the first resource chunk matches the second data span information in the second resource chunk, fuse the first resource chunk and the second resource chunk to generate the resource chunk.

[0062] Step S1132: When the first data span information in the first resource block is different from the second data span information in the second resource block and there is an overlapping part, decompose the first resource block and the second resource block according to the first front-end position information and the first end position information in the first data span information and the second front-end position information and the second end position information in the second data span information, to generate multiple data resource sub-blocks.

[0063] Step S1133: Fuse two data resource sub-blocks with matching data span information among the multiple data resource sub-blocks to generate a fused data resource sub-block.

[0064] Step S1134: Fuse the data resource sub-blocks with different data span information among the multiple data resource sub-blocks and the fused data resource sub-block according to the resource transfer order to generate the resource block.

[0065] In this embodiment, taking a product R & D project of an enterprise as an example, the first coded document is the technical specification in the product R & D process, and the second coded document is the test report document based on this technical specification. First, extract the first coded document to determine multiple first resource blocks. For the first coded document of the technical specification, since its content has a certain structure and logical hierarchy, the first resource blocks can be divided according to different technical modules or functional characteristics. For example, according to the technical specifications of different components of the product, such as the electronic component part, the mechanical structure part, the software function module part, etc., they are respectively determined as different first resource blocks. Each first resource block has its corresponding data span information, and this data span information clarifies the position range of this resource block in the entire technical specification. For example, the data span information of the first resource block of the electronic component part may indicate the content from line 100 to line 300 of the document.

[0066] Next, extract the second coded document to determine multiple second resource blocks. For the second coded document of the test report, divide the second resource blocks according to different test stages or test objects. For example, according to the tests on electronic components, mechanical structures, software functions, etc., they are respectively determined as different second resource blocks. Similarly, each second resource block also has its corresponding data span information. Taking the second resource block of the test on electronic components as an example, its data span information may indicate the content from line 50 to line 150 of the test report document.

[0067] For each first resource block, it is necessary to fuse the first resource block and the corresponding second resource block to generate a resource block. When the first data span information in the first resource block matches the second data span information in the second resource block, the first resource block and the second resource block are fused to generate a resource block. For example, the first resource block is the software function module part in the technical specification manual, and its data span information indicates that it contains the detailed specifications of a specific software function, from line 800 to line 1000 of the document; while the corresponding second resource block is the test part for this specific software function in the test report document, and its data span information is from line 100 to line 300 of the document, and the software functions involved in both are logically completely corresponding, that is, the data span information matches. At this time, these two resource blocks are directly fused to generate a new resource block. This new resource block contains the complete information of this software function from specification definition to test results, and its data span information combines the data span information ranges of the two original resource blocks. All relevant information about this software function, whether it is specification or test-related content, can be clearly found in the new resource block.

[0068] When the first data span information in the first resource block is different from the second data span information in the second resource block and there is an overlapping part, based on the first front-end position information and the first end position information in the first data span information and the second front-end position information and the second end position information in the second data span information, the first resource block and the second resource block are decomposed to generate multiple data resource sub-blocks. For example, the first resource block is the mechanical structure part in the technical specification, and its first data span information indicates from line 400 to line 600 of the document, which includes the framework of the overall mechanical structure of the product and the specifications of key components; while the corresponding second resource block is the mechanical structure test part in the test report document, and its second data span information is from line 200 to line 400 of the document, and there is an overlapping part here because the mechanical structure test part in the test report may only test some of the mechanical structure content in the technical specification. Assume that the first front-end position information is line 400 (the starting line of the first resource block), the first end position information is line 600 (the ending line of the first resource block), the second front-end position information is line 200 (the starting line of the second resource block), and the second end position information is line 400 (the ending line of the second resource block). In this case, decomposition is performed according to relevant rules. If the first front-end position information precedes the second front-end position information and the first end position information precedes the second end position information, based on the first front-end position information and the second end position information, the first data resource sub-block is decomposed from the first resource block, and the second data resource sub-block is decomposed from the second resource block. For example, the first data resource sub-block containing the content from line 400 to line 400 is decomposed from the first resource block (mechanical structure specification part), and the second data resource sub-block containing the content from line 200 to line 400 is decomposed from the second resource block (mechanical structure test part). Based on the first end position information and the second end position information, the third data resource sub-block is decomposed from the first resource block, that is, the third data resource sub-block containing the content from line 400 to line 600 is decomposed from the first resource block. Based on the first front-end position information and the second front-end position information, the fourth data resource sub-block is decomposed from the second resource block, that is, the fourth data resource sub-block containing the content from line 200 to line 200 is decomposed from the second resource block.

[0069] Fuse two data resource sub-blocks with matching data span information among multiple data resource sub-blocks to generate a fused data resource sub-block. For example, the first data resource sub-block and the fourth data resource sub-block may be relevant in content, such as both involving the relevant part of the specification and test of a specific component in the mechanical structure, then these two data resource sub-blocks are fused to generate a fused data resource sub-block.

[0070] For data resource sub - chunks with different data span information among multiple data resource sub - chunks, the data resource sub - chunks and the fused data resource sub - chunks are fused according to the resource transfer order to generate resource chunks. The resource transfer order can be determined according to the logical order of the document content or the processing order of the data. For example, in the order of the mechanical structure from the overall framework to the specific components, first fuse the fused data resource sub - chunk with the third data resource sub - chunk, and then fuse it with the second data resource sub - chunk to finally generate a complete resource chunk. This resource chunk contains comprehensive information about the mechanical structure part from specifications to tests and is integrated in a reasonable logical order, facilitating subsequent compression operations as well as data management and utilization. In this way, when processing the resource data to be compressed containing the first encoded document and the second encoded document, the resource chunks can be accurately determined according to the actual situation of the document content, improving the efficiency and accuracy of data processing.

[0071] In a possible implementation, when the first front - end position information precedes the second front - end position information and the first tail - end position information precedes the second tail - end position information.

[0072] Step S1132 may include:

[0073] According to the first front - end position information and the second tail - end position information, decompose and obtain the first data resource sub - chunk from the first resource chunk, and decompose and obtain the second data resource sub - chunk from the second resource chunk.

[0074] According to the first tail - end position information and the second tail - end position information, decompose and obtain the third data resource sub - chunk from the first resource chunk.

[0075] According to the first front - end position information and the second front - end position information, decompose and obtain the fourth data resource sub - chunk from the second resource chunk.

[0076] Or, when the first front - end position information precedes the second front - end position information and the first tail - end position information follows the second tail - end position information.

[0077] Step S1132 may also include:

[0078] According to the first front - end position information and the second front - end position information, decompose and obtain the fifth data resource sub - chunk from the second resource chunk.

[0079] According to the first front - end position information and the first tail - end position information, decompose and obtain the sixth data resource sub - chunk from the second resource chunk.

[0080] Decompose and obtain a seventh data resource sub-block from the second resource sub-block according to the first tail position information and the second tail position information.

[0081] Output the first resource sub-block as an eighth data resource sub-block.

[0082] Or, when the first front-end position information precedes the second front-end position information and the first tail position information is the same as the second tail position information, or when the first front-end position information is the same as the second front-end position information and the first tail position information follows the second tail position information.

[0083] Step S1132 may further include:

[0084] Decompose the second resource sub-block according to the first front-end position information and the first tail position information to generate a ninth data resource sub-block and a tenth data resource sub-block.

[0085] Output the first resource sub-block as an eleventh data resource sub-block.

[0086] In this embodiment, in the product R & D and management scenario of an enterprise, relevant operations are further elaborated based on the technical specification manual (the first coded document) and the test report document (the second coded document) in the previously mentioned product R & D project.

[0087] When the first front-end position information precedes the second front-end position information and the first tail position information precedes the second tail position information, take the relevant part of a certain function module (such as the power management module) in the technical specification manual as the first resource sub-block, and its first data span information ranges from line 200 to line 300 of the document; take the relevant part of the power management module test in the test report document as the second resource sub-block, and its second data span information ranges from line 350 to line 500 of the document.

[0088] Decompose and obtain a first data resource sub-block from the first resource sub-block and a second data resource sub-block from the second resource sub-block according to the first front-end position information and the second tail position information. For the first resource sub-block, since the first front-end position information is line 200 and the second tail position information is line 500, then decompose the content from line 200 to line 300 (because the overall range of the first resource sub-block is 200 - 300 lines) from the first resource sub-block as the first data resource sub-block; for the second resource sub-block, decompose the content from line 350 to line 500 as the second data resource sub-block from the range of line 350 to line 500.

[0089] Decompose and obtain the third data resource sub-block from the first resource block according to the first tail position information and the second tail position information. Here, the first tail position information is line 300, and the second tail position information is line 500. Therefore, the content from line 200 to line 300 (the overall range of the first resource block) is decomposed from the first resource block as the third data resource sub-block.

[0090] Decompose and obtain the fourth data resource sub-block from the second resource block according to the first front position information and the second front position information. Since the first front position information is line 200 and the second front position information is line 350, the content from line 350 to line 350 is decomposed from the second resource block as the fourth data resource sub-block.

[0091] When the first front position information precedes the second front position information and the first tail position information follows the second tail position information, assume that the part related to the cooling system in the technical specification is the first resource block, and its first data span information ranges from line 400 to line 600 of the document; the part related to the cooling system test in the test report document is the second resource block, and its second data span information ranges from line 450 to line 550 of the document.

[0092] Decompose and obtain the fifth data resource sub-block from the second resource block according to the first front position information and the second front position information. The first front position information is line 400, and the second front position information is line 450. The content from line 450 to line 450 is decomposed from the second resource block as the fifth data resource sub-block.

[0093] Decompose and obtain the sixth data resource sub-block from the second resource block according to the first front position information and the first tail position information. The first front position information is line 400, and the first tail position information is line 600. Therefore, the content from line 450 to line 550 (the overall range of the second resource block) is decomposed from the second resource block as the sixth data resource sub-block.

[0094] Decompose and obtain the seventh data resource sub-block from the second resource block according to the first tail position information and the second tail position information. The first tail position information is line 600, and the second tail position information is line 550. The content from line 550 to line 550 is decomposed from the second resource block as the seventh data resource sub-block.

[0095] Output the first resource block as the eighth data resource sub-block, that is, directly use the part related to the cooling system (the first resource block, from line 400 to line 600 of the document) as the eighth data resource sub-block.

[0096] When the first front-end position information precedes the second front-end position information and the first end position information is the same as the second end position information, or when the first front-end position information is the same as the second front-end position information and the first end position information follows the second end position information. For example, the part related to structural strength in the technical specification is the first resource block, and its first data span information ranges from line 500 to line 700 of the document; the part related to structural strength testing in the test report document is the second resource block, and its second data span information ranges from line 500 to line 700 of the document.

[0097] Decompose the second resource block according to the first front-end position information and the first end position information to generate a ninth data resource sub-block and a tenth data resource sub-block. The first front-end position information is line 500, and the first end position information is line 700. Decompose the second resource block (from line 500 to line 700) into the content from line 500 to line 600 as the ninth data resource sub-block and the content from line 600 to line 700 as the tenth data resource sub-block.

[0098] Output the first resource block as an eleventh data resource sub-block, that is, directly use the part related to structural strength (the first resource block, from line 500 to line 700 of the document) as the eleventh data resource sub-block.

[0099] Thus, it is possible to accurately decompose different first resource blocks and second resource blocks into multiple data resource sub-blocks under different data span relationships of the first resource block and the second resource block, so as to perform subsequent operations such as fusing the data resource sub-blocks, thereby realizing the effective management and processing of the entire resource block, improving the data integration efficiency, and laying a foundation for subsequent operations such as data compression. This precise decomposition method according to the data span information helps to maintain the integrity and logic of the data when dealing with complex document relationships, ensuring that the data resources can be reasonably organized and processed in different business scenarios.

[0100] In a possible implementation manner, the method further includes:

[0101] Step A110, when the compression progress is not within the interval of the data span information, determine the interval parameter between the front-end position information in the data span information and the compression progress.

[0102] Step A120, after an interval of the interval parameter, determine the compression stream feature according to the candidate resource block.

[0103] Step A130, apply the compression stream feature to the compression unit and load the compression unit onto the cloud storage medium.

[0104] Among them, step A130 may include:

[0105] Step A131: Determine the compatibility requirements of the compression unit for the compression stream features according to the type of the compression unit, and generate an adaptation analysis result.

[0106] Step A132: If the adaptation analysis result shows an incompatibility situation, convert the compression stream features according to the compression requirements of the compression unit to generate converted compression stream features.

[0107] Step A133: Perform an initialization operation on the compression unit, and load the compression stream features into the compression unit after the initialization operation, so that the compression unit compresses each resource content data block according to the internal defined processing logic, based on the storage configuration information and compression algorithm information in the implemented compression stream features, to generate compressed data processed by the compression unit.

[0108] Step A134: Query the status information of the cloud storage medium, where the status information includes the current available capacity of the cloud storage medium, the distribution of storage areas, and the performance parameters of the storage medium.

[0109] Step A135: Determine a storage strategy for storing the compressed data according to the type of the cloud storage medium, and preprocess the compressed data according to the storage strategy. The preprocessing includes chunking the compressed data, adding metadata to the compressed data, and encrypting the compressed data according to the storage strategy.

[0110] Step A136: Determine the storage location of the compressed data in the cloud storage medium according to the status information of the cloud storage medium, and store the preprocessed compressed data in the storage location of the cloud storage medium.

[0111] In this embodiment, in the scenario of product R & D and management of an enterprise, the description continues based on the situations of various documents (such as technical specification manuals, test report documents, etc.) in the previous product R & D projects during the compression process.

[0112] When compressing the document data related to the R & D of enterprise products, the relationship between the compression progress and the data span information of each resource block can be continuously monitored. When it is found that the compression progress is not within the interval of the data span information, it is necessary to determine the interval parameter between the front-end position information in the data span information and the compression progress. For example, when compressing the overall data including the technical specification manual and the test report document, assume that the data span information of a certain candidate resource block is from line 800 to line 1200 of the document, and the current compression progress is line 600. Then the front-end position information of this data span information is line 800, and the calculated interval parameter is 200 lines (800 - 600).

[0113] Continuing with the above example, after reaching the 800th line (i.e., the interval parameter with an interval of 200 lines), the compression stream characteristics are determined based on this candidate resource block. This candidate resource block contains various types of resource content data, such as the text description part in the technical specification, the chart data in the test report, etc., and each resource content data has its resource attributes. Assume that this candidate resource block also contains resource node information. For example, this resource block is related to the data of a specific functional module of the product (such as the communication module). First, determine the compression storage mapping partition based on the resource node information in the candidate resource block. For the data related to the communication module, a compression storage mapping partition dedicated to storing the compressed data of the communication module can be determined. Then, each resource content data in the candidate resource block is compressed into the compression storage mapping partition according to the resource attributes of the resource content data, generating the compression stream characteristics. For example, for the resource content data with text attributes such as the text description part, the Huffman coding algorithm in the lossless compression algorithm is selected for compression according to the characteristics of the text; for the resource content data with image attributes such as the chart data, a suitable image compression algorithm (such as the JPEG compression algorithm) is selected for compression. In this process, it is necessary to traverse all the resource content data in the candidate resource block. According to the resource attributes of each resource content data, the resource content data with the same or similar resource attributes are grouped into one group, generating a group of resource content data classified according to resource attributes. For each group of resource content data, analyze the resource attribute characteristics corresponding to this group of resource content data to determine the corresponding compression algorithm. For each group of resource content data, perform preprocessing operations according to the selected compression algorithm, generating a preprocessed group of resource content data and a list of compression algorithms matching each group of resource content data. For each preprocessed group of resource content data, use the corresponding compression algorithm for compression operations, generating a compressed resource content data block. Add marking information to each compressed resource content data block. The marking information includes the resource attribute identifier of the original resource content data and the sequence identifier in the original candidate resource block. Analyze the structural characteristics of the compression storage mapping partition. According to the resource attribute identifier in the marking information of each compressed resource content data block, determine the reference storage area of each compressed resource content data block in the compression storage mapping partition. In each reference storage area, determine the specific target storage location according to the sequence identifier or preset rules, so that each compressed resource content data block has a unique storage location in the implemented compression storage mapping partition. Finally, configure the storage of each compressed resource content data block according to its target storage location in the compression storage mapping partition, generating the compression stream characteristics.

[0114] The process of applying the compressed stream feature to the compression unit and loading the compression unit onto the cloud storage medium is as follows: First, according to the type of the compression unit, determine the compatibility requirements of the compression unit for the compressed stream feature, and generate an adaptation analysis result. Suppose the compression unit is a compression device customized for processing enterprise document data, which has specific requirements for compression algorithms, data formats, etc. If some compression algorithms or data storage configuration methods in the compressed stream feature do not exactly match the requirements of the compression unit, then the adaptation analysis result will show an incompatibility. If the adaptation analysis result shows an incompatibility, convert the compressed stream feature according to the compression requirements of the compression unit to generate a converted compressed stream feature. For example, if the compression algorithm required by the compression unit is a specific encoding format, and the algorithm in the compressed stream feature is another format, the compression algorithm needs to be converted to the format that meets the requirements of the compression unit.

[0115] Perform an initialization operation on the compression unit. This initialization operation may include setting some initial parameters, clearing the cache, etc. Then load the compressed stream feature into the compression unit after the initialization operation, so that the compression unit, according to the internal defined processing logic, based on the storage configuration information and compression algorithm information of the resource content data blocks in the implemented compressed stream feature, compresses each resource content data block to generate compressed data after being processed by the compression unit. For example, the compression unit finds the corresponding resource content data block according to the storage configuration information of the resource content data block, and then compresses it according to the compression algorithm information (such as Huffman coding algorithm or JPEG compression algorithm) to obtain the compressed data.

[0116] Then query the status information of the cloud storage medium. The cloud storage medium is a place for storing compressed data related to enterprise product R & D. Its status information includes the current available capacity of the cloud storage medium, the distribution of storage areas, and the performance parameters of the storage medium. For example, the cloud storage medium may be a cluster composed of multiple storage servers, with a current available capacity of 500TB, the storage areas are distributed in different data centers (such as Data Center 1 and Data Center 2 located in different geographical locations), and the performance parameters of the storage medium include read and write speeds, data transfer bandwidth, etc.

[0117] According to the type of the cloud storage medium, determine the storage strategy for storing the compressed data, and preprocess the compressed data according to the storage strategy. Suppose the cloud storage medium is based on object storage, the storage strategy may require splitting the compressed data into blocks, each block being 1MB in size, adding metadata to the compressed data. The metadata can include the source of the data (such as from a technical specification or a test report document), the compression time, the product module to which the data belongs (such as a communication module), etc., and encrypt the compressed data according to the storage strategy, using the AES encryption algorithm to encrypt the compressed data.

[0118] Determine the storage location of the compressed data in the cloud storage medium according to the status information of the cloud storage medium, and store the preprocessed compressed data in the storage location of the cloud storage medium. For example, if the current available capacity of Data Center 1 is large and the read / write speed is fast, and there is a partition in the storage area dedicated to storing product R & D related data, then the preprocessed compressed data is stored in this storage partition of Data Center 1. In this way, through a series of rigorous operation processes, it can be ensured that when the compression progress is not within the data span information interval, the data can still be effectively compressed and reasonably stored in the cloud storage medium, ensuring the integrity and efficiency of the enterprise product R & D document data management.

[0119] In a possible implementation manner, the resource block further includes resource node information, multiple resource content data, and resource attributes of each of the resource content data.

[0120] Step A120 includes:

[0121] Step A121, determine the compression storage mapping partition according to the resource node information in the candidate resource block.

[0122] Step A122, compress each of the resource content data in the candidate resource block into the compression storage mapping partition according to the resource attributes of the resource content data to generate the compression stream feature.

[0123] The loading of the compression unit onto the cloud storage medium includes:

[0124] Obtain the first address information and the first capacity information of the cloud storage medium.

[0125] Determine the second address information of the compression unit according to the first address information, the first capacity information, and the resource node information of the compression stream feature.

[0126] Load the cloud storage medium according to the first address information, and load the compression unit according to the second address information.

[0127] Among them, step A122 includes:

[0128] Step A1221, traverse all the resource content data in the candidate resource block, and group the resource content data with the same or similar resource attributes according to the resource attributes of each resource content data to generate a resource content data group classified by resource attributes.

[0129] Step A1222: For each resource content data group, analyze the resource attribute characteristics corresponding to the resource content data group to determine the corresponding compression algorithm. Specifically, for a resource content data group with text attributes, select the Huffman coding algorithm or the LZW algorithm in the lossless compression algorithms. For a resource content data group with image attributes, select the JPEG compression algorithm or the PNG compression algorithm. For a resource content data group with audio attributes, select the MP3 compression algorithm or the FLAC compression algorithm.

[0130] Step A1223: For each resource content data group, perform a preprocessing operation according to the selected compression algorithm to generate a preprocessed resource content data group and a list of compression algorithms matching each resource content data group.

[0131] Step A1224: For each preprocessed resource content data group, perform a compression operation using the corresponding compression algorithm to generate a compressed resource content data block.

[0132] Step A1225: Add marker information to each compressed resource content data block. The marker information includes the resource attribute identifier of the original resource content data and the sequence identifier in the original candidate resource block.

[0133] Step A1226: Analyze the structural characteristics of the compression storage mapping partition. According to the resource attribute identifier in the marker information of each compressed resource content data block, determine the reference storage area of each compressed resource content data block in the compression storage mapping partition. Within each reference storage area, determine the specific target storage location according to the sequence identifier or a preset rule, so that each compressed resource content data block has a unique storage location in the implemented compression storage mapping partition.

[0134] Step A1227: Configure the storage of each compressed resource content data block according to its target storage location in the compression storage mapping partition to generate the compressed stream feature.

[0135] In this embodiment, the description continues based on the data resource block operations related to various documents (such as technical specification manuals, test report documents, etc.) in the previously mentioned product R & D project.

[0136] For resource chunking, it includes resource node information, multiple resource content data, and the resource attributes of each resource content data. For example, in a certain document resource chunk in a product R & D project, the resource node information may identify that this chunk is a data set about a specific subsystem of the product (such as the power supply subsystem). The resource content data is the specific data content in this chunk, which may include text descriptions of the power supply subsystem, relevant circuit diagram images, some audio annotations about the working sound of the power supply, etc. Each resource content data has its corresponding resource attribute. For example, the text description has text attributes, the circuit diagram image has image attributes, and the audio annotation has audio attributes.

[0137] The process of determining the compression stream feature based on the candidate resource chunk is as follows:

[0138] First, determine the compression storage mapping partition according to the resource node information in the candidate resource chunk. Take the candidate resource chunk related to the power supply subsystem mentioned above. Since the resource node information indicates that this is the data of the power supply subsystem, a dedicated compression storage mapping partition is determined for this power supply subsystem. This partition will be specifically used to store the compressed data related to the power supply subsystem, and it is logically separated from the data of other subsystems, which helps the subsequent management and retrieval of data.

[0139] Then, compress each resource content data in the candidate resource chunk into the compression storage mapping partition according to the resource attributes of the resource content data to generate the compression stream feature. This process is relatively complex and involves multiple steps.

[0140] Traverse all the resource content data in the candidate resource chunk. According to the resource attributes of each resource content data, group the resource content data with the same or similar resource attributes into one group to generate the resource content data groups classified by resource attributes. For example, in the candidate resource chunk of this power supply subsystem, all the text description parts will be grouped into a resource content data group with text attributes, all the circuit diagram images will be grouped into a resource content data group with image attributes, and the audio annotations will be grouped into a resource content data group with audio attributes.

[0141] For each resource content data group, analyze the resource attribute characteristics corresponding to the resource content data group to determine the corresponding compression algorithm. For a resource content data group with text attributes, select the Huffman coding algorithm or the LZW algorithm in the lossless compression algorithms. In this example of the power supply subsystem, if the text description mainly consists of relatively regular data such as some specification parameters, the Huffman coding algorithm may be selected for compression. For a resource content data group with image attributes, select the JPEG compression algorithm or the PNG compression algorithm. If the circuit diagram image requires relatively high clarity and the color information is relatively simple, the PNG compression algorithm may be selected. For a resource content data group with audio attributes, select the MP3 compression algorithm or the FLAC compression algorithm. If the audio annotation is just some simple beeps, the MP3 compression algorithm may be sufficient to meet the requirements.

[0142] For each resource content data group, perform preprocessing operations according to the selected compression algorithm to generate a preprocessed resource content data group and a list of compression algorithms matching each resource content data group. Taking the resource content data group with image attributes as an example, the preprocessing operations may include operations such as adjusting the color mode of the image and cropping unnecessary blank areas, and then generate a preprocessed image resource content data group and a matching PNG compression algorithm list.

[0143] For each preprocessed resource content data group, perform compression operations using the corresponding compression algorithm to generate a compressed resource content data block. For the preprocessed image resource content data group, use the PNG compression algorithm for compression to obtain a compressed image resource content data block.

[0144] Add marker information to each compressed resource content data block. The marker information includes the resource attribute identifier of the original resource content data and the sequential identifier in the original candidate resource block. For example, for the compressed image resource content data block, the resource attribute identifier in the marker information is image, and the sequential identifier may be a number determined according to the appearance order in the original candidate resource block, such as 3, indicating that this is the third resource content data processed.

[0145] Analyze the structural characteristics of the compressed storage mapping partition. According to the resource attribute identifier in the marker information of each compressed resource content data block, determine the reference storage area of each compressed resource content data block in the compressed storage mapping partition. For example, in the compressed storage mapping partition, there may be areas dedicated to storing image data, text data, and audio data. If the resource attribute identifier is for an image, then determine that the compressed image resource content data block is in the reference storage area for image data. Within each reference storage area, determine the specific target storage location according to the sequence identifier or preset rules, so that each compressed resource content data block has a unique storage location in the implemented compressed storage mapping partition. If in the reference storage area for image data, they are stored in ascending order of the sequence identifier, then the compressed image resource content data block with a sequence identifier of 3 will be stored in the corresponding location.

[0146] Finally, configure the storage of each compressed resource content data block according to its target storage location in the compressed storage mapping partition to generate the compressed stream feature. In this way, the entire compressed stream feature contains information about all resource content data blocks compressed and stored according to specific rules, and this information is of great significance for subsequent operations in the compression unit and the cloud storage medium.

[0147] The process of loading the compression unit onto the cloud storage medium is as follows:

[0148] First, obtain the first address information and the first capacity information of the cloud storage medium. In an enterprise's cloud storage system, the cloud storage medium may be a storage cluster composed of multiple servers. The first address information may be a network address, such as 192.168.100.10, which is used to locate the cloud storage medium. The first capacity information may be 500TB, indicating the total current storage capacity size of the cloud storage medium.

[0149] Based on the first address information, the first capacity information, and the resource node information of the compressed stream feature, determine the second address information of the compression unit. Suppose the resource node information of the compressed stream feature indicates data of the power subsystem, and according to its storage policy, the cloud storage medium has a dedicated storage area allocated for data of the power subsystem. Based on the first address information, the first capacity information, and this storage policy, determine the second address information of the compression unit, such as 192.168.100.10:8080 / power_subsystem, indicating the specific address related to the power subsystem of the compression unit in the cloud storage medium.

[0150] Load the cloud storage medium according to the first address information and load the compression unit according to the second address information. Through the network connection, access and load the cloud storage medium according to the first address information of 192.168.100.10, and at the same time load the compression unit according to the second address information of 192.168.100.10:8080 / power_subsystem, ensuring that the compression unit can perform data compression and storage operations at the correct position in the cloud storage medium, so as to realize the effective management and storage of the entire data resource, and ensure the accuracy, efficiency and orderliness of the data in operations such as compression and storage during the enterprise product R & D process.

[0151] In a possible implementation manner, the first resource block corresponds to a first resource type, and the second resource block corresponds to a second resource type. The determining the compression flow characteristics according to the candidate resource blocks includes:

[0152] Determine an instruction based on the resource type, and obtain the reference resource type of the resource data to be compressed in the Internet data resource.

[0153] When the reference resource type is the first resource type, determine the compression flow characteristics according to the first resource block in the candidate resource blocks.

[0154] When the reference resource type is the second resource type, determine the compression flow characteristics according to the second resource block in the candidate resource blocks.

[0155] When the reference resource type is the first resource type and the second resource type, determine the compression flow characteristics according to the candidate resource blocks.

[0156] In this embodiment, it is assumed that the first resource block corresponds to the first resource type of text type, and the second resource block corresponds to the second resource type of image type. These two resource types have different uses and characteristics in the documents during the product R & D process, and it is necessary to determine the compression flow characteristics according to different situations.

[0157] Determine an instruction based on the resource type, and obtain the reference resource type of the resource data to be compressed in the Internet data resource. In the enterprise information management system, the resource type determination instruction may be issued by a dedicated resource management module. This module will scan and analyze the Internet data resources related to product R & D to determine the reference resource type of the resource data to be compressed. For example, for a document containing product design sketches (images) and design descriptions (text), the resource management module will determine the reference resource type of the resource data to be compressed by parsing the document structure and data format.

[0158] When the reference resource type is the first resource type (i.e., text type), the compression stream characteristics are determined based on the first resource block in the candidate resource blocks. For example, in the technical specification manual during product R & D, a large part of the content is in text form. Suppose this document is divided into multiple resource blocks, and a certain candidate resource block contains the part describing the product's performance parameters, which belongs to the first resource block of the text type.

[0159] First, this first resource block contains resource node information, which may indicate that this text part is a description of a certain functional module of the product (such as the processor performance). The resource content data is the specific performance parameter text, and its resource attribute is the text attribute. The compression storage mapping partition is determined based on the resource node information. Since it is text content related to processor performance, a compression storage mapping partition dedicated to storing processor performance text data can be determined.

[0160] Then, traverse all the resource content data in this first resource block. Since they are all text attributes, the entire resource content data forms a resource content data group. Analyze the resource attribute characteristics corresponding to this resource content data group. Since it is of the text type, the Huffman coding algorithm in the lossless compression algorithm is selected. Then, preprocessing operations are performed, which may include removing spaces in the text, standardizing some special symbols, etc., to generate the preprocessed resource content data group and the corresponding Huffman coding algorithm list.

[0161] The Huffman coding algorithm is used to compress the preprocessed resource content data group to generate the compressed resource content data block. Marking information is added to this compressed resource content data block. The marking information includes the resource attribute identifier (text) of the original resource content data and the sequence identifier in the original candidate resource block (assumed to be 1, indicating that this is the first resource content data in the candidate resource block).

[0162] Analyze the structural characteristics of the compression storage mapping partition. According to the resource attribute identifier (text) in the marking information, determine the reference storage area (text data storage area) of the compressed resource content data block in the compression storage mapping partition. Within this reference storage area, determine the specific target storage location according to the sequence identifier (1). Finally, configure the storage of the compressed resource content data block according to this target storage location, thereby determining the compression stream characteristics based on this first resource block.

[0163] When the reference resource type is the second resource type (i.e., image type), the compression stream characteristics are determined based on the second resource block in the candidate resource blocks. For example, in the product's appearance design document, there is a design sketch image of the product's appearance, and this part of the image data is divided into the second resource block in a candidate resource block.

[0164] The resource node information of this second resource block may indicate that this is a design sketch of the front appearance of the product. The resource content data is the pixel data of the image and other content, and its resource attribute is the image attribute. Determine the compression storage mapping partition based on the resource node information, and determine a dedicated compression storage mapping partition for the design sketch of the front appearance of the product.

[0165] Traverse the resource content data in this second resource block. The entire image resource content data forms a resource content data group. Analyze the resource attribute characteristics corresponding to this resource content data group. Since it is of the image type, select the JPEG compression algorithm. Perform preprocessing operations, such as adjusting the resolution and color mode of the image, to generate the preprocessed resource content data group and the corresponding JPEG compression algorithm list.

[0166] Use the JPEG compression algorithm to compress the preprocessed resource content data group to generate the compressed resource content data block. Add marker information, which includes the resource attribute identifier (image) of the original resource content data and the sequence identifier in the original candidate resource block (assumed to be 1).

[0167] Analyze the structural characteristics of the compression storage mapping partition. According to the resource attribute identifier (image) in the marker information, determine the reference storage area (image data storage area) of the compressed resource content data block in the compression storage mapping partition. Within this reference storage area, determine the specific target storage location according to the sequence identifier (1). Finally, configure the storage of the compressed resource content data block according to the target storage location, so that the compression stream characteristics are determined based on this second resource block.

[0168] When the reference resource types are the first resource type and the second resource type, determine the compression stream characteristics based on the candidate resource block. For example, in a comprehensive design document of a product, there are both text descriptions of the internal structure of the product (the first resource type - text type) and schematic diagrams (the second resource type - image type), and these contents form a candidate resource block.

[0169] The resource node information of this candidate resource block may indicate that this is the content of the overall structure part of the product. For the text part in the resource content data, its resource attribute is the text attribute, and the resource attribute of the image part is the image attribute. Determine the compression storage mapping partition based on the resource node information, and determine a unified compression storage mapping partition for the overall structure part of the product.

[0170] Traverse all the resource content data in the candidate resource block, group the text resource content data into a resource content data group with text attributes, and group the image resource content data into a resource content data group with image attributes.

[0171] For the text resource content data group, analyze its resource attribute characteristics, and select the LZW algorithm in the lossless compression algorithms. Perform preprocessing operations such as text format standardization, etc., to generate a preprocessed resource content data group and a matching list of LZW algorithms. Use the LZW algorithm for compression operation to generate a compressed text resource content data block, and add marker information (resource attribute identifier is text, sequence identifier is assumed to be 1).

[0172] For the image resource content data group, analyze its resource attribute characteristics, and select the PNG compression algorithm. Perform preprocessing operations such as color optimization, etc., to generate a preprocessed resource content data group and a matching list of PNG compression algorithms. Use the PNG compression algorithm for compression operation to generate a compressed image resource content data block, and add marker information (resource attribute identifier is image, sequence identifier is assumed to be 2).

[0173] Analyze the structural characteristics of the compressed storage mapping partition. According to the resource attribute identifier (text) in the marker information of the compressed text resource content data block, determine its reference storage area (text data storage area) in the compressed storage mapping partition, and determine the specific target storage location according to the sequence identifier (1); according to the resource attribute identifier (image) in the marker information of the compressed image resource content data block, determine its reference storage area (image data storage area) in the compressed storage mapping partition, and determine the specific target storage location according to the sequence identifier (2).

[0174] Finally, configure the storage of the compressed text resource content data block and the compressed image resource content data block according to their respective target storage locations in the compressed storage mapping partition, so as to determine the compression stream characteristics based on this candidate resource block containing the first resource type and the second resource type. By determining the compression stream characteristics in this way according to different reference resource types, effective compression processing can be performed on different types of data, improving the efficiency and accuracy of data compression, and facilitating the management and operation of data during subsequent storage, transmission, and use.

[0175] Figure 2 FIG. shows the hardware structure diagram of the Internet data compression storage system 100 provided by an embodiment of the present invention, as Figure 2 shown, the Internet data compression storage system 100 may include a processor 110, a machine-readable storage medium 120, a bus 130, and a communication unit 140.

[0176] The machine-readable storage medium 120 may store data and / or instructions. In some embodiments, the machine-readable storage medium 120 may store data obtained from an external terminal. In some embodiments, the machine-readable storage medium 120 may store the data and / or instructions that the Internet data compression storage system 100 uses to execute or complete the exemplary methods described in the present invention.

[0177] In a specific implementation process, one or more processors 110 execute the computer-executable instructions stored in the machine-readable storage medium 120, so that the processors 110 can execute the Internet data compression storage method in the above method embodiments. The processors 110, the machine-readable storage medium 120, and the communication unit 140 are connected through the bus 130, and the processors 110 can be used to control the transceiver actions of the communication unit 140.

[0178] For the specific implementation process of the processors 110, reference may be made to the various method embodiments executed by the above Internet data compression storage system 100. Their implementation principles and technical effects are similar, and will not be elaborated herein.

[0179] In addition, an embodiment of the present invention further provides a readable storage medium, in which computer-executable instructions are preset. When the processor executes the computer-executable instructions, the above Internet data compression storage method is implemented.

[0180] It should be noted that, in order to simplify the description of the present invention disclosure and thus help the understanding of one or more embodiments of the invention, in the previous description of the embodiments of the present invention, sometimes multiple features are merged into one embodiment, drawing, or description thereof.

Claims

1. An Internet data compression and storage method, characterized in that, The method includes: Obtaining resource data to be compressed corresponding to Internet data resources, and determining a plurality of resource chunks according to the resource data to be compressed, each of the resource chunks including at least data span information; When compressing the Internet data resources, obtaining candidate resource chunks and obtaining the current compression progress of the Internet data resources; When the compression progress is within the interval of the data span information included in the candidate resource chunks, determining a compression stream feature according to the candidate resource chunks; Applying the compression stream feature to a compression unit and loading the compression unit onto a cloud storage medium for compressing and storing the Internet data resources; The resource chunks further include resource node information, a plurality of resource content data, and resource attributes of each of the resource content data; The determining the compression stream feature according to the candidate resource chunks includes: Determining a compression storage mapping partition according to the resource node information in the candidate resource chunks; Compressing each of the resource content data in the candidate resource chunks into the compression storage mapping partition according to the resource attributes of the resource content data to generate the compression stream feature; The loading the compression unit onto the cloud storage medium includes: Obtaining first address information and first capacity information of the cloud storage medium; Determining second address information of the compression unit according to the first address information, the first capacity information, and the resource node information of the compression stream feature; Loading the cloud storage medium according to the first address information and loading the compression unit according to the second address information.

2. The Internet data compression and storage method according to claim 1, wherein When the resource data to be compressed is an integrated coding document, the obtaining the resource data to be compressed corresponding to the Internet data resources and determining a plurality of resource chunks according to the resource data to be compressed includes: Obtaining a resource file corresponding to the Internet data resources; Decomposing the resource file to generate structured attribute data, unstructured attribute data, and resource data to be compressed of the Internet data resources, the resource data to be compressed including a plurality of compression storage coding blocks; Extracting the plurality of compression storage coding blocks to generate a plurality of resource chunks.

3. The Internet data compression and storage method according to claim 1, characterized in that When the resource data to be compressed is an independent coding document, the obtaining the resource data to be compressed corresponding to the Internet data resources and determining a plurality of resource chunks according to the resource data to be compressed includes: Decomposing the resource data to be compressed to generate a plurality of compression storage coding blocks; Extracting the plurality of compression storage coding blocks to generate a plurality of resource chunks.

4. The Internet data compression and storage method according to claim 1, wherein When the resource data to be compressed includes a first coding document and a second coding document, the determining a plurality of resource chunks according to the resource data to be compressed includes: Extracting the first coding document to determine a plurality of first resource chunks; Extracting the second coding document to determine a plurality of second resource chunks; For each first resource chunk, fusing the first resource chunk and the second resource chunk corresponding to the first resource chunk to generate the resource chunk.

5. The Internet data compression and storage method according to claim 4, wherein Fusing the first resource block and the corresponding second resource block to generate the resource block includes: When the first data span information in the first resource block matches the second data span information in the second resource block, fusing the first resource block and the second resource block to generate the resource block; When the first data span information in the first resource block is different from the second data span information in the second resource block and there is an overlapping part, decomposing the first resource block and the second resource block according to the first front-end position information and the first end position information in the first data span information and the second front-end position information and the second end position information in the second data span information to generate a plurality of data resource sub-blocks; Fusing two data resource sub-blocks with matching data span information in the plurality of data resource sub-blocks to generate a fused data resource sub-block; Fusing the data resource sub-blocks with different data span information in the plurality of data resource sub-blocks and the fused data resource sub-block according to the resource transfer order to generate the resource block.

6. The Internet data compression and storage method according to claim 5, wherein When the first front-end position information precedes the second front-end position information and the first end position information precedes the second end position information; The decomposing the first resource block and the second resource block according to the first front-end position information and the first end position information in the first data span information and the second front-end position information and the second end position information in the second data span information to generate a plurality of data resource sub-blocks includes: Decomposing to obtain a first data resource sub-block from the first resource block and a second data resource sub-block from the second resource block according to the first front-end position information and the second end position information; Decomposing to obtain a third data resource sub-block from the first resource block according to the first end position information and the second end position information; Decomposing to obtain a fourth data resource sub-block from the second resource block according to the first front-end position information and the second front-end position information; Or, when the first front-end position information precedes the second front-end position information and the first end position information follows the second end position information; The decomposing the first resource block and the second resource block according to the first front-end position information and the first end position information in the first data span information and the second front-end position information and the second end position information in the second data span information to generate a plurality of data resource sub-blocks includes: Decomposing to obtain a fifth data resource sub-block from the second resource block according to the first front-end position information and the second front-end position information; Decomposing to obtain a sixth data resource sub-block from the second resource block according to the first front-end position information and the first end position information; Decompose and obtain a seventh data resource sub-block from the second resource sub-block according to the first tail end position information and the second tail end position information; Output the first resource sub-block as an eighth data resource sub-block; Or, when the first front end position information precedes the second front end position information and the first tail end position information is the same as the second tail end position information, or when the first front end position information is the same as the second front end position information and the first tail end position information follows the second tail end position information; Decompose the first resource sub-block and the second resource sub-block according to the first front end position information and the first tail end position information in the first data span information and the second front end position information and the second tail end position information in the second data span information to generate a plurality of data resource sub-blocks, including: Decompose the second resource sub-block according to the first front end position information and the first tail end position information to generate a ninth data resource sub-block and a tenth data resource sub-block; Output the first resource sub-block as an eleventh data resource sub-block.

7. The Internet data compression and storage method according to claim 1, characterized in that The method further includes: When the compression progress is not within the interval of the data span information, determine the interval parameter between the front end position information in the data span information and the compression progress; After an interval of the interval parameter, determine the compression stream feature according to the candidate resource sub-block; Apply the compression stream feature to a compression unit and load the compression unit onto a cloud storage medium; Among them, the step of applying the compression stream feature to a compression unit and loading the compression unit onto a cloud storage medium includes: According to the type of the compression unit, determine the compatibility requirement of the compression unit for the compression stream feature and generate an adaptation analysis result; If the adaptation analysis result shows an incompatibility situation, convert the compression stream feature according to the compression requirement of the compression unit to generate a converted compression stream feature; Perform an initialization operation on the compression unit and load the compression stream feature into the initialized compression unit, so that the compression unit compresses each resource content data block according to the internal defined processing logic based on the storage configuration information and the compression algorithm information in the implemented compression stream feature to generate compressed data after being processed by the compression unit; Query the status information of the cloud storage medium, where the status information includes the current available capacity of the cloud storage medium, the distribution of the storage area, and the performance parameters of the storage medium; According to the type of the cloud storage medium, determine the storage policy for storing the compressed data, and preprocess the compressed data according to the storage policy. The preprocessing includes splitting the compressed data, adding metadata to the compressed data, and encrypting the compressed data according to the storage policy; Determine the storage location of the compressed data in the cloud storage medium according to the status information of the cloud storage medium, and store the preprocessed compressed data in the storage location of the cloud storage medium.

8. The Internet data compression and storage method according to any one of claims 1-7, characterized in that, Among them, The step of compressing each resource content data in the candidate resource chunks into the compressed storage mapping partition according to the resource attributes of the resource content data to generate the compressed stream feature includes: Traverse all resource content data in the candidate resource chunks, and group the resource content data with the same or similar resource attributes according to the resource attributes of each resource content data to generate a group of resource content data classified according to resource attributes; For each group of resource content data, analyze the resource attribute characteristics corresponding to the group of resource content data to determine the corresponding compression algorithm. Among them, for the group of resource content data with text attributes, select the Huffman coding algorithm or the LZW algorithm in the lossless compression algorithm; for the group of resource content data with image attributes, select the JPEG compression algorithm or the PNG compression algorithm; for the group of resource content data with audio attributes, select the MP3 compression algorithm or the FLAC compression algorithm; For each group of resource content data, perform preprocessing operations according to the selected compression algorithm to generate a preprocessed group of resource content data and a list of compression algorithms matching each group of resource content data; For each preprocessed group of resource content data, perform compression operations using the corresponding compression algorithm to generate compressed resource content data blocks; Add marking information to each compressed resource content data block, and the marking information includes the resource attribute identifier of the original resource content data and the sequence identifier in the original candidate resource chunks; Analyze the structural characteristics of the compressed storage mapping partition, and determine the reference storage area of each compressed resource content data block in the compressed storage mapping partition according to the resource attribute identifier in the marking information of each compressed resource content data block. In each reference storage area, determine the specific target storage location according to the sequence identifier or preset rules, so that each compressed resource content data block has a unique storage location in the implemented compressed storage mapping partition; Configure the storage of each compressed resource content data block according to its target storage location in the compressed storage mapping partition to generate the compressed stream feature.

9. The Internet data compression and storage method according to claim 4, characterized in that, The first resource chunk corresponds to the first resource type, and the second resource chunk corresponds to the second resource type; The determining of the compressed stream feature according to the candidate resource chunks includes: Based on the resource type determination instruction, obtain the reference resource type of the resource data to be compressed in the Internet data resource; When the reference resource type is the first resource type, determine the compressed stream feature according to the first resource chunk in the candidate resource chunks; When the reference resource type is the second resource type, determine the compressed stream feature according to the second resource chunk in the candidate resource chunks; When the reference resource type is the first resource type and the second resource type, determine the compressed stream feature according to the candidate resource chunks.

10. An Internet data compression and storage system, characterized in that, The Internet data compression and storage system includes a processor and a memory. The memory is connected to the processor. The memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the Internet data compression and storage method described in any one of claims 1-9 above.

Citation Information

Patent Citations

  • Data compression management method and device

    CN113742335A

  • File compression method and electronic equipment

    CN117708070A