Data storage method for cloud business intelligent big data service system
By type division and logical block segmentation of data in the cloud business intelligence big data service system, combined with deep compression network model and dynamic compression rules, the problems of low data storage efficiency and high access delay in the existing technology are solved, and more efficient data storage and access are achieved.
Patent Information
- Application Number
- CN202510003449.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-02
AI Technical Summary
The data storage technology of existing cloud business intelligence big data service systems is difficult to dynamically adapt to data characteristics, resulting in low storage efficiency and high access latency, and data compression technology cannot adapt to changes in data access frequency.
By dividing the input data into logical blocks by data type, a deep compression network model is designed for processing, and a compression rule with dynamic priority is set according to the data access frequency, and the storage medium is finally selected based on the compressed data.
It realizes the flexibility of data storage and improves resource utilization, significantly improves the system's storage performance and access efficiency, and balances storage space and access speed.
Smart Images

Figure CN119917023A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cloud business intelligence data storage, and in particular to a data storage method for a cloud business intelligence big data service system. Background Art
[0002] With the rapid development of big data and cloud computing technologies, business intelligence (BI) systems have become increasingly important in enterprise decision-making and operations. The system can provide users with valuable business insights and decision support by analyzing and processing massive amounts of data. However, with the continuous growth of data volume and the diversification of data types, traditional data storage technologies have been unable to meet the requirements of modern business intelligence systems in terms of efficiency, scalability, and stability. In addition, existing cloud big data service systems usually rely on centralized or distributed storage solutions, but such solutions lack sufficient flexibility in processing different data types, especially in the allocation and compression of hot data and cold data, which often leads to low storage efficiency and increased access latency. In addition, existing data compression technologies are mostly based on fixed rules and cannot be adaptively optimized for dynamic changes in data access frequency, which further limits the performance of the system. First, conventional storage methods do not fully consider the differences in data types, access frequencies, and characteristics of hot and cold data, resulting in difficulty in quickly extracting high-frequency access data, while low-frequency data occupies a large amount of storage resources. Second, data compression technology has significant deficiencies in balancing storage space and access speed. Fixed compression ratios and single storage medium selection strategies show obvious limitations in the face of dynamically changing usage scenarios.
[0003] Based on the above problems, how to design an efficient storage method that can dynamically adapt to data characteristics and improve storage utilization and access efficiency through intelligent compression technology has become a technical problem that needs to be solved urgently. Summary of the invention
[0004] The purpose of this section is to summarize some aspects of embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the specification abstract and the invention title of this application to avoid blurring the purpose of this section, the specification abstract and the invention title, and such simplifications or omissions cannot be used to limit the scope of the present invention.
[0005] In view of the above existing problems, the present invention is proposed. Therefore, the present invention provides a data storage method for a cloud-based business intelligence big data service system to solve the problems raised in the background technology.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: a data storage method for a cloud-based business intelligence big data service system, comprising:
[0007] The data input into the cloud business intelligence big data service system is divided into data types in different regions of the cloud node, and the data in each data type is segmented using data blocks;
[0008] Design a deep compression network model, input the segmented data blocks into the network encoder in the deep compression network model for processing, and set a dynamic priority compression rule based on the processed data to obtain compressed data;
[0009] The storage medium is selected based on the compressed data, thereby realizing data storage in the cloud-based business intelligence big data service system.
[0010] As a preferred solution of the data storage method for the cloud-based business intelligence big data service system described in the present invention, the data input into the cloud-based business intelligence big data service system is divided into data types, including:
[0011] Data is obtained from cloud nodes in various regions, and delayed grouped according to each cloud node. At the same time, based on the data type identifier and content characteristics of the data in each cloud node, the data types are divided into text type, video type and log type, and the access frequency of each data type is recorded.
[0012] As a preferred solution of the data storage method for the cloud-based business intelligence big data service system described in the present invention, the data in each type of data is divided into data blocks, including:
[0013] According to the access frequency of each data type, the data type is marked as hot data and cold data respectively;
[0014] The hot data and cold data are divided into a plurality of logic blocks, each logic block takes the current data type as a unit, and the size of the logic block of the hot data is set not to exceed the size of the logic block of the cold data.
[0015] As a preferred solution of the data storage method for the cloud-based business intelligence big data service system described in the present invention, a deep compression network model is designed, including:
[0016] Network encoder module, network decoder module, and feature extraction layers corresponding to data types.
[0017] As a preferred solution of the data storage method for the cloud-based business intelligence big data service system described in the present invention, the data of the segmented data blocks is input into the network encoder in the deep compression network model for processing, including:
[0018] The input layer in the network encoder calls the corresponding feature extraction layer according to the data type of the data block and extracts the data features of the hot data and the cold data, and deduplicates the extracted data features of the hot data and the cold data through multiple hidden layers, while reducing the size of the logical blocks in the hot data and the cold data.
[0019] As a preferred solution of the data storage method for the cloud-based business intelligence big data service system described in the present invention, wherein: a compression rule with a dynamic priority is set according to the processed data to obtain compressed data, including:
[0020] The processed hot data and cold data are restored to their data types through a decoder to regain the access frequency of the data type. By comparing the updated access frequency with the initial access frequency, the compression rate of the data under the data type can be reduced or increased.
[0021] As a preferred solution of the data storage method for the cloud-based business intelligence big data service system described in the present invention, wherein: by comparing the updated access frequency with the initial access frequency, selecting to reduce or increase the compression depth of the data under the data type includes:
[0022] If the updated access frequency is greater than the initial access frequency, the compression rate of the data under the data type is reduced;
[0023] If the updated access frequency is less than the initial access frequency, the compression rate of the data under this data type is improved.
[0024] As a preferred solution of the data storage method for the cloud-based business intelligence big data service system described in the present invention, wherein: selecting a storage medium based on compressed data includes:
[0025] If the compression rate of the compressed data is lower than 50%, it is preferentially stored in the optical storage medium, otherwise it is stored in the solid state drive.
[0026] Compared with the prior art, the invention has the following beneficial effects:
[0027] 1. The present invention divides the input data according to the data type and divides each type of data into logical blocks, thereby realizing hierarchical management and storage optimization of data, solving the problem of low storage efficiency caused by failure to flexibly divide and store data according to data characteristics in the background technology; hot data and cold data are processed differently by dividing the data into logical blocks, thus significantly improving the flexibility and resource utilization of the service system;
[0028] 2. By designing a deep compression network model, the segmented data logic blocks are input into the network encoder for processing. Combined with the priority compression rules, the problems of low data compression efficiency, fixed compression depth and high access delay in the background technology are solved; by adjusting the compression rate, the compression ratio of hot data is reduced to increase the access speed, and the compression ratio of cold data is increased to save storage space, effectively balancing the storage performance and access efficiency of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. Among them:
[0030] Figure 1 This is an overall flow chart of a data storage method for a cloud-based business intelligence big data service system according to an embodiment of the present invention. DETAILED DESCRIPTION
[0031] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in the art without creative work should fall within the scope of protection of the present invention.
[0032] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0033] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.
[0034] The present invention is described in detail with reference to schematic diagrams. When describing the embodiments of the present invention, for the sake of convenience, the cross-sectional diagrams showing the device structure will not be partially enlarged according to the general scale, and the schematic diagrams are only examples, which should not limit the scope of protection of the present invention. In addition, in actual production, the three-dimensional dimensions of length, width and depth should be included.
[0035] At the same time, in the description of the present invention, it should be noted that the directions or positional relationships indicated by the terms "upper, lower, inner and outer" are based on the directions or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore cannot be understood as limiting the present invention. In addition, the terms "first, second or third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0036] In the present invention, unless otherwise clearly specified and limited, the terms "install, connect, connect" should be understood in a broad sense, for example: it can be a fixed connection, a detachable connection or an integral connection; it can also be a mechanical connection, an electrical connection or a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0037] Example 1
[0038] Reference Figure 1 , which is the first embodiment of the present invention, and provides a data storage method for a cloud-based business intelligence big data service system, comprising:
[0039] S1. Classify the data input into the cloud business intelligence big data service system in each region’s cloud node by data type, and divide the data in each data type by data blocks;
[0040] Furthermore, data is obtained from cloud nodes in various regions, and delayed grouping is performed according to cloud nodes in various regions. At the same time, based on the data type identifier and content characteristics of the data in cloud nodes in various regions, the data types are divided into text type, video type and log type, and the access frequency of each data type is recorded;
[0041] It should be noted that grouping cloud nodes in different regions according to low latency and high latency can not only give priority to low-latency transmission in cloud nodes, but also share the load pressure of high-latency cloud nodes and maintain the stability of the service system;
[0042] Specifically, the data type identifier refers to the mark of the category to which the data belongs, that is, the external or intrinsic properties of the data, such as file extensions: .txt and .csv belong to the text type, .avi belongs to the video type, and .log belongs to the log type. File extensions belong to external properties, and intrinsic properties such as data formats such as JSON, XML, and HTML represent structured text data and belong to the text type;
[0043] Specifically, content features refer to the intrinsic characteristics of the data itself, which can be judged by the content of the data itself, such as the number of lines and characters represented by the text type, and the frame rate, resolution, bit rate, etc. represented by the video type;
[0044] Furthermore, according to the access frequency of each data type, the data type is marked as hot data and cold data;
[0045] Specifically, hot data refers to data that is frequently accessed, modified, or used, that is, data with high timeliness. The higher the timeliness, the higher the access frequency.
[0046] Specifically, cold data refers to data with low access frequency and low timeliness requirements. That is, the access frequency is low, and cold data is usually used as historical data for reference;
[0047] Furthermore, the hot data and the cold data are divided into multiple logical blocks, each logical block takes the current data type as a unit, and the size of the logical block of the hot data cannot exceed the size of the logical block of the cold data;
[0048] It needs to be explained that although the division of logic blocks can improve the overall processing efficiency of the subsequent model, it will correspondingly increase the processing volume of the model;
[0049] It should be noted that the logic block of cold data is set larger than that of hot data because the timeliness of cold data is not high and the access frequency is low. When the access frequency changes, the service system can get a quick response because the logic block of hot data is small and the logic block of cold data is large, thereby reducing the processing volume of the model.
[0050] S2. Design a deep compression network model, input the data of the segmented data blocks into the network encoder in the deep compression network model for processing, and set a compression rule with dynamic priority according to the processed data to obtain compressed data;
[0051] Specifically, the deep compression network model includes a network encoder module, a network decoder module, and a feature extraction layer corresponding to the data type;
[0052] Furthermore, the processing in the deep compression network model is as follows:
[0053] First, the input layer in the network encoder calls the corresponding feature extraction layer according to the data type of the logic block itself and extracts the data features of the hot data and cold data. Then the feature extraction layer removes the redundancy of the extracted data features of the hot data and cold data through multiple hidden layers, while reducing the size of the logic blocks in the hot data and cold data, and sends them to the network decoder.
[0054] Specifically, de-redundancy uses a low-rank decomposition method, including removing invalid information, that is, eliminating blank characters and useless fields in the data and merging duplicate entries;
[0055] It should be noted that since the redundancy of hot data and cold data has been removed, the corresponding logical block size will be compressed;
[0056] Furthermore, the processed hot data and cold data are restored to their data types through a network decoder to obtain the access frequency of the data type again. By comparing the updated access frequency with the initial access frequency, the compression rate of the data under the data type is reduced or increased.
[0057] It should be noted that the network decoder is mainly used to restore the encoded data, and in the restoration process, the structure and access records of the data are analyzed and the access frequency is recalculated;
[0058] Furthermore, if the updated access frequency is greater than the initial access frequency, the compression rate of the data under this data type (hotspot data) is reduced;
[0059] Furthermore, if the updated access frequency is less than the initial access frequency, the compression rate of the data (cold data) under this data type is improved;
[0060] S3, selecting a storage medium based on the compressed data, thereby realizing data storage in the cloud-based business intelligence big data service system;
[0061] Furthermore, if the compression rate of the compressed data is less than 50%, it is preferentially stored in the optical storage medium, otherwise it is stored in the solid state drive;
[0062] It should be noted that since commercial data usually involves privacy and security, cloud data is usually stored in local physical media, thereby effectively reducing the risk of data leakage caused by cloud service failure or hacker attacks.
[0063] Example 2
[0064] Referring to Table 1, a second embodiment of the present invention provides a data storage method for a cloud-based business intelligence big data service system, including: verifying the beneficial effects of the data storage solution of the present invention compared with the traditional solution through simulation experiments;
[0065] The simulation test environment is as follows:
[0066] Hardware configuration: Storage medium: Three storage devices are used, representing different solutions: Prior art A: Mechanical hard disk (HDD), capacity 1TB, rotation speed 7200RPM; Prior art B: Solid state drive (SSD), capacity 1TB, support NVMe protocol; Inventive method: Hybrid storage solution (SSD+optical storage), SSD is used for hot data, and optical storage is used for cold data;
[0067] Data server: Intel Xeon E5-2690 v4@2.6GHz, 64GB memory; Network configuration: Gigabit Ethernet;
[0068] Software and Algorithms:
[0069] Data compression tools: Prior art A uses a traditional compression algorithm (LZ77); Prior art B uses an improved dictionary compression method; the method of the present invention uses a deep compression network model and designs dynamic compression rules for different data types;
[0070] Data access management: Redis-based cache scheduling strategy, test access latency;
[0071] Dataset: The test data includes different types (text, video, log), with a total data volume of about 100GB, divided into hot data (20GB) and cold data (80GB);
[0072] Among them, the access mode simulates real application scenarios, including high-frequency reading and low-frequency access scenarios; data is divided into logical blocks by type, hot data is stored in high-performance media first, and cold data is compressed and stored in optical storage media; a network encoder is used to extract data features and process redundancy, dynamically adjust the data block size, and set different compression rules; each technical solution measures parameters such as storage time, access delay, storage cost, and data integrity. All tests are run multiple times under the same load to ensure data reliability. The data in Table 1 are as follows:
[0073] Table 1
[0074]
[0075] Among them, from the comparison of the table data, it can be clearly seen that the method of the present invention is superior to the prior art in the following aspects:
[0076] Data compression rate: The compression rate of the method of the present invention is 20%, which is significantly lower than that of the prior art A (30%) and the prior art B (40%), indicating that the method of the present invention achieves a more efficient data compression effect through a deep compression network model and a dynamic priority compression rule;
[0077] Storage time: In the storage process of a 100GB data set, the method of the present invention takes 8 seconds, which is only 53.3% of the prior art A and 66.7% of the prior art B, proving that it has higher storage efficiency after processing data segmentation, classification and compression;
[0078] Access delay: The access delay of the method of the present invention is 50ms, which is significantly lower than 150ms of prior art A and 120ms of prior art B, which is attributed to the dynamic storage mechanism and data priority access method of the present invention;
[0079] Storage cost: The storage cost per GB of the method of the present invention is RMB 0.3641, which is only half of that of the prior art A and 62.5% of that of the prior art B. This is because the cost of the storage device is reduced by compressing and storing the data on the optical storage medium.
[0080] Data integrity: The data integrity of the method of the present invention reaches 99%, which is much higher than the prior art A (92%) and prior art B (95%), indicating that it can still well retain the accuracy of data content while performing redundancy removal and compression processing;
[0081] Through the above data comparison, the method of the present invention shows advantages in data compression rate, storage time, access delay and storage cost, and data integrity indicators; it illustrates that the scheme of the present invention can make up for the shortcomings of the prior art in storage efficiency and performance optimization, and is suitable for large-scale data storage scenarios.
[0082] Those skilled in the art will appreciate that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the application can adopt the form of complete hardware embodiments, complete software embodiments, or embodiments in combination with software and hardware. Moreover, the application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program codes. The scheme in the embodiments of the present application can be implemented in various computer languages, for example, object-oriented programming language Java and literal scripting language JavaScript, etc.
[0083] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0084] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0085] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0086] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0087] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A data storage method for a cloud-based business intelligence big data service system, characterized in that: include: The data input into the cloud business intelligence big data service system is divided into data types in different regions of the cloud node, and the data in each data type is segmented using data blocks; Design a deep compression network model, input the segmented data blocks into the network encoder in the deep compression network model for processing, and set a dynamic priority compression rule based on the processed data to obtain compressed data; The storage medium is selected based on the compressed data, thereby realizing data storage in the cloud-based business intelligence big data service system.
2. The data storage method for a cloud-based business intelligence big data service system according to claim 1, characterized in that: The data input into the cloud business intelligence big data service system in each region of the cloud node is divided into data types, including: Data is obtained from cloud nodes in various regions, and delayed grouped according to each cloud node. At the same time, based on the data type identifier and content characteristics of the data in each cloud node, the data types are divided into text type, video type and log type, and the access frequency of each data type is recorded.
3. The data storage method for a cloud-based business intelligence big data service system according to claim 2, characterized in that: The data in each data type is segmented using data blocks, including: According to the access frequency of each data type, the data type is marked as hot data and cold data respectively; The hot data and cold data are divided into a plurality of logic blocks, each logic block takes the current data type as a unit, and the size of the logic block of the hot data is set not to exceed the size of the logic block of the cold data.
4. The data storage method for a cloud-based business intelligence big data service system according to claim 1, characterized in that: Design deep compression network models, including: Network encoder module, network decoder module, and feature extraction layers corresponding to data types.
5. The data storage method for a cloud-based business intelligence big data service system according to claim 3 or 4, characterized in that: The data of the segmented data blocks are input into the network encoder in the deep compression network model for processing, including: The input layer in the network encoder calls the corresponding feature extraction layer according to the data type of the data block and extracts the data features of the hot data and the cold data, and deduplicates the extracted data features of the hot data and the cold data through multiple hidden layers, while reducing the size of the logical blocks in the hot data and the cold data.
6. The data storage method for a cloud-based business intelligence big data service system according to claim 5, characterized in that: A compression rule with a dynamic priority is set according to the processed data to obtain compressed data, including: The processed hot data and cold data are restored to their data types through a decoder to regain the access frequency of the data type. By comparing the updated access frequency with the initial access frequency, the compression rate of the data under the data type can be reduced or increased.
7. The data storage method for a cloud-based business intelligence big data service system according to claim 6, characterized in that: By comparing the updated access frequency with the initial access frequency, you can choose to reduce or increase the compression depth of the data under this data type, including: If the updated access frequency is greater than the initial access frequency, the compression rate of the data under the data type is reduced; If the updated access frequency is less than the initial access frequency, the compression rate of the data under this data type is improved.
8. The data storage method for a cloud-based business intelligence big data service system according to claim 6, characterized in that: Select storage media based on the compressed data, including: If the compression rate of the compressed data is lower than 50%, it is preferentially stored in the optical storage medium, otherwise it is stored in the solid state drive.