Optimization processing system based on big data
By designing an optimization processing system based on big data, the problem of low storage efficiency caused by redundant data in the cloud platform is solved, efficient data compression and storage optimization are achieved, and user satisfaction and user experience are improved.
Patent Information
- Application Number
- CN202510129207.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-02-05
AI Technical Summary
The traditional stand-alone data storage model cannot meet the new needs brought by massive data. There is a large amount of redundant data in the cloud platform, which reduces the utilization rate of storage space resources and affects user satisfaction.
Design an optimization processing system based on big data, and control and manage cloud platform registration and data access, compress and backup cloud platform data, perform storage optimization analysis and control of cloud platform data, and review and display data information and store control management.
It improves the efficiency of data compression processing, reduces the storage of redundant data, improves the utilization rate of data storage space, and greatly improves user satisfaction and user experience.
Smart Images

Figure CN120200910A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data, and particularly to an optimization processing system based on big data. Background Art
[0002] In recent years, with the popularization of 5G technology and the emergence of new concepts such as smart cities, smart transportation, and smart homes, people use smart devices to view surrounding information in real time, conduct online communication and sharing, and network data has experienced explosive growth. The traditional single-machine data storage mode can no longer meet the new requirements brought by the huge and rapidly growing mass data in the system. The existing method is to migrate single-machine data to the cloud platform for storage. However, the applications deployed in the cloud platform have characteristics such as diversity and complexity, resulting in a large amount of redundant data in the cloud platform, reducing the utilization rate of storage space resources, and seriously affecting the user satisfaction. Therefore, it is necessary to design an optimization processing system based on big data with high storage efficiency and intelligent management. Summary of the Invention
[0003] The purpose of the present invention is to provide an optimization processing system based on big data to solve the problems raised in the above background art.
[0004] To solve the above technical problems, the present invention provides the following technical solutions: An optimization processing method based on big data, including: Controlling and managing cloud platform registration and data access; Controlling and processing compression and backup of cloud platform data; Performing storage optimization analysis control of cloud platform data; Consulting, displaying, and controlling and managing data information storage.
[0005] According to the above technical solution, the controlling and managing cloud platform registration and data access includes: After a user registers in the big data system facing the cloud platform, fills in identity information, passes the background review and logs in to the system, service control and management are performed.
[0006] According to the above technical solution, the controlling and processing compression and backup of cloud platform data includes: By setting a time window threshold and adopting a multi-threaded processing method for data compression control and management; Using a data replication file thread to simultaneously copy and backup multiple files, making each thread correspond to a data file, using a data storage copy to correspond to the data file, and having another thread write it into the target file.
[0007] According to the above technical solution, the performing storage optimization analysis control of cloud platform data includes: After setting the fixed window and sliding window thresholds, data block optimization management is performed on the cloud platform data file according to different low-entropy strings; After receiving the data block optimization processing result, further duplicate data detection processing is performed on it; After marking the data blocks determined to be duplicates, a warning notice is sent to the user for this result, waiting for the user to check and process.
[0008] According to the above technical solution, the access, display, and storage control management of data information includes: When the user views the warning notice, the display content can be viewed through conditional filtering. The displayed content includes the reason for the warning notice, the corresponding warning data, and the warning time; The user can manage the data blocks determined to be duplicates. Through the options after viewing the warning notice, the user can store or filter the data blocks determined to be duplicates.
[0009] According to the above technical solution, an optimization processing system based on big data includes: A data management module for controlling and managing cloud platform data; An analysis and processing module for performing storage optimization analysis and control of cloud platform data; A query and storage module for accessing, displaying, and storing control management of data information.
[0010] According to the above technical solution, the data management module includes: A registration and access module for managing the registration and access of users to the cloud platform; A data compression module for controlling and managing data compression; A backup management module for analyzing and controlling cloud platform data backups.
[0011] According to the above technical solution, the analysis and processing module includes: A data block module for optimizing data block processing; An analysis and determination module for analyzing and determining duplicate data; A warning notice module for managing warning notices to users.
[0012] According to the above technical solution, the query and storage module includes: A query and display module for querying and displaying data information; A storage management module for storing and managing data information.
[0013] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: By providing a data management module, an analysis and processing module, and a query and storage module, the present invention can make the compression processing of data more efficient and accurate, reduce the storage accumulation of redundant data, make more efficient use of the data storage space, and greatly improve the user satisfaction and usage experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The accompanying drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, but do not constitute a limitation to the present invention. In the drawings: Figure 1 FIG. is a flowchart of an optimization processing method based on big data provided in Embodiment 1 of the present invention; Figure 2 FIG. is a block diagram of the module composition of an optimization processing system based on big data provided in Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0016] Embodiment 1: Figure 1 FIG. is a flowchart of an optimization processing method based on big data provided in Embodiment 1 of the present invention. This embodiment can be applied to an optimization processing system, and this method can be executed by an optimization processing system provided in the embodiments of the present invention. The system consists of multiple software and hardware modules, such as Figure 1 shown, and the method specifically includes the following steps: S101. Control and manage the cloud platform registration and data access; Exemplarily, in an embodiment of the present invention, after a user registers in a big data system facing a cloud platform, fills in identity information, passes background review and logs in to the system, service control and management are performed; in this step, after the user logs in to the cloud platform, the user views the service types provided by the system through the front-end page. The user can view the database service configuration provided by the used cloud platform. When viewing the current system resource usage situation through the system monitoring page, the front-end of the cloud platform is controlled to send a request to the back-end. The back-end retrieves the system resource monitoring data, performs data encapsulation processing and returns it to the front-end for display. Moreover, the user can set the backup time of the database according to needs. After the setting is completed, the back-end will be notified to store the backup time in the database for recording, and then the message queue will be notified to perform regular backup of the database. Through this processing, the user can use the cloud platform more efficiently and conveniently, effectively improving the user's satisfaction.
[0017] S102. Control and process the compression and backup of cloud platform data; Exemplarily, in an embodiment of the present invention, by setting a time window threshold and adopting a multi-threaded processing method for data compression control and management; in this step, for data requests within a write-only instance, a time window is set to accumulate access requests, and then the data within its time window is sampled to detect its data content for data compression. Specifically, after setting the time window threshold and the expected data length threshold within the time window respectively, when within the threshold time range and the expected data length threshold is not met, the data within the time window is controlled not to be persistently stored, so that it is accumulated into a data file. Only when the data accumulation time is greater than the time window threshold or the data length within the time window is greater than the expected data length threshold, the sampler is controlled to detect the data within the time window, and then the CPU and memory of the current system workload are obtained through the system monitor for corresponding data compression processing. Among them, for the data within the time window, a multi-threaded method is used for data processing. When the number of currently active working threads exceeds the set threshold, the threads without work are put into the thread waiting queue to sleep, otherwise these threads are destroyed. Through this processing, the data compression processing can be made more efficient and accurate, effectively reducing the space required for data storage.
[0018] Use the data replication file thread to simultaneously copy and backup multiple files, and make each thread correspond to a data file. Use a data storage copy to correspond to the data file, and have another thread write it into the target file; effectively reduce the backup time of thread queuing writes.
[0019] S103. Perform storage optimization analysis and control of cloud platform data; Exemplarily, in the embodiments of the present invention, after setting the fixed window and the sliding window threshold, data block optimization management is performed on the data files in the cloud platform according to different low-entropy strings. In this step, the low-entropy strings in the data files are mainly in two modes, which are composed of the same strings and a series of repeated substrings respectively. After detecting and identifying that the current data file to be processed is composed of the same strings, when the maximum value of the data in the fixed window is continuously the same as the maximum value of the data in the sliding window, and the sum of the lengths of the fixed window and the sliding window is not less than the set threshold, the data within the window is used as the data block point; otherwise, the sliding window is continued to search for the data block point byte by byte backward. After detecting and identifying that the current data file to be processed is composed of a series of repeated substrings, the data in the fixed window and the sliding window is processed asynchronously to determine whether it is composed of low-entropy strings, and at the same time, another working thread is controlled to continue to block the subsequent data, which can effectively reduce the additional overhead caused by detecting low-entropy strings. When it is confirmed that the data is a low-entropy string, it is merged with the result of the working thread; otherwise, the maximum value in the subsequent fixed window is updated by the working thread.
[0020] After receiving the result of the data block optimization processing, further perform duplicate data detection processing on it. In this step, duplicate data is detected according to the metadata of the data block, the data block fingerprint, the data block boundary, and the data block length. The metadata includes the number of times the data block is repeated and the window size during data block division. The data block fingerprint represents the unique block identifier of the physical block. The data block boundary value represents the byte value of the data block point. The data block length represents the length of the current data block. When matching with the metadata, data block fingerprint, data block boundary, and data block length of the data stream in the subsequent same data file, if the match is successful, it is determined that the subsequent data block of this data block is a duplicate data block; otherwise, it is determined that there is no duplication.
[0021] After marking the data blocks determined to be duplicates, send a warning notice of this result to the user and wait for the user to check and process.
[0022] S104. Conduct access display and storage control management on data information; Exemplarily, in the embodiments of the present invention, when the user views the warning notice, the content to be displayed can be viewed through conditional filtering. The displayed content includes the reason for the warning notice, the corresponding warning data, and the warning time. The user can manage the data blocks determined to be duplicates, and can perform storage or filtering processing on the data blocks determined to be duplicates through the options after viewing the warning notice.
[0023] Embodiment 2: Embodiment 2 of the present invention provides an optimization processing system based on big data. Figure 2Schematic diagram of the module composition of an optimization processing system based on big data provided in the second embodiment of the present invention, as Figure 2 shown, the system includes: A data management module for controlling and managing cloud platform data; An analysis and processing module for performing storage optimization analysis and control of cloud platform data; A query and storage module for viewing, displaying, and storing and controlling data information.
[0024] In some embodiments of the present invention, the data management module includes: A registration and access module for managing user registration and access to the cloud platform; A data compression module for controlling and managing data compression; A backup management module for analyzing and controlling cloud platform data backup.
[0025] In some embodiments of the present invention, the analysis and processing module includes: A data chunking module for optimizing data chunking processing; An analysis and determination module for analyzing and determining duplicate data; An early warning notification module for managing early warning notifications to users.
[0026] In some embodiments of the present invention, the query and storage module includes: A query and display module for querying and displaying data information; A storage management module for storing and managing data information.
[0027] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variation thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0028] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements on some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An optimization processing method based on big data, characterized in that: include: Control and manage cloud platform registration and data access; Control the compression and backup of cloud platform data; Perform storage optimization, analysis and control of cloud platform data; Review, display, store, control and manage data information.
2. The optimization processing method based on big data according to claim 1, characterized in that: The control and management of cloud platform registration and data access includes: After registering in the big data system for the cloud platform, the user fills in the identity information, passes the background review and logs in to the system to carry out service control management.
3. The optimization processing method based on big data according to claim 1, characterized in that: The control process of compressing and backing up the cloud platform data includes: By setting the time window threshold and adopting multi-threaded processing, data compression control management is performed; Use data copy file threads to copy and back up multiple files simultaneously, and make each thread correspond to a data file. Use a data storage copy to copy the corresponding data file, and another thread writes it to the target file.
4. The optimization processing method based on big data according to claim 1, characterized in that: The storage optimization analysis and control of cloud platform data includes: After completing the setting of the fixed window and sliding window thresholds, perform data block optimization management on the cloud platform data files according to different low entropy strings; After receiving the data block optimization processing result, further perform duplicate data detection processing on it; After marking the data blocks determined to be duplicate, an early warning notification of the result is sent to the user, waiting for the user to review and process.
5. The optimization processing method based on big data according to claim 1, characterized in that: The data information review, display, storage control and management includes: When users view warning notifications, they can filter by conditions to view the displayed content, including the reason for the warning notification, the corresponding warning data and the warning time; The user can manage the data blocks that are determined to be duplicates, and can store or filter out the data blocks that are determined to be duplicates by checking the options after the warning notification.
6. An optimization processing system based on big data, characterized in that: include: Data management module, used to control and manage cloud platform data; Analysis and processing module, used for storage optimization, analysis and control of cloud platform data; The query storage module is used to query, display, store, control and manage data information.
7. The optimization processing system based on big data according to claim 6, characterized in that: The data management module comprises: The registration access module is used to manage the user's registration access to the cloud platform; Data compression module, used for control and management of data compression; The backup management module is used to analyze and control the cloud platform data backup.
8. The optimization processing system based on big data according to claim 6, characterized in that: The analysis and processing module comprises: Data block module, used to optimize data block processing; An analysis and determination module, used for analyzing and determining duplicate data; The early warning notification module is used to manage early warning notifications for users.
9. The optimization processing system based on big data according to claim 6, characterized in that: The query storage module includes: Query and display module, used to query and display data information; The storage management module is used to store and manage data information.
Citation Information
Patent Citations
Data backup method based on cloud computing
CN104536849A
Content-based partitioning method and system applied to duplicated data deletion and medium
CN114625316A
Internet of Things data optimization processing method and system
CN118733338A
High efficiency data storage system through data redundancy elimination based on parallel processing compression
KR102175094B1
Processing System of Data De-Duplication
US20120150824A1