Batch data storage method, system and device and medium
Through intelligent resource management and optimization of write performance methods, combined with asynchronous processing and data preprocessing technology, the performance bottleneck problem in large-scale data writing scenarios is solved, and efficient, stable and secure data storage processing is achieved.
Patent Information
- Application Number
- CN202510227151.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-20
AI Technical Summary
Existing data storage technology faces performance bottlenecks in large-scale data writing scenarios, resulting in extended system response time and even lag and crashes, affecting the real-time business processing capabilities and system stability.
Through intelligent resource management, optimized write performance, strict data verification, instant feedback mechanism and adaptive model training, asynchronous processing is used to write to the database in batches, and data integrity verification, compression processing and encryption processing are performed before each batch is written.
It significantly improves system performance and stability, reduces the waiting time during data writing, improves data writing efficiency, and ensures data security and integrity.
Smart Images

Figure CN120179708A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data storage, and more specifically, relates to a method, system, device and medium for batch data storage. Background Art
[0002] In today's digital age, big data, cloud computing, and Internet of Things technologies are booming at an unprecedented speed. With the wide spread of various intelligent devices and the accelerating digital transformation of enterprises, the data scale shows an explosive growth trend. For example, in the field of Internet e-commerce, there are massive amounts of information such as daily transaction records and user browsing data; in industrial Internet of Things, the amount of device operation status data collected by sensors in real time is also extremely large.
[0003] Facing such a large-scale data, the current data storage technologies expose many significant pain points. When performing large-scale data writing operations, the database system will encounter severe challenges. The database needs to frequently process a large number of I / O operations, which makes disk I / O a performance bottleneck. At the same time, the CPU in the system needs to continuously schedule and process these data, and the memory also needs to frequently read and write data, resulting in hardware resources such as CPU and memory being in a tense state due to carrying a large amount of data processing tasks. For example, in some traditional relational databases, when facing tens of thousands of data writes per second, the system response time will be greatly extended, and even situations such as freezing and crashing may occur. This performance degradation not only affects the real-time processing ability of the business, such as an online trading system being unable to complete order storage in time, resulting in transaction delays; but also reduces the overall stability of the system and increases the risk of data loss.
[0004] Therefore, there is an urgent need for a new technical solution to solve the problem of how to effectively utilize system resources and optimize writing performance in the scenario of large batches of data, so as to improve the overall performance and stability of the system. Summary of the Invention
[0005] Aiming at the above problems, the purpose of the present invention is to provide a method, system, device and medium for batch data storage, which provides an efficient, stable and reliable solution for the storage and processing of large amounts of data through intelligent resource management, optimized writing performance, strict data verification, instant feedback mechanism and adaptive model training, and is of great significance for improving system performance and user experience.
[0006] To achieve the above object, the present invention is realized through the following technical solutions: In a first aspect, an embodiment of the present application provides a method for batch data storage, including: Regularly obtain system load data through a monitoring tool and save it as historical load data; Establishing a data table structure on the server side, the data table structure includes a data target table and a batch information table; Listen to the data write request from the client, parse the data structure and expected total write amount in the data write request, and add a new data in the batch information table according to the expected total write amount; Based on historical load data, use the preset prediction model to analyze historical data writing, determine the optimal batch processing quantity, and adjust the data writing strategy according to the current system load data; Start the data writing strategy and perform data integrity check, data compression and encryption before writing each batch; Use asynchronous processing to write data into the database in batches; Update the writing progress information to the batch information table in real time.
[0007] In an optional implementation, the periodically acquiring system load data through a monitoring tool and saving it as historical load data includes: Introduced the monitoring tool Monitorix to obtain CPU statistics, memory statistics, and disk I / O statistics as system load data and save and output them to file reports; The acquired system load data is updated every five seconds at a set frequency. Based on the sliding time window mechanism, the average utilization rate, peak utilization rate, and valley utilization rate information of system resources are obtained for the system load data within the current thirty seconds. The collected data is saved in a file as historical load data.
[0008] In an optional implementation, the data target table is used as a table for data insertion; the batch information table is used to record the batch processing data volume and provide real-time feedback to the user, and the batch information table records the total amount of submitted data and the amount of written data.
[0009] In an optional implementation, the method of analyzing the historical data writing situation based on the historical load data using a preset prediction model, determining the optimal batch processing quantity, and adjusting the data writing strategy according to the load data of the current system includes: Obtain historical load data, continuously write data by setting the batch size, and monitor changes in system resources; The prediction model is constructed by using the difference in data writing amount and system resource utilization between different batches; the prediction model adopts the support vector machine regression method in the support vector machine method; When writing batch data, the optimal batch processing quantity is given according to the prediction model based on the current system utilization rate; Obtain the current system load data in real time and adjust the batch processing quantity using the preset rule model.
[0010] In an optional implementation manner, the startup data writing policy performs data integrity verification, data compression processing, and encryption processing before each batch writing, including: Before each batch writing, perform field format check and uniqueness verification of the data; For fields with non-null settings, verify whether the data is null and check the field format of the data; For fields with uniqueness settings, perform uniqueness verification by querying the database; Identify special fields in the data and encrypt the special fields using the SM4 encryption algorithm before storing the relevant data; For data that fails the field format check and uniqueness verification, record it in an error data file and return it to the client for resubmission; For large text or binary data, perform compression processing through a compression algorithm.
[0011] In an optional implementation manner, the asynchronous processing method is adopted to write data into the database in batches, including: Use the Executors class to create an ExecutorService instance to manage the thread pool; Create an implementation of Callable<List <result>>The task class of the interface uses its call() method to execute SQL statements and return a result list, while throwing a checked exception for multi-threading; In the call() method of Callable, use the syntax of INSERT INTO... VALUES (), (),... to construct SQL statements, and insert data in batches according to the batch size determined by the data writing strategy; Obtain the load data of the current system in real time, and adjust the batch processing quantity using a preset rule model.
[0012] In an optional implementation manner, the construction process of the prediction model includes: S401: Collect historical load data, extract CPU usage rate, memory usage rate, disk usage rate, and thread status features, as well as the corresponding written data volume, and construct a feature matrix; S402: Obtain the size of the written data under different system resource metrics and different data volumes, and generate corresponding labels; S403: Use the normalization method in the SVM method to perform data standardization; S404: Select features related to the adjustment of the written data volume based on the feature matrix and the historical written data volume; S405: Construct a prediction model based on the support vector machine regression method, and use the selected features and the standardized data to train the prediction model; the prediction model uses a Gaussian kernel function, and introduces 5-fold cross-validation during training to adjust the model parameters to optimize the prediction results; S406: After training, use an independent test data set to evaluate the performance of the model, and save the trained model in a loadable format, PMML format, and integrate it into the Java program.
[0013] In a second aspect, the embodiments of the present application further provide a batch data storage system, including: A data monitoring module for regularly obtaining system load data through a monitoring tool and saving it as historical load data; A data table construction module for establishing a data table structure on the server side, and the data table structure includes a data target table and a batch information table; A request listening module for listening to data writing requests from the client, parsing the data structure and the expected total written amount in the data writing request, and adding a new piece of data to the batch information table according to the expected total written amount; A policy prediction module for analyzing the historical data writing situation based on the historical load data using a preset prediction model, determining the optimal batch processing quantity, and adjusting the data writing policy according to the load data of the current system; A data processing module, which is used to start a data writing strategy and perform data integrity verification, data compression processing, and encryption processing before each batch of data is written. A storage module, which is used to write data into a database in batches by using an asynchronous processing method. A progress information update module, which is used to update the writing progress information into the batch information table in real time.
[0014] Thirdly, an embodiment of the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the batch data storage method described in any one of the above are implemented.
[0015] Fourthly, an embodiment of the present application further provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the batch data storage method described in any one of the above are implemented.
[0016] It can be seen from the above technical solutions that the present invention has the following advantages: In the batch data storage method provided by the present application, this method is specifically optimized for large-scale and high-concurrency data writing scenarios, aiming to significantly improve system performance through a series of strategies, while strictly ensuring the accuracy and integrity of data. This method combines an intelligent batching strategy to dynamically divide data into multiple batches for storage. An asynchronous processing strategy decouples the data writing operation from the client's response request, enabling the data to be processed asynchronously in the background. A database access strategy reduces the number of accesses to the database and the access time by optimizing query statements and establishing indexes. An instant feedback strategy immediately feeds back the writing progress to the client, effectively solving the performance bottleneck problem in the data review process and significantly improving the client response speed, thus improving the user experience.
[0017] By real-time monitoring and analyzing the usage of system resources, such as CPU usage rate, memory occupancy, disk I / O, etc., the present application can intelligently predict and adjust the data writing rate, thereby avoiding system overload and ensuring the effective utilization of system resources.
[0018] The present application jointly constructs an efficient data processing chain through data reception, storage, verification, and instant feedback. The present application greatly speeds up the data storage speed, reduces the processing delay, and improves the user experience by using dynamic SQL generation, intelligent batch writing, asynchronous execution, and real-time monitoring strategies.
[0019] This application uses an asynchronous processing method to write data into the database in batches, reducing the waiting time during the data writing process and improving the data writing efficiency. At the same time, through data integrity verification, data compression processing, and encryption processing, the security and integrity of the data during the writing process are ensured.
[0020] By establishing data table structures on the server side in this application, including a data target table and a batch information table, it can provide real-time feedback on the data writing progress to the user, enabling the user to clearly understand the data writing situation and enhancing the user experience.
[0021] This application supports dynamically adjusting the data writing strategy through a preset prediction model, and can adapt to changes in different system loads and data writing requirements. At the same time, by introducing technical means such as compression algorithms and encryption algorithms, more flexibility and scalability are provided for the storage and processing of data.
[0022] This application performs integrity verification, compression processing, and encryption processing before data writing, effectively preventing data loss, damage, and leakage during transmission and storage, and improving data security. Brief Description of the Drawings
[0023] To more clearly illustrate the technical solutions of the present invention, the accompanying drawings required for description will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0024] Figure 1 It is a schematic flowchart of the batch data storage method provided by this application.
[0025] Figure 2 It is a schematic flowchart of the construction method of the prediction model provided by this application.
[0026] Figure 3 It is a schematic structural diagram of the batch data storage system provided by this application.
[0027] Figure 4 It is a schematic structural diagram of the electronic device provided by this application. Detailed Embodiments
[0028] In the following, the specific steps of the batch data storage method will be described in detail, and various embodiments of the present disclosure will be described more comprehensively. The present disclosure can have various embodiments, and adjustments and changes can be made therein. However, it should be understood that there is no intention to limit the various embodiments of the present disclosure to the specific embodiments disclosed herein, but the present disclosure should be understood to cover all adjustments, equivalents, and / or alternative solutions falling within the spirit and scope of the various embodiments of the present disclosure.
[0029] In the following, the term "comprising" or "may comprise" that can be used in various embodiments of the present disclosure indicates the presence of the disclosed functions, operations, or elements, and does not limit the addition of one or more functions, operations, or elements. Further, as used in various embodiments of the present disclosure, the terms "comprising", "having", and their cognates are only intended to indicate a specific feature, number, step, operation, element, component, or combination of the foregoing items, and should not be construed as precluding the existence or addition of one or more other features, numbers, steps, operations, elements, components, or combinations of the foregoing items.
[0030] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0031] Please refer to Figure 1 The following is a flowchart of a method for batch data storage in a specific embodiment. The method includes: S1: Regularly obtain system load data through a monitoring tool and save it as historical load data.
[0032] In the specific implementation, first, introduce the monitoring tool Monitorix to obtain CPU statistics, memory statistics, and disk I / O statistics as system load data and save and output them to a file report. Among them, the obtained system load data is updated every five seconds at a set frequency. Based on the sliding time window mechanism, obtain the average usage rate, peak usage rate, and valley usage rate information of system resources for the system load data within the current thirty seconds, and save the collected data to a file as historical load data.
[0033] For example, introduce the lightweight monitoring tool Monitorix, configure monitoring items including the usage of system resources such as CPU, memory, and disk I / O, obtain CPU statistics and I / O statistics and save and output them to a file report, set to update every five seconds at a set frequency, and based on the sliding time window mechanism, obtain the average usage rate, peak usage rate, and valley usage rate information of system resources for the system load data within the current thirty seconds, and save the collected data to a file.
[0034] S2: Establish a data table structure on the server side, and the data table structure includes a data target table and a batch information table.
[0035] Exemplarily, establish a data table structure on the server side, including a data target table and a batch information table. Among them, the data target table is the table for data insertion. The batch information table is used to record the real-time feedback of the batch processing data volume to the user, including the total submitted data volume and the written data volume.
[0036] In this step, when designing the data target table, consider the rationality of fields and the creation of indexes. For common query conditions, such as the name field, establish a B-tree index to accelerate queries. Before inserting a large amount of data, select a deferred index. When creating and rebuilding indexes, choose a low-traffic period (such as early morning) to reduce the impact on online services. At the same time, set indexes for common fields such as the name field, and avoid unnecessary and excessive indexes to improve data query efficiency.
[0037] S3: Listen for data write requests from the client, parse the data structure and the expected total write volume in the data write request, and add a new piece of data in the batch information table according to the expected total write volume.
[0038] S4: Analyze the historical data write situation using a preset prediction model based on historical load data, determine the optimal batch processing quantity, and adjust the data write strategy according to the current system's load data.
[0039] The purpose of this step is to analyze the historical data write situation based on historical data, intelligently predict the appropriate write data volume size according to the current system's usage rate, and thus dynamically adjust the data write strategy to ensure that it can adapt to the system resource status.
[0040] Exemplarily, first, obtain historical load data, continuously write data by setting the batch size, and at the same time monitor the changes in system resources.
[0041] Then, construct a prediction model using the difference in data write volume and the difference in system resource utilization rate between different batches; the prediction model uses the method of support vector regression in the support vector machine method. Among them, the input of the prediction model is the usage situation of the current system resources, the output is the expected data write speed, that is, the data write volume size that can be achieved for each batch of data, and the kernel function selects the radial basis kernel function. Use the cross-validation method to evaluate the model to ensure the generalization ability of the model on unknown data.
[0042] Specifically, the input indicators of the prediction model cover multiple aspects, including the overall CPU usage rate; memory usage, including total memory, used memory, free memory, and the usage of the swap partition; disk I / O activities, such as read and write speeds, I / O wait times; process / thread status, including the number of active processes, the number of threads, and the number of blocked processes.
[0043] Next, when writing batch data, based on the utilization rate of the current system, the optimal batch processing quantity is given according to the prediction model.
[0044] Finally, the load data of the current system is obtained in real time, and the batch processing quantity is adjusted using the preset prediction model. For example: when it is detected that the system resources are tense (such as indicators like CPU utilization rate exceeding 80% and memory being less than 10% of the total memory), the writing rate is reduced by decreasing the amount of data written; on the contrary, if the system resources are sufficient, the amount of data written in batches and the writing rate are increased, so as to perform intelligent batch processing.
[0045] It should be particularly noted that when using system resource data for optimizing batch data, the system resources are monitored simultaneously according to the data, and a hierarchical alarm strategy is used. When the CPU utilization rate and IO utilization rate are too high, an alarm is issued, and different levels of alarm information are sent according to the degree of tension of the system resources.
[0046] S5: Start the data writing strategy. Before writing each batch, perform data integrity verification, data compression processing, and encryption processing.
[0047] For example, before writing each batch, perform basic data integrity verification. For fields with non-null settings, verify whether the data is null, check the field format of the data. For fields with uniqueness settings, perform uniqueness verification by querying the database. To ensure data security, set encryption storage requirements for special fields, and use the SM4 encryption algorithm to encrypt the fields before storage, etc. For data that does not meet the requirements, record it in an error data file and return it to the client for resubmission. For large text or binary data, apply an appropriate compression algorithm to reduce the storage space requirements and improve the writing efficiency.
[0048] S6: Use an asynchronous processing method to write the data to the database in batches.
[0049] For example, use an asynchronous processing method, and use the ExecutorService and Callable interfaces to start multiple threads to write the data to the database in batches.
[0050] Specifically, first use the Executors class to create an ExecutorService instance to manage the thread pool, and create an implementation of Callable<List <result>>The task class of the interface uses its call() method to execute SQL statements and return a result list, while throwing a checked exception for multi-threading.
[0051] Then, in the call() method of Callable, use the syntax of INSERT INTO... VALUES (), (),... to construct the SQL statement, and insert data in batches according to the batch size determined by the data writing strategy.
[0052] At the same time, obtain the load data of the current system in real time, and use a preset prediction model to adjust the batch processing quantity. That is, monitor the real-time resource status of the system during the writing process, and adjust the writing strategy according to the real-time model prediction when necessary.
[0053] S7: Update the writing progress information to the batch information table in real time.
[0054] Exemplarily, update the writing progress information to the batch table in real time for easy monitoring and querying. When all data is successfully written according to the strategy, send a notification of writing completion to the client through the server-side interface to inform that the data writing has been completed.
[0055] In this embodiment, by adaptively adjusting the batch processing quantity through dynamic load monitoring and intelligent prediction models, combined with asynchronous multi-threaded writing, data preprocessing mechanisms, and dual-table collaborative management, while ensuring data security and consistency, using sliding window real-time resource analysis and rule-model dual-driven strategies to dynamically optimize the writing efficiency, finally achieving high-throughput and low-latency data batch storage, and relying on progress visualization feedback and abnormal data retry mechanisms to significantly improve system stability and resource utilization.
[0056] In an embodiment of the present invention, based on step S4, the following will give a possible embodiment to non-restrictively elaborate on its specific implementation.
[0057] In the present invention, since it is impossible to obtain a large amount of historical data and it belongs to small-sample prediction, random forest, gradient boosting regression tree, neural network, and appropriate kernel methods can be selected for prediction. Here, the support vector machine regression method in the support vector machine method is selected as the prediction model. Among them, the input of the prediction model is the usage of the current system resources, and the output is the expected data writing speed, that is, the data writing volume that can be achieved for each batch of data. The radial basis kernel function is selected as the kernel function. The cross-validation method is used to evaluate the model to ensure the generalization ability of the model on unknown data.
[0058] Reference Figure 2 As shown, this embodiment discloses a method for constructing a prediction model, which specifically includes the following steps: S401: Collect historical load data, extract CPU usage rate, memory usage rate, disk usage rate, and thread status features, as well as the corresponding written data volume, and construct a feature matrix.
[0059] Exemplarily, collect historical data on system resource usage and written data volume. Include system resource metrics such as CPU usage rate, memory occupancy, disk I / O, network bandwidth, etc., and the corresponding written data volume for constructing the feature matrix. Among them, X is the feature matrix, C represents the CPU usage rate feature, M represents the memory usage rate feature, D represents the disk usage rate feature, and R represents the thread status feature.
[0060] S402: Obtain the size of the written data under different system resource metrics and different data volumes, and generate corresponding labels.
[0061] Exemplarily, obtain the size of the written data under different system resource metrics and different data volumes as the label Y = S. Among them, Y represents the target variable, and S is the data volume size.
[0062] S403: Use the normalization method in the SVM method to standardize the data.
[0063] Exemplarily, before using the aforementioned data for model training, some preprocessing steps are required. Include: Use the normalization method in the SVM method to standardize the data. This can help eliminate noise and outliers in the data and improve the prediction performance of the model.
[0064] S404: Select features related to the adjustment of the written data volume based on the feature matrix and historical written data volume.
[0065] Exemplarily, select features related to the adjustment of the written data volume from the collected data. In this method, the selected features include various system resource metrics, timestamps, and statistical data of historical written data volume.
[0066] S405: Construct a prediction model based on the support vector machine regression method, and use the selected features and the standardized data to train the prediction model.
[0067] Among them, the prediction model uses the Gaussian kernel function, and 5-fold cross-validation is introduced during training to adjust the model parameters to optimize the prediction results.
[0068] Specifically, when performing cross-validation, select 5-fold cross-validation k = 5, and set the regularization parameter C = [0.1, 1, 10, 100]. Since the RBF kernel function is used, the value of gamma is also important. Here, gamma is selected from 0.01 to 0.1 with a step size of 0.01, and the optimal parameters for cross-validation are selected.
[0069] S406: After training is completed, use an independent test data set to evaluate the performance of the model, and save the trained model in a loadable format, PMML format, and integrate it into a Java program.
[0070] Exemplarily, after training is completed, use an independent test data set to evaluate the performance of the model. Evaluation metrics include prediction accuracy, error rate, recall rate, F1 score, etc. According to the evaluation results, make necessary adjustments and optimizations to the model to improve its prediction ability, and save the trained model in a loadable format, PMML format, and integrate it into a Java program.
[0071] As Figure 3 shown, the following is an embodiment of a batch data storage system provided by the present disclosure. This system and the batch data storage methods of the above embodiments belong to the same inventive concept. For details not described in detail in the embodiment of the batch data storage system, reference may be made to the embodiments of the above batch data storage methods.
[0072] A batch data storage system includes: a data monitoring module, a data table construction module, a request listening module, a policy prediction module, a data processing module, a storage module, and a progress information update module.
[0073] The data monitoring module is used to regularly obtain system load data through a monitoring tool and save it as historical load data.
[0074] The data table construction module is used to establish a data table structure on the server side. The data table structure includes a data target table and a batch information table.
[0075] The request listening module is used to listen for data writing requests from the client, parse the data structure and the expected total amount to be written in the data writing request, and add a new piece of data to the batch information table according to the expected total amount to be written.
[0076] The policy prediction module is used to analyze the historical data writing situation based on the historical load data using a preset prediction model, determine the optimal batch processing quantity, and adjust the data writing policy according to the current system load data.
[0077] The data processing module is used to start the data writing policy, and before each batch is written, perform data integrity verification, data compression processing, and encryption processing.
[0078] The storage module is used to asynchronously write data into the database in batches.
[0079] The progress information update module is used to update the writing progress information to the batch information table in real time.
[0080] The batch data storage system provided in this embodiment collects system load data in real time through a data monitoring module, and combines a policy prediction module to dynamically optimize the batch processing quantity by using historical load data and a machine learning model, so as to achieve intelligent adjustment of the writing policy; the data processing module performs data integrity verification, compression, and encryption before writing to ensure the security and reliability of the data; the storage module uses an asynchronous processing method to efficiently write data in batches, and at the same time, the progress information update module real-time feedbacks the writing status, forming a full-process closed-loop management from data reception, processing to storage, significantly improving data writing efficiency, system stability, and data security.
[0081] Figure 4 Schematic diagram of the hardware structure of an electronic device for implementing each embodiment of the present invention.
[0082] The batch data storage method provided in the embodiments of the present application can be applied to an electronic device. Those skilled in the art can understand that the structure of the electronic device involved in the embodiments of the present invention does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements. In the embodiments of the present invention, the electronic device includes, but is not limited to, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present application described herein and / or claimed.
[0083] The electronic device may include a processor, an external memory interface, an internal memory, a universal serial bus (USB) interface, a charging management module, a power management module, a battery, a wireless communication module, an audio module, a speaker, a microphone, a sensor module, a key, a camera, a display screen, and a SIM card interface, etc.
[0084] The processor may include one or more processing units. For example, the processor may include a central processing unit (CPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0085] Among them, the processor may be the nerve center and command center of the electronic device. The controller can generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.
[0086] A memory may also be provided in the processor for storing instructions and data. In some embodiments, the memory in the processor is a cache memory. This memory can store the instructions or data that the processor has just used or recycled. If the processor needs to use the instruction or data again, it can directly call it from this memory. This avoids repeated accesses, reduces the waiting time of the processor, and thus improves the system efficiency.
[0087] The external memory interface can be used to connect an external memory card, such as a MicroSD card, to expand the storage capacity of the electronic device. The external memory card communicates with the processor through the external memory interface to implement the data storage function. For example, files such as music and videos are saved in the external memory card.
[0088] The internal memory can be used to store computer-executable program code, and the computer-executable program code includes instructions. The processor executes various functional applications and data processing of the electronic device by running the instructions stored in the internal memory. The internal memory may include a program storage area and a data storage area. The internal memory may include a high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0089] The wireless communication function of the electronic device can be implemented through an antenna, a wireless communication module, a modem processor, a baseband processor, etc.
[0090] The wireless communication module can provide solutions for wireless communications applied to electronic devices, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite systems (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc.
[0091] The electronic device can implement audio functions through an audio module, speaker, receiver, microphone, headphone jack, and application processor, etc.
[0092] The electronic device can implement a shooting function through an ISP, camera, video codec, GPU, display screen, and application processor, etc.
[0093] The electronic device can implement a display function through a GPU, display screen, and application processor, etc.
[0094] The GPU is a microprocessor for image processing, connecting the display screen and the application processor. The GPU is used to execute mathematical and geometric calculations for graphics rendering. The processor may include one or more GPUs, which execute program instructions to generate or change display information.
[0095] The display screen is used to display images, videos, etc. The display screen includes a display panel.
[0096] The above-mentioned electronic device realizes the batch data storage method of the present application by monitoring system load data in real time and constructing a historical data set, dynamically analyzing the correlation between resource characteristics and writing performance using a prediction model, and intelligently determining the optimal batch processing quantity; performing data integrity verification, compression, and encryption processing before writing to ensure data security and reliability; adopting an asynchronous processing method to efficiently write data in batches, and combining with a batch information table to update the writing progress in real time, forming a full-process closed-loop management from data reception, processing to storage, achieving the beneficial effects of significantly improving data writing efficiency, system stability, and data security.
[0097] In the storage medium provided by the present application, there is a program product capable of implementing the batch data storage method.
[0098] The batch data storage method includes: Regularly obtain system load data through a monitoring tool and save it as historical load data; Establish a data table structure on the server side, where the data table structure includes a data target table and a batch information table; Listen for data writing requests from the client, parse the data structure and the expected total amount to be written in the data writing request, and add a new piece of data to the batch information table according to the expected total amount to be written; Analyze the historical data writing situation based on the historical load data using a preset prediction model, determine the optimal batch processing quantity, and adjust the data writing strategy according to the load data of the current system; Start the data writing strategy. Before each batch is written, perform data integrity verification, data compression processing, and encryption processing; Use an asynchronous processing method to write the data to the database in batches; Update the writing progress information to the batch information table in real time.
[0099] In some possible implementation manners, the batch data storage method of the present disclosure may be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Method" section above of this specification.
[0100] The storage medium of the present disclosure may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0101] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.< / result> < / result>
Claims
1. A batch data storage method, characterized in that: include: Obtain system load data regularly through monitoring tools and save it as historical load data; Establishing a data table structure on the server side, the data table structure includes a data target table and a batch information table; Listen to the data write request from the client, parse the data structure and expected total write amount in the data write request, and add a new data in the batch information table according to the expected total write amount; Based on historical load data, use the preset prediction model to analyze historical data writing, determine the optimal batch processing quantity, and adjust the data writing strategy according to the current system load data; Start the data writing strategy and perform data integrity check, data compression and encryption before writing each batch; Use asynchronous processing to write data into the database in batches; Update the writing progress information to the batch information table in real time.
2. The method for storing batch data according to claim 1, characterized in that: The system load data is obtained regularly through the monitoring tool and saved as historical load data, including: Introduced the monitoring tool Monitorix to obtain CPU statistics, memory statistics, and disk I / O statistics as system load data and save and output them to file reports; The acquired system load data is updated every five seconds at a set frequency. Based on the sliding time window mechanism, the average utilization rate, peak utilization rate, and valley utilization rate information of system resources are obtained for the system load data within the current thirty seconds. The collected data is saved in a file as historical load data.
3. The method for storing batch data according to claim 1, characterized in that: The data target table is used as a table for data insertion; the batch information table is used to record the batch processing data volume and provide real-time feedback to the user. The batch information table records the total amount of submitted data and the amount of written data.
4. The method for storing batch data according to claim 1, characterized in that: The method of analyzing the historical data writing situation based on the historical load data using a preset prediction model, determining the optimal batch processing quantity, and adjusting the data writing strategy according to the load data of the current system includes: Obtain historical load data, continuously write data by setting the batch size, and monitor changes in system resources; The prediction model is constructed by using the difference in data writing amount and system resource utilization between different batches; the prediction model adopts the support vector machine regression method in the support vector machine method; When writing batch data, the optimal batch processing quantity is given according to the prediction model based on the current system utilization rate; Obtain the current system load data in real time and use the preset prediction model to adjust the batch processing quantity.
5. The method for storing batch data according to claim 1, characterized in that: The startup data writing strategy performs data integrity verification, data compression processing and encryption processing before each batch is written, including: Before writing each batch, perform field format check and uniqueness verification on the data; For fields that are set to be non-empty, verify whether the data is empty and check the field format of the data; For fields that have been set to be unique, uniqueness verification is performed by querying the database; Identify special fields in the data and encrypt them using the SM4 encryption algorithm before storing the relevant data; For data that fails to pass the field format check and uniqueness verification, it is recorded in the error data file and returned to the client for resubmission; For large text or binary data, compression is performed using a compression algorithm.
6. The method for storing batch data according to claim 1, characterized in that: The method of writing data into the database in batches in an asynchronous processing manner includes: Use the Executors class to create an ExecutorService instance to manage the thread pool; Create a class that implements Callable <List <result> >The task class of the interface uses its call() method to execute the SQL statement and return the result list, while throwing a checked exception multithreading;< / result> In the call() method of Callable, use the INSERT INTO ... VALUES (), (), ... syntax to construct SQL statements and insert data in batches according to the batch size determined by the data writing strategy; Obtain the current system load data in real time and use the preset prediction model to adjust the batch processing quantity.
7. The batch data storage method according to claim 4, characterized in that: The construction process of the prediction model includes: S401: Collect historical load data, extract CPU usage, memory usage, disk usage and thread state features, as well as the corresponding written data volume, and construct a feature matrix; S402: Obtain the size of written data under different system resource indicators and different data volumes, and generate corresponding tags; S403: using the normalization method in the SVM method to perform data standardization; S404: Selecting features related to adjustment of the amount of written data based on the feature matrix and the amount of historical written data; S405: constructing a prediction model based on a support vector machine regression method, and using the selected features and the standardized data to train the prediction model; the prediction model adopts a Gaussian kernel function, and introduces 5-fold cross validation in the training to adjust the model parameters to optimize the prediction results; S406: After the training is completed, an independent test data set is used to evaluate the performance of the model, and the trained model is saved in a loadable PMML format and integrated into the Java program.
8. A batch data storage system, characterized in that: The system adopts the batch data storage method as claimed in any one of claims 1 to 7; The system comprises: The data monitoring module is used to obtain system load data regularly through the monitoring tool and save it as historical load data; A data table construction module is used to establish a data table structure on the server side, wherein the data table structure includes a data target table and a batch information table; The request monitoring module is used to monitor the data writing request from the client, parse the data structure and expected total writing amount in the data writing request, and add a new data in the batch information table according to the expected total writing amount; The strategy prediction module is used to analyze the historical data writing situation based on the historical load data using the preset prediction model, determine the optimal batch processing quantity, and adjust the data writing strategy according to the current system load data; The data processing module is used to start the data writing strategy and perform data integrity verification, data compression processing and encryption processing before each batch is written; A storage module is used to write data into the database in batches using an asynchronous processing method; The progress information update module is used to update the writing progress information to the batch information table in real time.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the bulk data storage method according to any one of claims 1 to 7 are implemented.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the bulk data storage method according to any one of claims 1 to 7 are implemented.