Data management method, system and equipment under high-concurrency write operation and medium
By separating data storage and identification into two files, and combining regular scan and merge strategies, the performance bottlenecks and resource waste of traditional databases in high concurrent write operations are solved, and efficient data management and optimization are achieved.
Patent Information
- Application Number
- CN202510494045.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-19
- Publication Date
- 2025-07-29
AI Technical Summary
Traditional relational databases perform poorly in high concurrent write operations, locking mechanisms lead to performance bottlenecks and waste of computing resources, and existing distributed databases have limitations in handling high concurrent write operations.
By separating the data storage and data identification into two files, namely the first data file and the second data file, data or data identification is written to the respective files according to the type of the write operation instructions, and the identification list in the second data file is periodically scanned to delete the corresponding data in the first data file, combining merging and optimizing the storage strategy.
It improves data writing efficiency, reduces concurrent access conflicts, enhances the scalability and user experience of the system, and optimizes storage space utilization and data consistency.
Smart Images

Figure CN120386490A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data management, and particularly to a data management method, system, device, and medium under high-concurrency write operations. Background Art
[0002] With the development of Internet technology, more and more online services need to process a large number of concurrent access requests. Especially in terms of data writing, traditional relational databases perform poorly when dealing with high-concurrency write operations. The main reason lies in the performance bottleneck caused by the lock mechanism.
[0003] To break through this limitation, various new types of distributed databases have been developed. They have improved the writing performance through different mechanisms, such as memory-based operations, multi-version concurrency control, write-ahead logging, etc. These technologies have alleviated the problem of high-concurrency write operations to a certain extent. Although the above solutions can effectively meet the requirements of high-concurrency write operations under specific conditions, there are still some obvious limitations. Most technologies rely on complex lock mechanisms, which not only consume a large amount of computing resources but also may cause serious blocking phenomena.
[0004] Therefore, how to efficiently manage and optimize data write operations has become a key problem to be solved urgently. Summary of the Invention
[0005] This application provides a data management method, system, device, and medium under high-concurrency write operations. By optimizing data storage and access strategies, it improves data writing efficiency, resource utilization rate, and data consistency, while enhancing the scalability and user experience of the system.
[0006] In the first aspect of this application, a data management method under high-concurrency write operations is provided, which is applied to a data management platform. The method includes: Create a first data file and a second data file. The first data file is used to store data, and the second data file is used to store data identifiers; In response to receiving a write operation instruction sent by a client, analyze the write operation instruction to determine the type of the write operation instruction; According to the type of the write operation instruction, write target data to the first data file and / or write target data identifiers to the second data file, and return the change status information of the first data file and / or the second data file to the client; Scan the identifier list in the second data file every preset time, and delete the data at the target position in the first data file. The target position is the position of the data in the first data file corresponding to the identifier in the identifier list.
[0007] Optionally, writing target data to the first data file and / or writing a target data identifier to the second data file according to the type of the write operation instruction includes: When the type of the write operation instruction is an insert operation, extracting first target data to be inserted from the write operation instruction, and writing the first target data to the end of the first data file; When the type of the write operation instruction is a modification operation, extracting a first target data identifier of data to be modified and second target data after modification from the write operation instruction, writing the second target data to the end of the first data file, and writing the first target data identifier to the end of the second data file; When the type of the write operation instruction is a deletion operation, extracting a second target data identifier of data to be deleted from the write operation instruction, and writing the second target data identifier to the end of the second data file.
[0008] Optionally, scanning the identifier list in the second data file every preset time and deleting data at a target position in the first data file includes: Scanning the identifier list in the second data file to identify all third target data identifiers marked for deletion; Locating a corresponding target position in the first data file according to the third target data identifier; Removing the third target data corresponding to the target position from the first data file, and determining the size of the first data file after removing the third target data; Merging multiple first data files according to the size.
[0009] Optionally, merging multiple first data files according to the size includes: Dividing multiple first data files into a small file list, a medium file list, and a large file list according to a preset first threshold and a preset second threshold. The small file list contains first data files smaller than the preset first threshold, the large file list contains first data files greater than or equal to the preset second threshold, and the medium file list contains first data files greater than or equal to the preset first threshold and smaller than the preset second threshold; Selecting a first file from the large file list, and matching a second file from the small file list such that the difference between the size of a third file obtained by merging the first file and the second file and a preset size is less than an error threshold until the large file list is empty or the small file list is empty; When the large file list is empty, merge the fourth file in the medium file list with the remaining files in the small file list; When the small file list is empty, merge the files in the medium file list.
[0010] Optionally, the merging of the multiple first data files according to the size includes: Monitor the access records of the data, and divide the multiple first data files into high - heat - degree data files and low - heat - degree data files according to a preset heat - degree threshold. The high - heat - degree data files are data files whose access frequency is greater than the preset heat - degree threshold, and the low - heat - degree data files are data files whose access frequency is less than or equal to the preset heat - degree threshold; Merge the low - heat - degree data files in chronological order from earliest to latest to form multiple third data files, and merge the high - heat - degree data files in chronological order from earliest to latest to form multiple fourth data files. Both the third data files and the fourth data files are smaller than the preset size.
[0011] Optionally, the method further includes: When a query request is received, retrieve a candidate data set from the first data file according to the query request; Use an index to find the first candidate data marked as deleted from the second data file, and delete the first candidate data from the candidate data set to obtain a second candidate data set. The index is constructed based on the data identifiers in the second data file; Perform sorting and / or aggregation processing on the second candidate data set to meet the requirements of the query request, and encapsulate and return the processed result to the client.
[0012] Optionally, the method further includes: After writing the target data at the end of the first data file, generate and store a first timestamp. After writing the target data identifier at the end of the second data file, generate and store a second timestamp; When scanning the identifier list in the second data file, determine the processing priority of the write operation instruction according to the first timestamp and the second timestamp, and process the fourth target data identifier and the fourth target data corresponding to the write operation instruction in the order of the processing priority.
[0013] In the second aspect of the present application, a data management system under high - concurrency write operations is provided, including a creation module, an analysis module, an execution module, and a deletion module, where: A creation module, configured to create a first data file and a second data file, where the first data file is used to store data and the second data file is used to store data identifiers; An analysis module, configured to analyze the write operation instruction received from the client to determine the type of the write operation instruction in response to receiving the write operation instruction sent by the client; An execution module, configured to write target data to the first data file and / or write target data identifiers to the second data file according to the type of the write operation instruction, and return the change status information of the first data file and / or the second data file to the client; A deletion module, configured to scan the identifier list in the second data file at preset intervals and delete the data at the target position in the first data file, where the target position is the position of the data in the first data file corresponding to the identifier in the identifier list.
[0014] In a third aspect of the present application, an electronic device is provided, including a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, both the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the method described in any one of the above.
[0015] In a fourth aspect of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores instructions, and when the instructions are executed, the method described in any one of the above is executed.
[0016] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. By separating data storage and data identifier storage into two different files (the first data file and the second data file), the processing logic of each write operation can be simplified. Especially when write operations are frequent, this separation can reduce concurrent access conflicts to a single file, thereby improving the write efficiency; 2. By clearly distinguishing the types of write operation instructions (such as writing new data, updating data, or deleting data identifiers) and updating the data file and the identifier file accordingly, it is easier to maintain data consistency. Especially in a high-concurrency environment, this method helps to reduce the risk of data conflicts and inconsistencies; 3. Scanning the identifier list in the second data file at preset intervals and deleting the corresponding data in the first data file, this asynchronous cleaning mechanism can reduce interference with write operations. It allows the system to perform data cleaning during off-peak hours or when resources are relatively idle, thereby avoiding introducing additional latency during high concurrency; 4. Provides better scalability for the data management platform. As the amount of data increases, new needs can be accommodated by adding data files, optimizing storage strategies, or adjusting scanning frequencies without making major changes to the system architecture. 5. By promptly returning the status information of the first data file and / or the second data file to the client, the user's perception and control of the system status can be enhanced. This helps improve the user experience, especially in application scenarios that require real-time feedback. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 Schematic diagram of the process of the data management method under high concurrent write operation disclosed in the embodiment of the present application; Figure 2 This is a module diagram of a data management system under high concurrent write operations disclosed in an embodiment of the present application; Figure 3 This is a structural diagram of an electronic device disclosed in an embodiment of the present application.
[0018] Explanation of the accompanying drawings: 201, creation module; 202, analysis module; 203, execution module; 204, deletion module; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. DETAILED DESCRIPTION
[0019] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments.
[0020] In the description of the embodiments of this application, words such as "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "for example" or "for instance" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "for example" or "for instance" is intended to present the relevant concepts in a concrete manner.
[0021] In the description of the embodiments of the present application, the term "plurality" means two or more. For example, a plurality of systems means two or more systems, and a plurality of screen terminals means two or more screen terminals. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "comprise", "include", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0022] This embodiment discloses a data management method under high-concurrency write operations, which is applied to a data management platform. Figure 1 It is a schematic flowchart of the data management method under high-concurrency write operations disclosed in the embodiments of the present application, as Figure 1 shown. The method includes the following steps: S101. Create a first data file and a second data file. The first data file is used to store data, and the second data file is used to store data identifiers. S102. In response to receiving a write operation instruction sent by a client, analyze the write operation instruction to determine the type of the write operation instruction. S103. Write target data to the first data file and / or write target data identifiers to the second data file according to the type of the write operation instruction, and return the change status information of the first data file and / or the second data file to the client. S104. Scan the identifier list in the second data file at preset intervals, and delete the data at the target position in the first data file. The target position is the position of the data in the first data file corresponding to the identifier in the identifier list.
[0023] In a data management platform, two key files need to be created first: the first data file (named x for example) and the second data file (named y for example). These two files have different responsibilities: the first data file (x) is used to store actual data entries. These data entries can be any form of data submitted by users, such as text, pictures, videos, etc., depending on the design goals and application scenarios of the data management platform. In the embodiment of this application, the capacity of the first data file is set to 5GB, which means that all new data will be appended to the first data file before reaching this capacity limit. The second data file (y) is used as an auxiliary file to record the data IDs (identifications) that need to be deleted. These data identifications correspond to the data entries in the first data file, but it does not mean that the data is immediately physically deleted from the first data file. Instead, when a data entry needs to be deleted, the data identification of this data entry will be added to the second data file for subsequent logical deletion and physical cleaning. Similarly, the capacity of the second data file is also set to 5GB to accommodate enough deletion identifications. When the data management platform receives a write operation instruction sent by the client, it will first analyze these instructions to determine their types. Write operation instructions usually include three types: insert, modify, and delete. Write the target data to the first data file and / or write the target data identification to the second data file according to the type of the write operation instruction. After completing the write operation, the data management platform will return the change status information of the first data file and / or the second data file to the client. The change status information may include whether the write operation is successful, the location or index of the data, etc., so that the client can correctly process subsequent operations. To maintain data consistency and optimize storage space, the data management platform needs to regularly scan the identification list in the second data file and delete the corresponding data in the first data file according to these identifications. This process is usually called merging or cleaning. The data management platform will scan the identification list in the second data file every preset time (such as daily, weekly, or monthly). These identifications represent all the logically deleted data entries. According to the scanned identifications, the data management platform will locate the corresponding data positions in the first data file and perform physical deletion operations. This process may require traversing the entire first data file or using indexes to accelerate the search process. After deleting the data, the data management platform also needs to perform optimization processing on the first data file. Since the deletion operation may cause file fragmentation or generate too much free space, the storage layout can be optimized by merging small files or reallocating storage space. These optimization operations help improve subsequent read and write performance and storage utilization.
[0024] By directly appending new data to the end of the first data file, the waiting problem caused by locking is avoided, thus improving the concurrent writing ability. This lock-free writing method allows multiple clients to write data to the data file simultaneously without interfering with each other. Regularly scanning the identification list in the second data file and deleting the corresponding data in the first data file based on these identifications helps to free up storage space, reduce fragmentation, and improve storage utilization. By merging overly small files or reallocating storage space, the storage layout can be further optimized to improve read and write performance. Although an additional check of the second data file is required during query to exclude deleted data, this process can usually be accelerated through indexing, thus reducing the impact on query performance. Regular merging operations reduce the amount of invalid data, further improving query efficiency. By taking measures such as regularly backing up important data, reasonably planning file system parameters, and monitoring system resource usage, the stability and reliability of the system can be ensured, and problems can be discovered and solved in a timely manner.
[0025] Optionally, the writing of the target data to the first data file and / or the writing of the target data identifier to the second data file according to the type of the write operation instruction includes: When the type of the write operation instruction is an insert operation, extract the first target data to be inserted from the write operation instruction and write the first target data to the end of the first data file; When the type of the write operation instruction is a modification operation, extract the first target data identifier of the data to be modified and the modified second target data from the write operation instruction, write the second target data to the end of the first data file, and write the first target data identifier to the end of the second data file; When the type of the write operation instruction is a deletion operation, extract the second target data identifier of the data to be deleted from the write operation instruction and write the second target data identifier to the end of the second data file.
[0026] When the type of the write operation instruction is an insert operation, the first target data to be inserted is parsed from the write operation instruction. The first target data can be any form of structured or unstructured data, depending on the application scenario of the system. The first target data is appended to the end of the first data file. Since the append method is used, the system does not need to traverse the entire file to find the insertion position, thus avoiding the waiting problem caused by locking and improving the writing efficiency. After the data writing is completed, the system returns the status information of the write operation to the client, including whether the writing is successful, the data position where the data is written (if necessary), etc. When the type of the write operation instruction is a modification operation, the first target data identifier of the data to be modified (i.e., the ID or index of the data to be modified) and the second target data after modification are parsed from the write operation instruction. The second target data (i.e., the modified data) is appended to the end of the first data file. Note that here the original data is not directly overwritten, but a new data record is created. The first target data identifier (i.e., the original ID or index of the data to be modified) is appended to the end of the second data file. This is done to mark that the old data has been modified or discarded, so that it can be correctly identified and processed in subsequent query or cleanup operations. After the data writing and marking are completed, the system returns the status information of the modification operation to the client. When the type of the write operation instruction is a deletion operation, the second target data identifier of the data to be deleted (i.e., the ID or index of the data to be deleted) is parsed from the write operation instruction. The second target data identifier is appended to the end of the second data file. This is done to mark that the data has been logically deleted, so that it can be excluded or deleted in subsequent query or cleanup operations. After the writing of the data identifier is completed, the system returns the status information of the deletion operation to the client.
[0027] When the write operation instruction is of the insert type, the system directly extracts the first target data to be inserted from the instruction and appends it to the end of the first data file. This append-write method avoids data overwrite and locking issues, improving write efficiency and concurrent performance. For modification operations, the system not only extracts the modified second target data but also the first target data identifier of the data to be modified. Then, the new data is appended to the end of the first data file, and the identifier of the old data is appended to the end of the second data file. This approach preserves the historical versions of the data, allowing for rollback or recovery when necessary, while avoiding potential concurrency issues associated with directly overwriting old data. When the write operation instruction is of the delete type, the system only extracts the second target data identifier of the data to be deleted and appends it to the end of the second data file. This logical deletion method avoids the risk of data loss that may result from immediate physical deletion of data, while retaining a record of the deletion operation for subsequent data cleanup and recovery operations. By writing the data identifiers of modification and deletion operations to the second data file, the system can maintain data consistency and integrity. When querying or processing data, the system can exclude deleted or modified data based on the identifier list in the second data file to ensure data accuracy and reliability. Although modification and deletion operations cause the first data file to grow continuously, the second data file records the identifiers of the old data, enabling the system to effectively release storage space during subsequent merge and cleanup operations. This design avoids fragmentation problems caused by frequent physical deletion operations and improves storage space utilization. The design of this process enables the system to easily handle data volume growth. By increasing the size or number of the first and second data files, the system can expand its capacity to adapt to changes in business requirements. At the same time, this process also supports multiple types of write operation instructions, offering strong flexibility and scalability.
[0028] Optionally, the step of scanning the identifier list in the second data file at every preset time and deleting the data at the target position in the first data file includes: Scanning the identifier list in the second data file to identify all third target data identifiers marked as deleted; Locating the corresponding target position in the first data file based on the third target data identifier; Removing the third target data corresponding to the target position from the first data file and determining the size of the first data file after removing the third target data; Merging multiple first data files based on the size.
[0029] The system scans the identification list in the second data file regularly (every preset time). The identification list contains the identifications of all the data marked for deletion, i.e., the third target data identifications. These identifications were written during the previous deletion operation to record which data has been logically deleted. During the scanning process, the system identifies all the third target data identifications marked for deletion. These identifications are the targets for subsequent data cleaning operations. Based on the identified third target data identifications, the system locates the corresponding target positions in the first data file. These positions store the data marked for deletion. When the target positions are located, the system removes this data from the first data file. This process is a physical deletion, i.e., truly erasing this data from the storage medium. After removing the data, the system recalculates the size of the first data file. This is because after data deletion, the file size usually decreases, and the system needs to know the new size for subsequent file management operations. If there are multiple first data files in the system, the system decides whether to merge them based on the sizes of these first data files and the current storage requirements. The merge operation can optimize the utilization of storage space, reduce file fragmentation, and improve the efficiency of data access.
[0030] By scanning the identification list in the second data file, the system can identify all the third target data identifications marked for deletion. The data corresponding to these identifications is regarded as invalid or outdated data in the first data file. Removing this data can free up storage space for subsequent data writing, thus improving the utilization rate of storage space. When deleting the third target data at the target positions in the first data file, the system ensures that only the data clearly marked for deletion in the second data file is deleted. This mechanism avoids the risk of accidentally deleting valid data, thus maintaining the consistency and integrity of the data. After the deletion operation, the first data file may become fragmented, i.e., there are multiple discontinuous data blocks in the file. By merging multiple first data files, the system can reorganize these data blocks into a continuous data stream, thereby reducing the number of disk I / O operations and improving the speed of data reading and writing. This process is designed to be automatically executed every preset time without manual intervention. This automated maintenance mechanism reduces the burden on system administrators and improves the stability and reliability of the system. By regularly cleaning invalid data and merging files, the system can utilize disk resources more effectively. This helps to reduce disk fragmentation, improve the read and write performance of the disk, and thus enhance the overall operating efficiency of the system.
[0031] Optionally, the merging of the multiple first data files according to the size includes: Divide the multiple first data files into a small file list, a medium file list, and a large file list according to a preset first threshold and a preset second threshold. The small file list contains first data files smaller than the preset first threshold, the large file list contains first data files greater than or equal to the preset second threshold, and the medium file list contains first data files greater than or equal to the preset first threshold and smaller than the preset second threshold; Select a first file from the large file list and match a second file from the small file list so that the difference between the size of the third file obtained by merging the first file and the second file and the preset size is less than the error threshold until the large file list is empty or the small file list is empty; When the large file list is empty, merge the fourth file in the medium file list with the remaining files in the small file list; When the small file list is empty, merge the files in the medium file list.
[0032] The system first sets a preset first threshold and a preset second threshold. The preset first threshold (e.g., 1G) is less than the preset second threshold (e.g., 4G). These two thresholds are used to divide multiple first data files into three different lists: a small file list, a medium file list, and a large file list. The small file list contains all first data files smaller than the preset first threshold. These files are usually small and may be generated due to frequent small-scale data writing operations. The large file list contains all first data files greater than or equal to the preset second threshold. These files are usually large and may contain a large amount of data. The medium file list contains all first data files with sizes between the preset first threshold and the preset second threshold. The system selects a file (referred to as the first file) from the large file list as the starting point for merging. Then, the system searches for one or more files (referred to as the second files) in the small file list. The sum of the sizes of these files and the size of the first file has a difference less than the preset error threshold (e.g., 200M) from the preset size (e.g., 5G). When suitable second files are found, the system merges the first file and these second files into a new file (referred to as the third file). This process will repeat until the large file list is empty (i.e., all large files have been merged) or the small file list is empty (i.e., there are no more small files available for merging with large files). If the large file list is empty but there are still files remaining in the small file list, the system will merge the files (referred to as the fourth files) in the medium file list with the remaining files in the small file list. This may be to utilize the remaining small file space or to adjust the size of the medium files to be closer to the preset size. If the small file list is empty but there are still files remaining in the large file list, the system will merge the files in the medium file list. This may be to reduce the number of files and improve the efficiency of storage management.
[0033] By dividing the files into three categories: small, medium, and large, the system can manage storage resources more flexibly. For large files, the system attempts to optimize the size of the merged files by matching small files to make it close to the preset size, thereby reducing the waste of storage space. The size of the merged files is more balanced, which helps to reduce the fragmentation of the file system and improve the read and write performance of the disk. At the same time, the smaller number of files also reduces the metadata overhead of the file system, further improving the file access efficiency. The system preferentially matches files from the large file list and the small file list for merging, and this strategy can reduce the number and complexity of the merge operations. When the large file list or the small file list is empty, the medium file list is processed, ensuring the efficiency of the merge process. By reasonably allocating the file merge tasks, the system can avoid generating excessive load on a single disk or storage node. This helps to maintain the overall performance and stability of the system and prevent performance bottlenecks or failures caused by uneven load.
[0034] Optionally, the merging of the multiple first data files according to the size includes: Monitoring the access records of the data, and dividing the multiple first data files into high-heat data files and low-heat data files according to a preset heat threshold. The high-heat data files are data files with an access frequency greater than the preset heat threshold, and the low-heat data files are data files with an access frequency less than or equal to the preset heat threshold; Merging the low-heat data files in chronological order from early to late to form multiple third data files, and merging the high-heat data files in chronological order from early to late to form multiple fourth data files. Both the third data files and the fourth data files are smaller than the preset size.
[0035] The system needs to have a mechanism to record the access frequency of each data file. This can be achieved by triggering a recording event during data access or by using a dedicated logging system to capture access information. Generate a database or log containing each data file and its access frequency. The preset popularity threshold is a value set according to system requirements and is used to distinguish between high-popularity and low-popularity data files. Data files with an access frequency greater than the preset popularity threshold are classified as high-popularity data files. High-popularity data files are usually those that are frequently accessed by users or are critical business data. Data files with an access frequency less than or equal to the preset popularity threshold are classified as low-popularity data files. Low-popularity data files may contain historical data, infrequently used reports, or backup files. Obtain two separate file lists for high-popularity and low-popularity data files. Low-popularity data file merging: Starting from the earliest file, merge them one by one in chronological order until the size of the merged file approaches or reaches the preset size limit. Then, create a new third data file and continue to merge the remaining low-popularity data files. High-popularity data file merging: Similar to low-popularity data files, merge high-popularity data files in chronological order starting from the earliest file to form multiple fourth data files.
[0036] By monitoring the access records of data and classifying data files into high - heat and low - heat based on a preset heat threshold, the system can identify the data frequently accessed by users. High - heat data files are processed and merged separately, which helps reduce the disk I / O operations required when accessing these popular data, thus improving the data access speed. Merging low - heat and high - heat data files in chronological order helps reduce storage fragmentation. As data is continuously written and deleted, a large amount of fragmented space may be generated on the storage device, which will affect storage efficiency and data access speed. By merging files, the data can be reorganized, the fragmented space can be released, and the storage utilization rate can be improved. Classifying and merging data files according to heat helps simplify storage management. The system can manage data of different heats more effectively. For example, for high - heat data, it can be stored on a storage device with higher performance to improve access speed; for low - heat data, it can be stored on a storage device with lower cost to reduce costs. By merging files, the amount of data during data backup and recovery can be reduced. When data needs to be backed up or recovered, the system can process the merged files more efficiently instead of dealing with a large number of scattered files. This helps shorten the backup and recovery time and improve the availability and reliability of the system. The merged data files are more ordered and structured, which helps support data analysis and mining work. The system can process and analyze the merged data more efficiently, extract valuable information, and provide support for business decisions. By merging high - heat data files and low - heat data files separately, the system can distribute the I / O load more evenly. This helps prevent some storage devices or processors from being overloaded and improves the overall performance and stability of the system.
[0037] Optionally, the method further includes: When a query request is received, retrieving a candidate data set from the first data file according to the query request; Using an index to find the first candidate data marked as deleted from the second data file, and deleting the first candidate data from the candidate data set to obtain a second candidate data set, where the index is constructed based on the data identifiers in the second data file; Performing sorting and / or aggregation processing on the second candidate data set to meet the requirements of the query request, and encapsulating and returning the processed result to the client.
[0038] The system receives a query request from a client. This request may contain specific query conditions, sorting requirements, aggregation needs, etc. After receiving the query request, the system first retrieves a candidate data set that matches the query conditions from the first data file. The first data file may contain a large amount of raw data, which is either unprocessed or only preliminarily processed. After obtaining the candidate data set, the system uses an index built based on the data identifiers in the second data file to find the first candidate data marked as deleted. This data may have been marked as deleted at some previous point in time but remains in the data file for various reasons (such as delayed deletion, transaction processing, etc.). The system deletes this data marked as deleted from the candidate data set to obtain the second candidate data. An index is a data structure that allows the system to quickly locate specific records in the data file. In this scenario, the index is built based on data identifiers, which can accelerate the process of finding data marked as deleted. After obtaining the second candidate data, the system processes this data according to the sorting requirements and / or aggregation needs in the query request. The sorting operation may involve arranging the data in ascending or descending order according to a specific field; the aggregation operation may involve calculating statistical data such as sum, average, maximum, minimum, etc. The system encapsulates the processed result into an appropriate format (such as JSON, XML, etc.) and returns it to the client via the network.
[0039] By first retrieving a set of candidate data from the first data file, the system can quickly locate the set of files that may contain the required data. This step helps reduce the amount of data to be processed subsequently, thereby improving the overall query efficiency. Using the index constructed based on the data identifiers in the second data file, the system can accurately find the first candidate data marked as deleted. This step ensures that even if the data is deleted, the system can promptly exclude these invalid data from the query results, thus maintaining data consistency. After obtaining the second candidate data, the system sorts and / or aggregates it to meet the specific requirements of the query request. This step not only helps improve the accuracy and relevance of the query results but also enables flexible processing and analysis of the data according to business needs. By encapsulating and returning the processed results to the client, the system can provide more intuitive and understandable query results. This step helps enhance the user experience, enabling users to more conveniently obtain the information they need. This method supports dynamic updates and management of the first data file and the second data file. As the data volume grows and business requirements change, the system can optimize the query process by adding new data files, adjusting the index construction strategy, etc., to ensure the scalability and flexibility of the system. By quickly locating the set of candidate data in the first stage and using the index to exclude invalid data in the second stage, the system can reduce unnecessary disk I / O operations and data processing tasks. This step helps reduce the system load and improve the system's response speed and stability.
[0040] Optionally, the method further includes: After writing the target data at the end of the first data file, generating and storing a first timestamp, and after writing the target data identifier at the end of the second data file, generating and storing a second timestamp; When scanning the list of identifiers in the second data file, determining the processing priority of the write operation instruction according to the first timestamp and the second timestamp, and processing the fourth target data identifier and the fourth target data corresponding to the write operation instruction in the order of the processing priority.
[0041] When the system receives a write operation instruction, it first writes the target data at the end of the first data file. This typically involves writing the data to a certain location on the disk and updating the file system metadata to reflect this change. After successfully writing the target data, the system generates a first timestamp that records the exact time when the data was written. The first timestamp is crucial for subsequent data consistency checks and recovery operations. The system stores this first timestamp in a secure location, usually in the metadata file associated with the first data file or in a dedicated log file. At the same time, the system writes the data identifier corresponding to the target data at the end of the second data file. This identifier is usually a unique identifier used to quickly locate the target data in subsequent operations. After successfully writing the data identifier, the system generates a second timestamp that records the exact time when the data identifier was written. Similar to the first timestamp, the second timestamp is also stored in a secure location for subsequent use. The system periodically or on demand scans the list of identifiers in the second data file to check if there are new write operation instructions to process. During the scan, the system checks the first timestamp and the second timestamp corresponding to each write operation instruction. By comparing these timestamps, the system can determine the processing priority of the write operation instruction. For example, if the first timestamp and the second timestamp of a certain write operation instruction are both earlier, then it may be an older instruction with a lower processing priority. On the contrary, if the timestamps are newer, then the instruction may be a newer instruction with a higher processing priority. According to the determined order of processing priorities, the system processes the fourth target data identifier and the fourth target data corresponding to each write operation instruction in sequence.
[0042] By generating and storing the first timestamp and the second timestamp respectively when writing the target data and the target data identifier, the system can accurately record the sequence of data writing. This helps to ensure data consistency and integrity in the concurrent write operation scenario, and avoid data conflicts and losses. When scanning the identifier list in the second data file, the system can determine the processing priority of the write operation instructions according to the first timestamp and the second timestamp. This strategy helps to ensure that the write operation instructions that arrive first are processed first, thus maintaining the order and correctness of the data. By processing the fourth target data identifier and the fourth target data corresponding to the write operation instructions in the order of the processing priority, the system can manage the write operation tasks more efficiently. This helps to reduce processing latency and improve the response speed and throughput of the system. This method can provide an effective control mechanism for concurrent write operations. By comparing timestamps, the system can judge the dependency relationships and conflict situations between write operations, and thus take appropriate measures to avoid data competition and inconsistencies. This method supports dynamic expansion of the first data file and the second data file. As the data volume grows, the system can optimize the write operation processing flow by adding new data files, adjusting the timestamp generation strategy, etc., to ensure the scalability and flexibility of the system. In case of a failure, the system can restore the consistent state of the data according to the timestamp information. By comparing the data and timestamps at different time points, the system can determine which data is valid and which data needs to be restored or deleted, thus ensuring the integrity and accuracy of the data.
[0043] This embodiment also discloses a data management system under high-concurrency write operations. Figure 2 It is a schematic diagram of the modules of the data management system under high-concurrency write operations disclosed in the embodiments of the present application. As Figure 2 shown, the system includes a creation module 201, an analysis module 202, an execution module 203, and a deletion module 204, where: The creation module 201 is configured to create a first data file and a second data file. The first data file is used to store data, and the second data file is used to store data identifiers. The analysis module 202 is configured to analyze the write operation instruction to determine the type of the write operation instruction in response to receiving the write operation instruction sent by the client. The execution module 203 is configured to write the target data to the first data file and / or write the target data identifier to the second data file according to the type of the write operation instruction, and return the change status information of the first data file and / or the second data file to the client. The deletion module 204 is configured to scan the identification list in the second data file at preset time intervals and delete the data at the target positions in the first data file, where the target positions are the positions of the data in the first data file corresponding to the identifications in the identification list.
[0044] Optionally, the execution module 203 is configured to: When the type of the write operation instruction is an insertion operation, extract the first target data to be inserted from the write operation instruction and write the first target data to the end of the first data file; When the type of the write operation instruction is a modification operation, extract the first target data identification of the data to be modified and the modified second target data from the write operation instruction, write the second target data to the end of the first data file, and write the first target data identification to the end of the second data file; When the type of the write operation instruction is a deletion operation, extract the second target data identification of the data to be deleted from the write operation instruction and write the second target data identification to the end of the second data file.
[0045] Optionally, the deletion module 204 is configured to: Scan the identification list in the second data file to identify all the third target data identifications marked for deletion; Locate the corresponding target positions in the first data file according to the third target data identifications; Remove the third target data corresponding to the target positions from the first data file and determine the size of the first data file after removing the third target data; Merge multiple first data files according to the size.
[0046] Optionally, the deletion module 204 is configured to: Divide multiple first data files into a small file list, a medium file list, and a large file list according to a preset first threshold and a preset second threshold. The small file list contains first data files smaller than the preset first threshold, the large file list contains first data files greater than or equal to the preset second threshold, and the medium file list contains first data files greater than or equal to the preset first threshold and smaller than the preset second threshold; Select a first file from the large file list and match a second file from the small file list so that the difference between the size of the third file obtained by merging the first file and the second file and the preset size is less than the error threshold until the large file list is empty or the small file list is empty; When the large file list is empty, merge the fourth file in the medium file list with the remaining files in the small file list; When the small file list is empty, merge the files in the medium file list.
[0047] Optionally, the deletion module 204 is configured to: Monitor the access records of the data, and divide multiple first data files into high-heat data files and low-heat data files according to a preset heat threshold. The high-heat data files are data files with access frequencies greater than the preset heat threshold, and the low-heat data files are data files with access frequencies less than or equal to the preset heat threshold; Merge the low-heat data files in chronological order from early to late to form multiple third data files, and merge the high-heat data files in chronological order from early to late to form multiple fourth data files. Both the third data files and the fourth data files are smaller than the preset size.
[0048] Optionally, the system further includes a query module, and the query module is configured to: When receiving a query request, retrieve a candidate data set from the first data file according to the query request; Use the index to find the first candidate data marked as deleted from the second data file, and delete the first candidate data from the candidate data set to obtain a second candidate data set. The index is constructed based on the data identifiers in the second data file; Sort and / or aggregate the second candidate data set to meet the requirements of the query request, and encapsulate and return the processed result to the client.
[0049] Optionally, the system further includes a processing module, and the processing module is configured to: After writing the target data at the end of the first data file, generate and store a first timestamp. After writing the target data identifier at the end of the second data file, generate and store a second timestamp; When scanning the identifier list in the second data file, determine the processing priority of the write operation instruction according to the first timestamp and the second timestamp, and process the fourth target data identifier and the fourth target data corresponding to the write operation instruction in the order of the processing priority.
[0050] It should be noted that: when the device provided in the above embodiments realizes its functions, only the division of the above function modules is used for illustration. In practical applications, the above functions can be allocated to different function modules according to needs, that is, the internal structure of the device is divided into different function modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments belong to the same concept. For the specific implementation process, please refer to the method embodiments and will not be elaborated here.
[0051] This embodiment also discloses an electronic device. Referring to Figure 3 , the electronic device may include: at least one processor 301, at least one communication bus 302, a user interface 303, a network interface 304, and at least one memory 305.
[0052] Among them, the communication bus 302 is used to realize the connection and communication between these components.
[0053] Among them, the user interface 303 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 303 may further include a standard wired interface and a wireless interface.
[0054] Among them, the network interface 304 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).
[0055] Among them, the processor 301 may include one or more processing cores. The processor 301 connects various parts within the entire server through various interfaces and lines, and executes various functions of the server and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 305, and by calling the data stored in the memory 305. Optionally, the processor 301 may be implemented in at least one of the hardware forms of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor 301 may integrate one or several combinations of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, the user interface, and application programs, etc.; the GPU is responsible for the rendering and drawing of the content to be displayed on the display screen; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor 301 and may be implemented separately by a single chip.
[0056] Among them, the memory 305 may include a Random Access Memory (RAM), or may include a Read-Only Memory. Optionally, the memory 305 includes a non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 305 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above method embodiments, etc.; the data storage area may store data involved in the above method embodiments. Optionally, the memory 305 may also be at least one storage device located far from the aforementioned processor 301. As Figure 3 shown, in the memory 305 as a computer storage medium, there may be included an operating system, a network communication module, a user interface module, and an application program of the data management method under high-concurrency write operations.
[0057] In Figure 3 the electronic device shown, the user interface 303 is mainly used to provide an input interface for the user to obtain user input data; while the processor 301 can be used to call the application program of the data management method under high-concurrency write operations stored in the memory 305. When executed by one or more processors 301, the electronic device is caused to execute the method of one or more of the above embodiments.
[0058] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0059] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0060] In several embodiments provided by this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some service interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.
[0061] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0062] In addition, each functional unit in various embodiments of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0063] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory 305 and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of this application. And the aforementioned memory 305 includes: various media such as USB flash drives, mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0064] The above are only exemplary embodiments of the present disclosure and should not be used to limit the scope of the present disclosure. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. Those skilled in the art will easily think of other implementation schemes of the present disclosure after considering the disclosure of the specification. This application aims to cover any variations, uses, or adaptive changes of the present disclosure, and these variations, uses, or adaptive changes follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not recorded in the present disclosure. The specification and embodiments are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.
Claims
1. A data management method under high-concurrency write operations, characterized in that, Applied to a data management platform, the method includes: Create a first data file and a second data file, where the first data file is used to store data and the second data file is used to store data identifiers; In response to receiving a write operation instruction sent by a client, analyze the write operation instruction to determine the type of the write operation instruction; Write target data to the first data file and / or write target data identifiers to the second data file according to the type of the write operation instruction, and return the change status information of the first data file and / or the second data file to the client; Scan the identifier list in the second data file every preset time, and delete the data at the target position in the first data file, where the target position is the position of the data in the first data file corresponding to the identifier in the identifier list.
2. The data management method under high-concurrency write operations according to claim 1, wherein The writing of target data to the first data file and / or writing of target data identifiers to the second data file according to the type of the write operation instruction includes: When the type of the write operation instruction is an insert operation, extract the first target data to be inserted from the write operation instruction, and write the first target data to the end of the first data file; When the type of the write operation instruction is a modification operation, extract the first target data identifier of the data to be modified and the modified second target data from the write operation instruction, write the second target data to the end of the first data file, and write the first target data identifier to the end of the second data file; When the type of the write operation instruction is a deletion operation, extract the second target data identifier of the data to be deleted from the write operation instruction, and write the second target data identifier to the end of the second data file.
3. The data management method under high-concurrency write operations according to claim 1, characterized in that, The scanning of the identifier list in the second data file every preset time and deleting the data at the target position in the first data file includes: Scan the identifier list in the second data file to identify all third target data identifiers marked for deletion; Locate the corresponding target position in the first data file according to the third target data identifier; Remove the third target data corresponding to the target position from the first data file, and determine the size of the first data file after removing the third target data; Merge multiple first data files according to the size.
4. The data management method under high-concurrency write operations according to claim 3, wherein The merging of multiple first data files according to the size includes: Divide multiple first data files into a small file list, a medium file list, and a large file list according to a preset first threshold and a preset second threshold. The small file list contains first data files smaller than the preset first threshold, the large file list contains first data files greater than or equal to the preset second threshold, and the medium file list contains first data files greater than or equal to the preset first threshold and smaller than the preset second threshold; Select a first file from the large file list, and match a second file from the small file list such that the difference between the size of the third file obtained by merging the first file and the second file and a preset size is less than an error threshold until the large file list is empty or the small file list is empty; When the large file list is empty, merge the fourth file in the medium file list with the remaining files in the small file list; When the small file list is empty, merge the files in the medium file list.
5. The data management method under high-concurrency write operations according to claim 3, characterized in that The merging of the multiple first data files according to the size includes: Monitoring the access records of the data, and dividing the multiple first data files into high-heat data files and low-heat data files according to a preset heat threshold. The high-heat data files are data files with an access frequency greater than the preset heat threshold, and the low-heat data files are data files with an access frequency less than or equal to the preset heat threshold; Merge the low-heat data files in chronological order from early to late to form multiple third data files, and merge the high-heat data files in chronological order from early to late to form multiple fourth data files. Both the third data files and the fourth data files are smaller than the preset size.
6. The data management method under high-concurrency write operations according to claim 1, characterized in that The method further includes: When a query request is received, retrieve a candidate data set from the first data file according to the query request; Use an index to find the first candidate data marked as deleted from the second data file, and delete the first candidate data from the candidate data set to obtain a second candidate data set. The index is constructed based on the data identifiers in the second data file; Perform sorting and / or aggregation processing on the second candidate data set to meet the requirements of the query request, and encapsulate and return the processed result to the client.
7. The data management method under high-concurrency write operations according to claim 1, characterized in that The method further includes: After writing the target data at the end of the first data file, generate and store a first timestamp. After writing the target data identifier at the end of the second data file, generate and store a second timestamp; When scanning the identifier list in the second data file, determine the processing priority of the write operation instruction according to the first timestamp and the second timestamp, and process the fourth target data identifier and the fourth target data corresponding to the write operation instruction in the order of the processing priority.
8. A data management system under high-concurrency write operations, characterized in that, Including a creation module, an analysis module, an execution module, and a deletion module, where: The creation module is configured to create a first data file and a second data file. The first data file is used to store data, and the second data file is used to store data identifiers; The analysis module is configured to analyze the write operation instruction to determine the type of the write operation instruction in response to receiving the write operation instruction sent by the client; The execution module is configured to write target data to the first data file and / or write target data identifiers to the second data file according to the type of the write operation instruction, and return the change status information of the first data file and / or the second data file to the client; A deletion module, configured to scan the identifier list in the second data file at preset time intervals, and delete the data at the target location in the first data file, where the target location is the location of the data in the first data file corresponding to the identifiers in the identifier list.
9. An electronic device, characterized in that, It includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. Both the user interface and the network interface are used to communicate with other devices. The processor is used to execute the instructions stored in the memory, so that the electronic device executes the method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions, and when the instructions are executed, the method according to any one of claims 1-7 is executed.