Data processing method and device, electronic equipment, storage medium and program product
By configuring index information for files and data blocks, efficient delayed deletion of files is achieved while reducing system overhead, solving the problems of high system overhead and accidental deletion in the prior art.
Patent Information
- Application Number
- CN202510363257.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-09-23
AI Technical Summary
The existing technology has a large system overhead in the data deletion process, cannot efficiently achieve delayed deletion of files, and is prone to accidental deletion.
By configuring the first index information and the second index information for the file, the status of the file and the status of the data block are marked respectively, and the file is delayed and deleted when the first time arrives, thereby reducing system overhead.
It achieves efficient delayed deletion of files while reducing system overhead, avoids accidental deletion, and simplifies the file management process.
Smart Images

Figure CN120687029A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to data storage technology, and in particular to a data processing method, device, electronic device, storage medium and program product. Background Art
[0002] With the advancement of computer technology, data management, including the storage, reading, writing, and deletion of massive amounts of data, has become a critical issue in data storage. Data deletion, in particular, is often performed periodically or on demand to save storage space. Therefore, efficient data deletion has become a hot topic of research. Summary of the Invention
[0003] The embodiments of the present application provide a data processing method, device, electronic device, storage medium and program product, which can achieve delayed deletion of files while reducing system overhead.
[0004] The technical solution of the embodiment of the present application is implemented as follows:
[0005] This embodiment of the present application provides a data processing method, the method comprising:
[0006] In response to a deletion request for a file, first index information of the file is determined, and a first state in the first index information is marked as a deletion state; a first data block storing the file is determined based on the first index information, and second index information of the first data block is determined; a second state in the second index information is marked as a deletion state, and a first time for deleting the file is configured in the second index information; when the first time arrives, the file stored in the first data block is deleted.
[0007] An embodiment of the present application provides a data processing device, including:
[0008] a first marking module, configured to determine, in response to a deletion request for a file, first index information of the file, and mark a first state in the first index information as a deleted state;
[0009] a determining module, configured to determine a first data block storing the file based on the first index information, and determine second index information of the first data block;
[0010] a second marking module, configured to mark the second state in the second index information as a deleted state, and configure a first time of deleting the file in the second index information;
[0011] A deleting module is configured to delete the file stored in the first data block when the first time arrives.
[0012] An embodiment of the present application provides an electronic device, comprising:
[0013] a memory for storing computer-executable instructions or computer programs;
[0014] The processor is used to implement the data processing method provided in the embodiment of the present application when executing the computer-executable instructions or computer programs stored in the memory.
[0015] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the data processing method provided in the embodiment of the present application when executed by a processor.
[0016] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, the data processing method provided in the embodiment of the present application is implemented.
[0017] The embodiments of the present application have the following beneficial effects:
[0018] In the above manner, since the first index information is configured for the file, and the first index information is configured with a first status that represents whether the file is deleted, in this way, when responding to a deletion request for the file, the first status of the first index information of the file is marked as a deletion status, so that the file can be made invisible to the user. However, in order to prevent accidental deletion of the file, the first data block storing the file can be determined by the first index information, and the second status in the second index information configured for the first data block can be marked as a deletion status, but at the same time, the first time for deleting the file is configured in the second index information, so that the file will be actually deleted in the first data block when the first time arrives. In this way, when delayed deletion of the file is required, delayed management of the file can be achieved by managing two index information, without the need to transfer the file to be delayed deleted to a designated area, thereby reducing the system overhead when processing data for the file and achieving the effect of delayed deletion. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is a schematic diagram of the data processing system architecture provided by an embodiment of the present application;
[0020] Figure 2 is a structural diagram of an electronic device provided in an embodiment of the present application;
[0021] Figure 3 This is a first flow chart of the data processing method provided in an embodiment of the present application;
[0022] Figure 4AThis is a first schematic diagram of data processing provided by an embodiment of the present application;
[0023] Figure 4B This is a second schematic diagram of data processing provided by the embodiment of the present application
[0024] Figure 5 It is a flowchart of the data space recovery method provided in an embodiment of the present application.
[0025] It should be pointed out that the above-mentioned "first" and "second" are only used to distinguish different solutions, and do not represent the degree of distinction between the advantages and disadvantages of the solutions or the priority in the implementation process. DETAILED DESCRIPTION
[0026] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0027] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0028] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0029] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0030] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0031] The relevant data collection and processing in the embodiments of this application should be strictly in accordance with the requirements of relevant laws and regulations when applied in examples, and the informed consent or separate consent of the personal information subject should be obtained. Subsequent data use and processing should be carried out within the scope of authorization of laws and regulations and the personal information subject.
[0032] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0033] 1) In response, it is used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more operations executed can be real-time or have a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations executed are executed.
[0034] 2) Files, also known as object files, are the basic units of data storage and are usually composed of a set of ordered bytes. These bytes can represent different types of data such as text, images, audio, and video. Depending on the content and purpose of the file, files can include but are not limited to the following types: text files, binary files, image files, audio files, table files, configuration files, etc.
[0035] 3) Data block (extent) is used to describe a contiguous area of storage space. It is commonly associated with file systems, databases, and disk storage. In file systems, a data block refers to one or more contiguous disk blocks (or sectors) that are allocated to a file to store its data. In disk storage solutions, a data block can refer to a contiguous area within a logical volume composed of multiple disks. This usage is common in disk array technology, where multiple physical disks are combined into a logical unit, and a data block is the contiguous space divided within this logical unit.
[0036] The embodiments of the present application provide a data processing method, device, electronic device, storage medium and program product, which can achieve delayed deletion of files while reducing system overhead and traffic.
[0037] See also Figure 1 , Figure 1 This is an architectural diagram of a data processing system 100 provided in an embodiment of the present application. To support a data processing application, a terminal 401 is connected to a server 200 via a network 300. The network 300 may be a wide area network or a local area network, or a combination of the two.
[0038] Terminal 401 is configured to receive a file deletion instruction initiated by a user, generate a file deletion request, and send the file deletion request to server 200. In response to the file deletion request, server 200 determines first index information of the file and marks a first state in the first index information as a deletion state; determines a first data block storing the file based on the first index information and determines second index information of the first data block; marks a second state in the second index information as a deletion state and configures a first time for deleting the file in the second index information; and deletes the file stored in the first data block when the first time arrives.
[0039] In some embodiments, the terminal 401 can be implemented as various types of terminals such as a laptop computer, a tablet computer, a desktop computer, a set-top box, a smart phone, a smart speaker, a smart watch, a smart TV, a car terminal, etc., and can also be implemented as a server.
[0040] In some embodiments, the server 200 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal and the server may be connected directly or indirectly via wired or wireless communication, which is not limited in the embodiments of the present application.
[0041] See also Figure 2 , Figure 2 is a structural diagram of an electronic device 400 provided in an embodiment of the present application, Figure 2 The electronic device 400 shown includes: at least one processor 410, a memory 450, at least one network interface 420 and a user interface 430. The various components in the electronic device 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 440 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 440 is not shown in FIG. Figure 2 Various buses are labeled as bus system 440 .
[0042] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0043] The user interface 430 includes one or more output devices 431 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0044] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.
[0045] The memory 450 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0046] In some embodiments, the memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.
[0047] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;
[0048] A network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include Bluetooth, Wi-Fi, and Universal Serial Bus (USB);
[0049] a presentation module 453 for enabling presentation of information via one or more output devices 431 (e.g., a display screen, a speaker, etc.) associated with the user interface 430 (e.g., a user interface for operating peripheral devices and displaying content and information);
[0050] The input processing module 454 is configured to detect one or more user inputs or interactions from one of the one or more input devices 432 and to translate the detected inputs or interactions.
[0051] In some embodiments, the data processing device provided in the embodiments of the present application can be implemented in software. Figure 2 The data processing device 455 stored in the memory 450 is shown. This device can be software in the form of a program or plug-in, and includes the following software modules: a first marking module 4551, a determination module 4552, a second marking module 4553, and a deletion module 4554. These modules are logical and can be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.
[0052] In other embodiments, the data processing device provided in the embodiments of the present application can be implemented in hardware. As an example, the data processing device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the data processing method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs) or other electronic components.
[0053] In some embodiments, the terminal or server can implement the data processing method provided by the embodiment of the present application by running various computer executable instructions or computer programs. For example, computer executable instructions can be commands, machine instructions or software instructions at the microprogram level. The computer program can be a native program or software module in the operating system; it can be a local (Native) application (APPlication, APP), that is, a program that needs to be installed in the operating system to run, such as a file management APP; it can also be a small program that can be embedded in any APP, that is, a program that only needs to be downloaded to a browser environment to run. In short, the above-mentioned computer executable instructions can be instructions in any form, and the above-mentioned computer program can be an application, module or plug-in in any form.
[0054] Below, the data processing method provided by the embodiment of the present application will be described with reference to the accompanying drawings. As mentioned above, the electronic device that implements the data processing method of the embodiment of the present application can be a terminal 401 or a server 200.
[0055] Below, the data processing method of the embodiment of the present application is explained by taking the execution subject as the server 200 as an example. A file management APP is installed in the terminal 401, and the files are stored in the server 200. The user can upload, download, read and write files in the file management APP of the terminal 401. If the user initiates an operation instruction for the file (such as a deletion instruction) to the file management APP of the terminal 401, the terminal 401 can send an operation request to the server 200 through the network 300 based on the user's operation instruction, so that the server 200 performs relevant processing on the file based on the operation request.
[0056] See also Figure 3 , Figure 3 This is a first flow chart of the data processing method provided in the embodiment of the present application, which will be combined with Figure 3 The steps shown are explained.
[0057] In step 101, in response to a deletion request for a file, first index information of the file is determined, and a first state in the first index information is marked as a deletion state.
[0058] Here, a file is the basic unit for storing data, generally an object file, which is a file compiled from the original file uploaded by the user or object. It is usually composed of a set of ordered bytes, which can represent different types of data such as text, pictures, audio, and video. Depending on the content and purpose of the file, the file can include but is not limited to the following types: text files, binary files, image files, audio files, table files, configuration files, etc.
[0059] In some embodiments, first index information can be constructed for a file in the following manner: in response to a save request for the file, the file is divided into multiple data shards; according to the principle of storing discontinuous data shards in the same data block, multiple data shards are stored in the data block; based on the correspondence between the multiple data shards and the data blocks, the first index information is constructed, and the first index information includes a first state marked as a normal state, and a first mapping relationship.
[0060] The first mapping relationship is used to indicate the corresponding relationship between each data slice and its corresponding data block.
[0061] In actual implementation, the storage area of a data processing system can be divided into multiple data nodes, each of which includes at least one data block, which can store data fragments of a file. Here, a data block is a unit storage area obtained by dividing the storage area. Each data block has a unique identifier, which is hereinafter referred to as a block identifier. When dividing the storage area, several data blocks can be evenly divided according to the storage space of the storage area, and the storage space (i.e., data space) of each data block is the same; or fixed-size data blocks can be set, and the storage space can be divided into multiple data blocks according to the fixed size.
[0062] As an example, assuming that the size of each data block is pre-set to 128MB, after dividing the storage area into multiple data blocks according to the fixed size and assigning a unique block identifier to each data block, multiple data nodes can be set and multiple data blocks can be assigned to each data node. The data blocks in each data node are not continuous data blocks.
[0063] In actual implementation, for file storage, a file can be divided into multiple parts, i.e., multiple data shards, according to a pre-set sharding strategy. Each data shard is assigned a unique identifier, which is hereinafter referred to as a shard identifier. The sharding strategy can be any of the following: uniform sharding, which evenly divides the file into fixed-size data shards; fixed-size sharding, which divides the file into multiple fixed-size data shards based on the size of the data shard; content-based sharding, which divides the file into multiple data shards based on the file's content or specific field values.
[0064] Then, the multiple data shards divided for the file are distributedly stored in the corresponding data blocks. Here, the distributed storage principle is based on the principle of storing non-contiguous data shards in the same data block. That is, the data shards of the same file can be stored in different data blocks of different data nodes as much as possible. For example, data shard 1 of the file is stored in data block 1 of data node 1, data shard 2 of the file is stored in data block 2 of data node 2, and data shard 3 of the file is stored in data block 3 of data node 3. In other words, multiple consecutive data shards of the same file are usually not stored in a data block, but multiple data shards of the same file that are separated may be stored in the same data block.
[0065] For example, assuming there are 16 data nodes in a cluster, each containing 12 hard disks, and each hard disk allocated 4 data blocks, the number of data blocks allocated for the data processing system's storage space is 4*16*12=768. If a 400MB file is divided into 120 data shards, each of these 120 data shards can be stored on different data blocks. If a 4GB file is divided into 1024 data shards, data shards 1 through 768 can be stored sequentially on the 768 data blocks. Then, the remaining data shards 769 through 1024 can be stored again, starting from the first data block. In other words, data shards 769 through 1024 can be stored sequentially on the 768 data blocks.
[0066] Through the above storage method, since the reading of data shards for files is generally batch sequential reading and the number of concurrency is also limited, there will generally not be simultaneous reading of intervening data shards in the same data block (such as data shard 1 and data shard 769). In this way, the data shards of the file can be discretely distributed to different data nodes, thereby achieving high concurrent access to the file.
[0067] In actual implementation, security attributes can also be set for each data block. The security attributes are used to characterize the importance of the data slices stored in the data block. When a file is divided into multiple data slices, the importance of each data slice can also be set according to the importance of the file content carried by each data slice. Based on the importance of the data slices and the security attributes of each data block, the data slices are stored in appropriate data blocks. For example, data slices with high importance are stored in data blocks with high security attributes, and data slices with low importance are stored in data blocks with low security attributes.
[0068] In actual implementation, for each file, when the file is saved based on the above method, a first mapping relationship between the data shards and the data blocks storing the data shards can be constructed based on the corresponding relationship between the data shards of the file and the data blocks storing the data shards. That is, a mapping relationship is established between the shard identifiers of the data shards and the block identifiers of the data blocks storing the data shards, and a corresponding first state is configured for the file. The first state is used to characterize the current state of the file, such as a normal state, a deleted state, etc. The first index information of the file is constructed based on the constructed first mapping relationship and the first state. Each file has corresponding first index information. Here, when storing a file, the first state in the first index information constructed for the file is usually marked as a normal state, indicating that the file is stored normally.
[0069] As an example, assume that the file a user wants to save is divided into 8 data slices, and the slice identifiers of each data slice are recorded as slice1, slice2, ..., slice8, respectively. Each data slice is stored in data blocks 1 to 8, and the block identifiers of data blocks 1 to 8 are respectively extent1, extent2, ..., extent8. Based on the correspondence between each data slice and the data block storing the data slice, the first mapping relationship in the constructed first index information includes:
[0070] (my_bucket / file1.txt / slice1->extent1,0MB,4MB)
[0071] (my_bucket / file1.txt / slice2->extent2,4MB,4MB)
[0072] (my_bucket / file1.txt / slice3->extent3,0MB,4MB)
[0073] (my_bucket / file1.txt / slice4->extent4,6MB,4MB)
[0074] (my_bucket / file1.txt / slice5->extent5,4MB,4MB)
[0075] (my_bucket / file1.txt / slice6->extent6,20MB,4MB)
[0076] (my_bucket / file1.txt / slice7->extent7,32MB,4MB)
[0077] (my_bucket / file1.txt / slice8->extent8,120MB, 2MB)
[0078] Take (my_bucket / file1.txt / slice1->extent1,0MB,4MB) as an example. In "my_bucket / file1.txt / slice1", "my_bucket" is the bucket ID, "file1.txt" is the file name, and "slice1" is the slice ID of the data slice. "Extent1" is the block ID of the data block. "0MB" indicates that the starting position of the storage space of data slice slice1 in data block extent1 is 0, and "4MB" indicates that the data size of data slice slice1 is 4MB. That is, the data space range of 0-4MB in data block extent1 is the data storage range of data slice slice1.
[0079] In actual implementation, corresponding file metadata can also be configured for the saved files. The file metadata is used to record the relevant attribute information of the file, which may include the name of the file (or file identifier), the creation time of the file, the expiration time configured for the file, etc.
[0080] Therefore, for a saved file, after the file is saved, the file has corresponding file metadata and first index information, wherein the first index information includes the first state and the first mapping relationship. In order to better manage files, the file metadata and first index information of each file can be stored in a specially constructed metadata database.
[0081] As an example, continuing with the previous example, the file metadata and first index information of the file file1.txt may be stored in the metadata database in the following format:
[0082] [object_id: my_bucket / file1.txt, created_time: 0, expired time: 72, is_deleted: false, object_index: [
[0084] (my_bucket / file1.txt / slice1->extent1,0MB,4MB),
[0085] (my_bucket / file1.txt / slice2->extent2,4MB,4MB),
[0086] (my_bucket / file1.txt / slice3->extent3,0MB,4MB),
[0087] (my_bucket / file1.txt / slice4->extent4,6MB,4MB),
[0088] (my_bucket / file1.txt / slice5->extent5,4MB,4MB),
[0089] (my_bucket / file1.txt / slice6->extent6,20MB,4MB),
[0090] (my_bucket / file1.txt / slice7->extent7,32MB,4MB),
[0091] (my_bucket / file1.txt / slice8->extent8,120MB, 2MB)... ] 】
[0094] Among them, "object_id:my_bucket / file1.txt, created_time:0, expired time:72" represents the file metadata of file1.txt, and "object_id:my_bucket / file1.txt" represents the file identifier of the file. "created_time:0" represents the creation time of the file, and "expired time:72" represents the expiration time configured for the file, which means that the file expires 72 hours after the creation time. "is_deleted:false" represents the first state, "false" means it has not been deleted, that is, the normal state, and "true" means it has been deleted, that is, the deleted state. "object_index:[]" represents the first mapping relationship constructed for each data shard.
[0095] As an example, Figure 4A This is the first schematic diagram of data processing provided by the embodiment of the present application, see Figure 4A The processing unit 110 is used to receive a save request for a file sent by the terminal 401, and in response to the save request, obtain the file uploaded by the user, divide the 400MB file into 120 data slices based on a preset sharding strategy, and store the 120 data slices in 120 data blocks respectively. Figure 4A, storage area 120 includes 4 data nodes, each data node includes several data blocks, data node 1 includes data blocks with block identifiers of extent1, extent5...extentx1, data node 2 includes data blocks with block identifiers of extent2, extent6...extentx2, data node 3 includes data blocks with block identifiers of extent3, extent7...extentx3, and data node 4 includes data blocks with block identifiers of extent4, extent8...extent4. Among them, data slice slice1 is stored in data block extent1, data slice slice2 is stored in extent2...data slice slice120 is stored in extent120. After all the data slices of the file are stored in the data blocks, a first mapping relationship is constructed based on the correspondence between each data slice and the data blocks storing the data slices, and the file metadata and the first status of the file are configured. Then, the file metadata of the file and the first index information composed of the first mapping relationship and the first status are stored in the metadata database.
[0096] As an example, Figure 4A The processing unit 110 shown may be part of an object storage gateway (OGW), which is used to receive and respond to various operation requests for files from users through the terminal 401, and perform corresponding operations on the files based on the multiple data blocks of the multiple data nodes divided in the storage area 120, such as reading, saving, deleting, etc. Figure 4B , Figure 4B This is a second schematic diagram of data processing provided by the embodiment of the present application, see Figure 4B In order to improve the concurrent processing capability, the object storage gateway 150 is configured with multiple processing units (e.g., processing unit 1, processing unit 2...processing unit n). Users can initiate various operation requests for files through the program 130 for data management installed in the terminal 401 (e.g., file management APP, file management applet, etc.) or the web page 140 for data management in the terminal 401. After receiving the operation request, the object storage gateway 150 will allocate the operation request to the appropriate processing unit according to the load capacity of each processing unit to process the file based on the operation request. Among them, if the processing unit receives a save request for a file, it can store the file in each data block based on the preset sharding strategy, and construct the first index information and file metadata for the file; if the processing unit receives other operation requests for the file (except the save request), it can obtain the file metadata and the first index information of the file in the metadata database based on the file name (or file identifier), and perform relevant operations on the file according to the content in the file metadata and the first index information.
[0097] It is understood that a data block can store data shards of multiple files. To facilitate the management of each data block, corresponding index information is also constructed for each data block. The index information corresponding to the data block is used to represent the mapping relationship between the block identifier of the data block and the data shards stored in the data block. Because the data shards stored in a data block may come from different files, each stored data shard needs to be managed independently. In other words, a corresponding second state is configured for each data shard stored in the data block. The second state is used to indicate whether the data shard has been deleted from the data block.
[0098] For example, for data block extent1, if the data block stores 7 data shards, and these 7 data shards come from 7 different files, then based on the data shards stored in data block extent1, the index information constructed for data block extent1 may include the following 7 index records:
[0099] (extent1,0MB)->[my_bucket / file1.txt / slice1,deleted:fase,delay_expired_time:0],
[0100] (extent1,4MB)->[my_bucket / file2.txt / slice1,deleted:fase,delay_expired_time:0],
[0101] (extent1,8MB)->[my_bucket / file3.txt / slice1,deleted:fase,delay_expired_time:0],
[0102] (extent1,16MB)->[my_bucket / file4.txt / slice1,deleted:fase,delay_expired_time:0],
[0103] (extent1,18MB)->[my_bucket / file5.txt / slice1,deleted:fase,delay_expired_time:0],
[0104] (extent1,22MB)->[my_bucket / file6.txt / slice1,deleted:fase,delay_expired_time:0],
[0105] (extent1,26MB)->[my_bucket / file7.txt / slice1,deleted:true,delay_expired_time:12h]
[0106] Among them, the index record "(extent1,0MB)->[my_bucket / file1.txt / slice1,deleted:fase,delay_expired_time:0]" is used as an example for explanation. Here, the index record can represent the mapping relationship between the data block and the data slice, and can also include a second identifier and the time set for delayed deletion. Specifically, "extent1" represents the block identifier of the data block; "0MB" means that the data slice is stored starting from 0 of the data space of the data block; "my_bucket / file1.txt / slice1" means that the data slice slice1 is stored; "deleted:true" represents the second state configured for the stored data slice slice1, "false" means it has not been deleted, that is, the normal state, and "true" means it has been deleted, that is, the deleted state; "delay_expired_time:12h" represents the time set for delaying the deletion of the data slice.
[0107] As an example, for each data block, data block metadata of the data block can be constructed based on the attribute information of the data block (such as storage space and writable status), and index information can be constructed for the data block based on the data shards stored in the data block. The data block metadata and index information of each data block are stored in the metadata database in the following format:
[0108] [extent_id:extent1,used_bytes:4800KB,slices_number:11,is_writable:false,is_deleted:false,extent1 reverse index: [
[0110] (extent1,offset1-1)->[my_bucket / my_file.txt / slice1,deleted:fase,delay_expired_time:0],
[0111] (extent1,offset1-2)->[my_bucket / f1.txt / slice1,deleted:fase,delay_expired_time:0],
[0112] (extent1,offset1-3)->[my_bucket / aa.txt / slice3,deleted:fase,delay_expired_time:0],
[0113] (extent1,offset1-3)->[my_bucket / cc.txt / slice12,deleted:fase,delay_expired_time:0], ...
[0115] (extent1,offset1-44)->[my_bucket / dd.txt / slice1,deleted:true,delay_expired_time:0] ] 】
[0118] Among them, "extent_id: extent1, used_bytes: 4800KB (used storage space), slices_number: 11, is_writable: false" is used to represent the data block metadata. Here, "extent_id: extent1" represents the block identifier of the data block; "used_bytes: 4800KB" represents the data space used by the data block; "slices_number: 11" represents the number of data slices stored in the data block; "is_writable: false" indicates the writable state of the data block, "false" indicates the non-writable state, and "true" indicates the writable state. "is_deleted: false" indicates the fourth state configured for the data block, "false" indicates that the data block has not been deleted (that is, the normal state), and "true" indicates that the data block has been deleted, that is, the deleted state. "extent1 reverse index: []" represents the index information configured for each data slice stored in the data block.
[0119] Here, the writable state of a data block is explained. After a data block is created, data shards can be stored in the data block. At this time, the data block is in a writable state, that is, the writable state in the data block metadata is "true"; and when the data space in the data block is filled with stored data shards or other content, that is, when the data block is full, the data block is in a non-writable state. At this time, the writable state in the data block metadata of the data block can be adjusted to "false".
[0120] Continue to see Figure 4BThe recycling unit 160 is used to determine whether a data block needs to be space recycled based on the index information corresponding to each data block in the metadata database.
[0121] In actual implementation, if the processing unit receives a deletion request for a file, it can obtain the file metadata and the first index information of the file from the metadata database based on the file name (or file identifier) of the file, and mark the first status in the first index information from the normal status to the deleted status, so that when the user initiates other requests for the file after deletion (such as a read request), the user is prompted that the file has been deleted and cannot be viewed based on the fact that the first status in the first index information of the file is the deleted status.
[0122] Here, the deletion request may be initiated by the user through the terminal 401 for the object, or may be generated due to the expiration time in the file metadata of the file.
[0123] Continue to see Figure 3 , continue with the above step 101 for explanation.
[0124] In step 102, a first data block of a storage file is determined based on the first index information, and second index information of the first data block is determined.
[0125] In actual implementation, marking the first status in the first index information of a file as deleted only indicates to the user that the file is unreadable and does not actually delete the file. Therefore, it is necessary to find the first data block storing the file based on the first index information of the file and determine the second index information corresponding to the first data block in order to delete the file stored in the first data block.
[0126] In some embodiments, Figure 3 In step 102, "determining the first data block storing the file based on the first index information" can be implemented in the following manner: for each data shard of the file, determining the second data block storing the data shard according to the first mapping relationship; and combining the second data blocks storing each data shard to obtain the first data block of the file.
[0127] In actual implementation, since the file is divided into multiple data shards, each data shard may be stored in a different data block. Therefore, when determining the first data block storing the file, the second data block storing each data shard can be determined based on the first mapping relationship of each data shard in the first index information of the file, and the second data block storing each data shard can be determined as the first data block storing the file.
[0128] As an example, for the file file1.txt, the first index information of the file file1.txt is obtained. If, based on the first mapping relationship of each data slice in the first index information, the 8 data slices of the file file1.txt are determined to be stored in the 8 second data blocks of extent1 to extent8 respectively, then these 8 second data blocks can be determined as the first data blocks storing the file file1.txt.
[0129] Continue to see Figure 3 , continue with the above step 102 for description.
[0130] In step 103, the second status in the second index information is marked as a deleted status, and the first time of deleting the file is configured in the second index information.
[0131] In actual implementation, the user's request to delete a file may be a case of accidental deletion. If, after determining the first data block storing the file, the data fragment stored in the first data block for the file is directly deleted, it will become extremely difficult for the user to recover the file. Therefore, here, when receiving the user's deletion request, the first status in the first index information of the file can be directly marked as a deletion status to prompt the user that the file has been deleted; and for the first data block storing the file, a first time for delaying the deletion of the file is configured in the second index information of the first data block, and the second status in the second index information is marked as a deletion status, so as to inform the data processing system that the file needs to be deleted through the second status being a deletion status, but to inform the data processing system through the first time that the file needs to be deleted when the first time is reached.
[0132] In some embodiments, if all data slices of a file are stored in one data block or the file is divided into only one data slice, then Figure 3 Step 103 may be implemented by: determining a fourth mapping relationship in the second index information, where the fourth mapping relationship is used to indicate a correspondence between the file and the first data block; marking the second state of the file indicated by the fourth mapping relationship as a deleted state, and configuring a first time to delete the file.
[0133] In some cases, all data shards of a file may be stored in one data block, or the file may be divided into only one data block due to a small data volume. In this case, for the file, the number of second data blocks storing the data shards of the file is one, and the second data block can be directly used as the first data block. In this case, based on the file name (or file identifier), an index record corresponding to the file can be determined in the second index information of the first data block. The index record represents a fourth mapping relationship between the file and the first data block. After determining the fourth mapping relationship corresponding to the file, the second state in the index record can be adjusted from a normal state to a deleted state, that is, the second state is assigned a value of "true", and a first time for delayed deletion is configured for the file in the index record.
[0134] In some embodiments, if a file is divided into multiple data fragments, and the multiple data fragments of the file are not stored in the same data block, the first data block includes a second data block storing each data fragment, the second index information includes third index information of each second data block, and the third index information includes a second mapping relationship indicating the second data block and the corresponding data fragment; then Figure 3 The "marking the second state in the second index information as a deleted state" in step 103 can be implemented in the following way: for each second data block, the second state of the data shard indicated by the second mapping relationship in the third index information corresponding to the second data block is marked as a deleted state.
[0135] Here, the data slices indicated by the second mapping relationship are the data slices of the above-mentioned file.
[0136] In this case, multiple data slices of the file are not stored in the same data block. For example, if the file file1.txt is divided into three data slices (respectively recorded as slice1, slice2, and slice3), based on the first index information of the file file1.txt, the second data block storing the data slice slice1 is determined to be data block extent1, the second data block storing the data slice slice2 is data block extent2, and the second data block storing the data slice slice3 is data block extent3. Then, the first data block storing the file file1.txt includes data block extent1, data block extent2, and data block extent3, and the second index information of the first data block includes the third index information of data block extent1 (including the second mapping relationship between data block extent1 and the stored data slice slice1), the third index information of data block extent2 (including the second mapping relationship between data block extent2 and the stored data slice slice2), and the third index information of data block extent3 (including the second mapping relationship between data block extent3 and the stored data slice slice3). After determining the above information, the second state of each data shard indicated by the second mapping relationship can be marked as a deleted state, that is, assigned a value of "true", and a first time for delaying deletion of the data shard can be configured. The first time can be set based on actual needs.
[0137] For example, if it is determined that the index record corresponding to data slice slice2 in the third index information of data block extent2 is (extent2,0MB)->[my_bucket / file1.txt / slice2,deleted:fase,delay_expired_time:0], then the index record can be adjusted to (extent2,0MB)->[my_bucket / file1.txt / slice2,deleted:true,delay_expired_time:12] to indicate that data slice slice2 will be deleted after 12 hours. Similarly, the index record corresponding to data slice slice1 in the third index information of data block extent1 and the index record corresponding to data slice slice3 in the third index information of data block extent3 are modified in the same way, so that all data slices of file file1.txt are deleted simultaneously after 12 hours.
[0138] Continue to see Figure 3 , continue with the above step 103 for explanation.
[0139] In step 104, when the first time arrives, the file stored in the first data block is deleted.
[0140] In actual implementation, in response to a file deletion request, after adjusting the second state of the second index information of the first data block storing the file and setting a delayed deletion time based on the above method, each data fragment of the file can be deleted from the corresponding data block when the set first time arrives. After the data fragment is deleted, the first time of the corresponding index record in the second index information can be adjusted to 0, and the second state remains in the "true" state, i.e., the deleted state.
[0141] In the above manner, since the first index information is configured for the file, and the first index information is configured with a first status that represents whether the file is deleted, in this way, when responding to a deletion request for the file, the first status of the first index information of the file is marked as a deletion status, so that the file can be made invisible to the user. However, in order to prevent accidental deletion of the file, the first data block storing the file can be determined by the first index information, and the second status in the second index information configured for the first data block can be marked as a deletion status, but at the same time, the first time for delaying the deletion of the file is configured in the second index information, so that the file will be actually deleted in the first data block when the first time arrives. In this way, when delayed deletion of the file is required, delayed management of the file can be achieved by managing two index information, without the need to transfer the file to be delayed deleted to a designated area, thereby reducing the system overhead when processing data for the file and achieving the effect of delayed deletion.
[0142] In some embodiments, before the first time arrives, the user can also restore the file, and the terminal 401 can achieve file recovery in the following manner: in response to the recovery operation on the file, the second state in the second index information is adjusted from the deleted state to the normal state, and the first state in the first index information is adjusted from the deleted state to the normal state.
[0143] In actual implementation, after setting the first time for delayed deletion for each data shard of the file, the data shard will not be actually deleted from the data block before the first time arrives. Therefore, in this stage, if the user initiates a recovery operation for the file, in response to the recovery operation for the file, since the data shard has not changed in the corresponding data block, it is only necessary to adjust the second state of the corresponding data shard in the second index information from the deleted state to the normal state, and at the same time delete the first time set for delayed deletion. After completing the update of the second index information corresponding to all the second data blocks of the file, the first state in the first index information of the file can also be adjusted from the deleted state to the normal state.
[0144] When the first state of the first index information of the file indicates a normal state, it indicates that data recovery of the file is complete, that is, the user can use the file normally, for example, read the file.
[0145] In some embodiments, after deleting the file stored in the first data block, the terminal 401 can also release the space of the first data block in the following manner: based on the second state in the second index information, perform status detection on each data shard in the first data block; when the second state of each data shard is a deletion state, release the data space occupied by the first data block; when there is a data shard in each data shard whose second state is a normal state, and the number of data shards in the normal state is lower than the quantity threshold, migrate the data shards in the normal state in the first data block to the third data block, and release the data space occupied by the first data block.
[0146] In view of the storage method of data fragments of files in the above-mentioned embodiment, each data block may store data fragments of files with different life cycles. In this way, if there is at least one data fragment of a file with a long life cycle in a data block, the data fragment will occupy the data block for a long time, resulting in that even if the amount of data of the valid data fragments stored in the data block is small, it cannot be recycled to release the data space occupied by the data block.
[0147] Therefore, data blocks can be continuously checked based on their index information to determine whether the amount of valid data in the data block is less than a preset data volume threshold, or whether the number of valid data fragments stored in the data block is less than a threshold. If either of these two situations occurs, space reclamation can be performed on the data block. The reclamation process involves migrating the valid data fragments in the data block to other data blocks, and after the migration is complete, the data space occupied by the data block is released.
[0148] In actual implementation, taking the first data block of the storage file as an example, the status of each data shard in the first data block is detected based on the second status in each index record in the second index information of the first data block. Assuming that the second status of all index records in the second index information of the first data block is a deletion status, and all have reached the first time of the set delayed deletion, it can be determined that there are no valid data shards in the first data block, and the data space occupied by the first data block can be directly released without data migration. If it is determined through the second index information of the first data block that there are still data shards in the first data block whose second status is a normal status, then it is determined whether the number of data shards in the normal status is lower than the quantity threshold, or whether the data volume of the data shards in the normal status is lower than the data volume threshold. If it is lower than the quantity threshold or lower than the data volume threshold, the data shards in the normal status can be migrated to the third data block, and the data space occupied by the first data block can be released.
[0149] In actual implementation, when migrating data blocks, the idea of migration can be: after determining the data block that needs data space recovery (that is, data migration is required), place the data block in the recovery queue, read out the valid data fragments in each data block in the recovery queue, and write them into a newly created data block (that is, the third data block). In this way, after the data migration is completed, the data space of this batch of data blocks can be released, thereby improving the data space recovery efficiency and data space utilization, and reducing the data space fragmentation rate.
[0150] In some embodiments, after the space of the first data block is released, the first index information of the file can also be adjusted accordingly in the following manner: construct a third mapping relationship between the third data block and the migrated data shard, and determine the third state of the migrated data shard; add the third mapping relationship and the third state to the fourth index information of the third data block; for each migrated data shard, based on the fourth index information, update the first mapping relationship corresponding to the data shard in the first index information to the third mapping relationship.
[0151] In actual implementation, since some data fragments of the file are migrated from the first data block to the third data block, in order to enable the user to normally access the file, the first index information of the file needs to be updated in a timely manner.
[0152] In actual implementation, after the valid data shards in the first data block are migrated to the third data block, a third mapping relationship is constructed between the third data block and the migrated data shards. That is, a third mapping relationship is established between the block identifier of the third data block and the shard identifiers of each migrated data shard, and the third state of the migrated data shards is determined (usually a normal state). Then, the constructed third mapping relationship and the corresponding third state are added to the fourth index information of the third data block in the form of index records. In addition, for each migrated data shard in the file, the first mapping relationship corresponding to the data shard in the first index information of the file needs to be updated based on the third mapping relationship corresponding to the data shard in the fourth index information. That is, the block identifier of the first data block corresponding to the data shard in the first mapping relationship is updated to the block identifier of the third data block.
[0153] For example, see Figure 4B The data space release process of the data block can be executed by the recycling unit 160. The recycling unit 160 can be executed offline and set a detection cycle. Based on the detection cycle, the status of the data slices in each data block in the storage area is detected. When it is determined that the data block meets the preset space recycling conditions, the data block is placed in the recycling queue. The data blocks in the recycling queue are detected based on the detection cycle. When it is determined that the data block meets the recycling execution conditions, the valid data slices in the data block are migrated and the data space of the data block is released.
[0154] Figure 5 This is a flow chart of the data space recovery method provided in the embodiment of the present application, see Figure 5 ,The recovery unit can reclaim data space for the data block through the following steps.
[0155] S201 , performing status detection on each data block, and when it is determined that the data block meets the space reclamation condition, placing the data block meeting the space reclamation condition into a reclamation queue.
[0156] S202: Determine whether the number of data blocks in the recycling queue is 0.
[0157] S203: If the number is 0, continue to perform status detection on the data block based on the detection cycle.
[0158] For example, the recycling unit can determine every 10 minutes whether the number of data shards in the index information of each data block is less than 15, or whether the amount of data stored in each data block is less than 1MB. If either condition is met, a space recycling task is set for the data block and the data block is placed in the recycling queue. Otherwise, it continues to wait for the next test.
[0159] S204: If the quantity is not 0, extract 1 data block from the recycling queue.
[0160] S205: Determine whether the extracted data block meets the recycling execution condition.
[0161] Here, the data amount threshold and the quantity threshold in the reclamation execution condition are smaller than the data amount threshold and the quantity threshold in the space reclamation condition.
[0162] S206 , when it is determined that the extracted data block meets the recycling execution condition, the recycling unit performs data migration and data space release on the data block.
[0163] For example, when the recycling unit performs a data block space reclamation task, it can request a new data block from the memory area, sequentially append the remaining valid data fragments in the current data block to the new data block, and update the first index information of the file for each valid data fragment. Here, the data space of the new data block is determined by the amount of data to be recycled. If the total amount of valid data in the data blocks to be migrated in the recycling queue exceeds 128MB, a newer data block will be allocated to handle the data write.
[0164] S207: When it is determined that the extracted data block does not meet the recycling execution condition, the data block is put back into the recycling queue and the state detection of each data block is continued in the next detection cycle.
[0165] In an embodiment of the present application, when the life cycle of a file is extended or shortened, since the data shards within the file are discretely distributed in the data blocks of different data nodes, the characteristic of canceling the association between the data blocks and the file life cycle is achieved. Once the life cycle configuration of the file changes, it is only necessary to modify the expiration time field in the file metadata attributes, without moving the underlying data blocks, thereby reducing the system overhead caused by data migration.
[0166] In addition, in order to ensure the available space ratio of the underlying data, the recovery unit can flexibly configure the space recovery task execution strategy for the data block, which can include the following strategies: setting the space recovery task execution time range, selecting the off-peak period to execute the space recovery task according to the business load, and reducing the impact of deleted data on the read and write data; setting the space recovery task execution trigger condition to trigger forced execution when the remaining capacity of the underlying storage area or the remaining space ratio is less than a certain threshold, to ensure the available space capacity of the cluster; configuring the number of concurrent executions of space recovery tasks and limiting the number of data blocks deleted per second according to the cluster system load and capacity water level, taking into account both space recovery efficiency and reduced overhead; for the space recovery task of the data block, differentiate the execution according to the effective number of data blocks and the effective data volume, and delete the data blocks with a data volume or effective data fragmentation of 0 in their entirety to improve the recovery efficiency, and merge the data blocks with different severe fragmentation and small effective data volume into new data blocks in sequence, and release the old data blocks as a whole, which not only improves the recovery efficiency, but also reduces the fragmentation rate of the underlying hard disk storage space and improves disk performance.
[0167] The following continues to describe the exemplary structure of the data processing device 455 provided in the embodiment of the present application implemented as a software module. In some embodiments, such as Figure 2 As shown, the software modules stored in the data processing device 455 of the memory 450 may include:
[0168] The first marking module 4551 is configured to determine first index information of a file in response to a deletion request for the file, and mark a first state in the first index information as a deletion state.
[0169] The determination module 4552 is configured to determine a first data block of a storage file based on the first index information, and determine second index information of the first data block.
[0170] The second marking module 4553 is configured to mark the second status in the second index information as a deleted status, and configure the first time of deleting the file in the second index information.
[0171] The deletion module 4554 is configured to delete the file stored in the first data block when the first time arrives.
[0172] In some embodiments, the data processing device 455 also includes an index construction module for dividing a file into multiple data shards in response to a save request for the file; storing multiple data shards in a data block according to the principle of storing discontinuous data shards in the same data block; and constructing first index information based on the correspondence between multiple data shards and data blocks, the first index information including a first state marked as a normal state, and a first mapping relationship, the first mapping relationship being used to indicate the correspondence between each data shard and its corresponding data block.
[0173] In some embodiments, the determination module 4552 is further configured to determine, for each data shard of the file, a second data block storing the data shard according to the first mapping relationship; and combine the second data blocks storing each data shard to obtain the first data block of the file.
[0174] In some embodiments, the second marking module 4553 is also used to determine a fourth mapping relationship in the second index information, where the fourth mapping relationship is used to indicate the correspondence between the file and the first data block; mark the second state of the file indicated by the fourth mapping relationship as a deleted state, and configure the first time to delete the file.
[0175] In some embodiments, a file is divided into multiple data shards, the first data block includes a second data block storing each data shard, the second index information includes third index information of each second data block, and the third index information includes a second mapping relationship between the second data block and its corresponding data shard; the second marking module 4553 is also used to mark the second state of the data shard indicated by the second mapping relationship in the third index information corresponding to the second data block as a deleted state for each second data block.
[0176] In some embodiments, the data processing device 455 also includes a recovery module for adjusting the second state in the second index information from a deleted state to a normal state, and adjusting the first state in the first index information from a deleted state to a normal state, in response to a recovery operation on the file before the first time arrives.
[0177] In some embodiments, the first data block stores multiple data shards of a file; the data processing device 455 also includes a release module for performing status detection on each data shard in the first data block based on the second status in the second index information; when the second status of each data shard is a deleted status, the data space occupied by the first data block is released; when there is a data shard whose second status is a normal status in each data shard, and the number of data shards in the normal status is lower than the quantity threshold, the data shards in the normal status in the first data block are migrated to the third data block, and the data space occupied by the first data block is released.
[0178] In some embodiments, the release module is also used to construct a third mapping relationship between the third data block and the migrated data shards, and determine the third state of the migrated data shards; add the third mapping relationship and the third state to the fourth index information of the third data block; for each migrated data shard, based on the fourth index information, update the first mapping relationship of the corresponding data shard in the first index information to the third mapping relationship.
[0179] An embodiment of the present application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the data processing method described in the embodiment of the present application.
[0180] The embodiment of the present application provides a computer-readable storage medium in which computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the data processing method provided in the embodiment of the present application, for example, Figure 3 The data processing method is shown.
[0181] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or may be various devices including one or any combination of the above memories.
[0182] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0183] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).
[0184] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.
[0185] To sum up, through the embodiments of the present application, since the first index information is configured for the file, and the first index information is configured with the first status representing whether the file is deleted, in this way, when responding to a deletion request for the file, the first status of the first index information of the file is marked as a deletion status, so that the file can be made invisible to the user. However, in order to prevent accidental deletion of the file, the first data block storing the file can be determined by the first index information, and the second status in the second index information configured for the first data block can be marked as a deletion status, but at the same time, the first time for delaying the deletion of the file is configured in the second index information, so that the file will be actually deleted in the first data block when the first time arrives. In this way, when delayed deletion of the file is required, delayed management of the file can be achieved by managing two index information, without the need to transfer the file to be delayed deleted to a designated area, thereby reducing the system overhead when processing data for the file and achieving the effect of delayed deletion.
[0186] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.
Claims
1. A data processing method, characterized in that: The method comprises: In response to a deletion request for a file, determining first index information of the file, and marking a first state in the first index information as a deletion state; Determining a first data block storing the file based on the first index information, and determining second index information of the first data block; Marking the second state in the second index information as a deleted state, and configuring the first time of deleting the file in the second index information; When the first time arrives, the file stored in the first data block is deleted.
2. The method according to claim 1, characterized in that Before determining the first index information of the file, the method includes: In response to a request to save the file, dividing the file into a plurality of data slices; According to the principle of storing discontinuous data fragments in the same data block, multiple data fragments are stored in multiple data blocks; Based on the correspondence between the multiple data shards and data blocks, first index information is constructed, wherein the first index information includes a first state marked as a normal state and a first mapping relationship, and the first mapping relationship is used to indicate the correspondence between each data shard and its corresponding data block.
3. The method according to claim 1, characterized in that The file is divided into a plurality of data shards, the first data block includes a second data block storing each data shard, the second index information includes third index information corresponding to each second data block, and the third index information includes a second mapping relationship between the second data block and its corresponding data shard; The marking the second status in the second index information as a deleted status includes: For each second data block, the second state of the data slice indicated by the second mapping relationship in the third index information corresponding to the second data block is marked as a deleted state.
4. The method according to claim 1, wherein Before the first time arrives, the method further includes: In response to a recovery operation on the file, the second state in the second index information is adjusted from the deleted state to the normal state, and the first state in the first index information is adjusted from the deleted state to the normal state.
5. The method according to claim 1, wherein The first data block stores multiple data fragments of the file; After deleting the file stored in the first data block, the method further includes: Performing status detection on each data fragment in the first data block based on the second status in the second index information; When the second states of all data shards are in the deletion state, releasing the data space occupied by the first data block; If there are data shards whose second state is a normal state in each data shard, and the number of data shards in the normal state is lower than the quantity threshold, the data shards in the normal state in the first data block are migrated to the third data block, and the data space occupied by the first data block is released.
6. The method according to claim 5, characterized in that After migrating the data shards in the first data block that are in a normal state to the third data block, the method further includes: Constructing a third mapping relationship between the third data block and the migrated data shard, and determining a third state of the migrated data shard; Adding the third mapping relationship and the third state to the fourth index information of the third data block; For each migrated data shard, based on the fourth index information, the first mapping relationship corresponding to the data shard in the first index information is updated to the third mapping relationship.
7. A data processing device, characterized in that: The device comprises: a first marking module, configured to determine, in response to a deletion request for a file, first index information of the file, and mark a first state in the first index information as a deleted state; a determining module, configured to determine a first data block storing the file based on the first index information, and determine second index information of the first data block; a second marking module, configured to mark the second state in the second index information as a deleted state, and configure a first time of deleting the file in the second index information; A deleting module is configured to delete the file stored in the first data block when the first time arrives.
8. An electronic device, characterized in that: The electronic device comprises: a memory for storing computer-executable instructions or computer programs; The processor is configured to implement the data processing method according to any one of claims 1 to 6 when executing the computer-executable instructions or computer programs stored in the memory.
9. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the data processing method according to any one of claims 1 to 6 is implemented.
10. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the data processing method according to any one of claims 1 to 6 is implemented.