Data management method and device, electronic equipment and storage medium
By determining the target data in the object storage file and generating a sub-file of verification information, the problem that additional writes in the prior art cannot guarantee data security is solved, and the effect of data addition writes is achieved and data security is improved.
Patent Information
- Application Number
- CN202510482486.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-16
AI Technical Summary
In the prior art, the method of appending and writing to object storage files cannot ensure data security, especially in WORM mode, it is easier to threaten data security.
By receiving the client's write request, the target data to be written is determined based on the end position of the target file and the position in the write request, and a sub-file containing the verification information of the user ID and data information is generated. Finally, the sub-file is written into the target file with the end position of the target file as the starting position.
It realizes that without modifying or overwriting the data written in the target file, add write data to the object storage file to ensure the consistency of the data, and improve the security of data management through verification information, preventing malicious writes from damaging data security.
Smart Images

Figure CN120029554A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of the present specification relate to the field of database technology, and in particular, to a data management method and device, an electronic device, and a storage medium. Background Art
[0002] In the database field, append writing (also known as append writing) is a common data operation method, which means adding new data to the end of an existing data set instead of overwriting or modifying existing data; append writing can bring data integrity, efficient writing and other effects. Object storage is a data storage model without a directory structure. It can store and manage data in the form of objects, which is very suitable for storing unstructured data such as images, videos, and documents.
[0003] In the database field, there is often a need to write additional data to object storage files. However, the related art of writing additional data to object storage files cannot guarantee data security, especially when writing additional data to object storage files in Write Once Read Many (WORM) mode, it is more likely to threaten data security. Summary of the invention
[0004] In view of this, one or more embodiments of the present specification provide a data management method and device, an electronic device, and a storage medium.
[0005] To achieve the above objectives, one or more embodiments of this specification provide the following technical solutions: According to a first aspect of one or more embodiments of this specification, a data management method is proposed, the method comprising: Receiving a write request sent by a client for instructing to write data to be written into a target file, and determining target data in the data to be written based on an end position of data stored in the target file and a requested write position in the write request, wherein the target data is part or all of the data to be written; Generating verification information based on the user identification of the client and the data information of the target data, and generating a sub-file to be written based on the verification information and the target data; The end position of the target file is taken as the starting position, and the sub-file to be written is written into the target file.
[0006] In a possible embodiment of the present specification, determining the target data in the data to be written based on the end position of the data stored in the target file and the requested write position in the write request includes: If the offset of the requested write position is consistent with the offset of the end position of the target file, the data to be written is determined as the target data, wherein the offset is used to represent the offset relative to the start position of the target file.
[0007] In a possible embodiment of the present specification, the write request includes the data length of the data to be written; The determining the target data within the data to be written based on the end position of the target file and the requested write position in the write request comprises: If the offset of the requested write position is less than the offset of the end position of the target file, then based on the offset of the requested write position and the data length of the data to be written, determining the offset of the end position of the data to be written, wherein the offset is used to represent the offset relative to the start position of the target file; If the offset of the end position of the data to be written is greater than the offset of the end position of the data stored in the target file, the data in the data to be written whose offset is greater than the offset of the end position of the target file is determined as the target data.
[0008] In a possible embodiment of the present specification, determining the offset data in the data to be written whose offset is greater than the end position of the target file as the target data includes: Determine data in the data to be written whose offset is not greater than the offset of the end position of the target file, and perform consistency verification with the stored data corresponding to the same offset in the target file; If the result of the consistency verification is that the verification is passed, the data whose offset in the data to be written is greater than the offset of the end position of the target file is determined as the target data.
[0009] In a possible embodiment of the present specification, the generating verification information based on the user identification of the client and the data information of the target data includes: The user identification of the client and the data information of the target data are input into a preset verification function, and the verification value output by the verification function is used as verification information.
[0010] In a possible embodiment of the present specification, the generating the to-be-written sub-file based on the verification information and the target data includes: An initial sub-file is generated based on the target data, and the verification information is added to the header of the initial sub-file to obtain a sub-file to be written.
[0011] In a possible embodiment of this specification, the method further includes: The starting position of the target data is determined as the name of the subfile to be written into the subfile, and the length of the target data is determined as the data length of the subfile to be written into the subfile.
[0012] In a possible embodiment of this specification, the method further includes: receiving a read request sent by a client for instructing to read data to be read from a target file, and determining at least one target sub-file in the target file based on a data position range in the read request; For each target subfile, generating verification information based on the user identification of the client and data information of the data in the target subfile, and determining whether the target subfile is valid based on the generated verification information and the verification information in the target subfile; Based on all valid target sub-files and the data location range in the read request, the data to be read is returned to the client.
[0013] According to a second aspect of one or more embodiments of this specification, a data management device is provided, the device comprising: a target data module, configured to receive a write request sent by a client for instructing to write data to be written into a target file, and determine target data in the data to be written based on the end position of the data stored in the target file and the requested write position in the write request, wherein the start position of the target data is the end position of the data stored in the target file; A sub-file module, used to generate verification information based on the user identification of the client and the data information of the target data, and to generate a sub-file to be written based on the verification information and the target data; A writing module is used to write the sub-file to be written into the target file.
[0014] According to a third aspect of one or more embodiments of this specification, a computer program product is provided, comprising a computer program / instruction, which implements the steps of the method described in the first aspect when executed by a processor.
[0015] According to a fourth aspect of one or more embodiments of this specification, an electronic device is provided, including: processor; a memory for storing processor-executable instructions; The processor implements the method as described in the first aspect by running the executable instructions.
[0016] According to a fifth aspect of one or more embodiments of the present specification, a computer-readable storage medium is provided, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method described in the first aspect are implemented.
[0017] The technical solutions provided by the embodiments of the present specification may include the following beneficial effects: For the data management method provided by the embodiments of the present specification, a write request sent by a client for instructing to write data to be written into a target file is received, and based on the end position of the data already stored in the target file and the requested write position in the write request, target data is determined within the data to be written; verification information is generated based on the user identifier of the client and the data information of the target data, and a sub-file to be written is generated based on the verification information and the target data; starting from the end position of the target file, the sub-file to be written is written into the target file. Since the target data in the data to be written is written in the form of a sub-file, there is no need to perform operations such as modifying or overwriting the data already written in the target file, so as to achieve the effect of adding data after the already stored data, that is, to achieve the effect of a data append write operation; moreover, since the target data is written starting from the end position of the target file, there is no conflict between the target data and the data already written in the target file, thus ensuring data consistency; more importantly, this method adds verification information related to the user identifier of the client and the target data to the file to be written, so that each written sub-file can verify its validity. For example, when querying, it can be determined whether the sub-file is valid through the verification information in the sub-file, thus improving the security of data management and preventing maliciously written sub-files from damaging data security. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a flowchart of a data management method provided by an exemplary embodiment.
[0019] Figure 2 is a schematic diagram of the directory structure of a target file provided by an exemplary embodiment.
[0020] Figure 3 is a schematic diagram of data to be written in one case provided by an exemplary embodiment.
[0021] Figure 4 is a schematic diagram of data to be written in another case provided by an exemplary embodiment.
[0022] Figure 5 is a schematic diagram of data to be written in yet another case provided by an exemplary embodiment.
[0023] Figure 6FIG. 4 is a schematic diagram of data to be written in another situation provided by an exemplary embodiment.
[0024] Figure 7 is a flow chart of a data reading process provided by an exemplary embodiment.
[0025] Figure 8 is a schematic diagram of a data reading process provided by an exemplary embodiment.
[0026] Fig. 9 It is a schematic diagram of a sub-file provided by an exemplary embodiment without adding verification information.
[0027] Fig.10 It is a schematic diagram of writing a forged sub-file into a target file without adding verification information to the sub-file provided by an exemplary embodiment.
[0028] Fig.11 It is a schematic diagram of writing a forged sub-file into a target file for adding verification information to a sub-file provided by an exemplary embodiment.
[0029] Fig.12 It is a structural schematic diagram of a device provided by an exemplary embodiment.
[0030] Fig.13 is a block diagram of a data management device provided by an exemplary embodiment. DETAILED DESCRIPTION
[0031] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with one or more embodiments of this specification. Instead, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.
[0032] It should be noted that: in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may be combined into a single step for description in other embodiments.
[0033] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0034] First of all, it is explained that compared with the patent application with publication number CN117130995A submitted by the applicant on August 4, 2023, the present disclosure further improves its technical solution and solves the technical problems existing in its technical solution. The specific effects will be described in detail below in conjunction with the specific solution content.
[0035] In the related art, object storage, as a data storage model without a directory structure, can store and manage data in the form of objects, which is very suitable for storing unstructured data such as images, videos, and documents. When storing data through the object storage model, in order to meet the data storage compliance requirements, avoid data tampering, and meet the long-term retention and archiving needs of data, the object storage WORM mode is often used. As a data protection strategy in the object storage model, the object storage WORM mode requires that once an object is successfully created, it is not allowed to be deleted or overwritten within a period of time. The append write operation, as a data writing method that appends data to the end of the storage object, will inevitably involve the modification of the stored data. Therefore, the object storage WORM mode generally does not allow append write operations. Even if the append write is simulated in some way, data security and data consistency cannot be guaranteed.
[0036] Based on the above technical problems mentioned in the background technology, at least one embodiment of this specification provides a data management method, which can realize continuous writing of data in object storage files that do not support additional write operations, such as object storage files in WORM mode, and can ensure data consistency and security.
[0037] The method can be executed by a computing device, which may be a terminal device, such as a desktop computer, a portable computer, a tablet computer, a laptop computer, a smart phone, etc., or the computing device may be a server, such as one server, multiple servers, a server cluster, a cloud computing platform, etc. This specification does not limit the device type of the computing device.
[0038] It should be noted that when a computing device receives a data write request, it can use the data writing method provided in this specification to write the data to be written after the stored data in the object storage file, so as to achieve a continuous data writing effect similar to an append write operation.
[0039] Optionally, the object storage file can be used to store various types of unstructured data, such as image data, video data, audio data, document data, report data, log data, time series data, etc. Therefore, the data to be written involved in this specification may be image data, video data, audio data, document data, report data, log data, time series data, etc. This specification does not limit the specific type of data to be written.
[0040] The above is only an exemplary description of the application scenarios of this specification and does not constitute a limitation on the application scenarios of this specification. In more possible implementations, the solutions provided in this specification can be applied to more scenarios that require data writing, and this specification does not limit the specific application scenarios.
[0041] After introducing the application scenarios of this specification, the specific implementation process of this specification is introduced next.
[0042] Please refer to the attached Figure 1 , which exemplarily shows a flow chart of the data management method, including steps S101 to S103.
[0043] In step S101, a write request sent by a client to instruct writing data to be written into a target file is received, and target data is determined in the data to be written based on the end position of the data stored in the target file and the requested write position in the write request, wherein the target data is part or all of the data to be written.
[0044] The target file may be an object storage file for storing image data, video data, audio data, document data, report data, log data, time series data, etc., and correspondingly, the data to be written may be video data, audio data, document data, report data, log data, time series data, etc. Optionally, the target file may also be other types of files, and the data to be written may also be other types of data. This specification does not limit the specific types of the target file and the data to be written.
[0045] It should be noted that a write request may carry the data to be written, the file identifier of the target file to which the data is to be written (such as the file name, file ID, etc.), and the requested write position of the data to be written in the target file (i.e., indicating the position in the target file from which the data to be written is to be written), so that after receiving the write request, the computing device may directly obtain the file identifier of the target file and the requested write position of the data to be written in the target file based on the write request, and may thereby determine the target file from multiple data files maintained by the computing device based on the obtained file identifier, and obtain the end position of the stored data in the target file.
[0046] For example, the requested write position has an offset, which is used to represent its offset relative to the starting point (starting position) of the target file; the end position of the target file is the end position of the data stored therein, and the end position has an offset, which is used to represent its offset relative to the starting point (starting position) of the target file. Bytes can usually be used as the unit of the above offset.
[0047] This step can split the data to be written based on the requested write position and the end position of the data stored in the target file, so as to obtain the data to be written with the end position of the stored data as the starting position (that is, the data whose offset is greater than the offset of the end position of the stored data) as the target data to be actually written to the target file, thereby ensuring that there is no conflict between the data written this time and the data already written in the target file.
[0048] In step S102, verification information is generated based on the user identification of the client and the data information of the target data, and a sub-file to be written is generated based on the verification information and the target data.
[0049] Exemplarily, the user identification of the client and the data information of the target data may be input into a preset verification function, and the verification value output by the verification function may be used as the verification information.
[0050] For example, the user identifier can be the identifier of the user logged in on the client that sends the write request. After the user logs in based on the user identifier, the client can generate a write request in response to the received operation instruction (i.e., the instruction entered by the user through operating the client) and send it to the device that executes the method. Therefore, the write request carries the user identifier.
[0051] For example, the data information may be a data path, data length, file data (ie, data content), etc. of the target data.
[0052] For example, the verification function may be a hash function, such as an md5 function, etc. Then the verification value output by the verification function is a hash value.
[0053] Exemplarily, an initial sub-file may be generated based on the target data, and the verification information may be added to a header of the initial sub-file to obtain a sub-file to be written.
[0054] For example, the number of bytes occupied by the verification information is fixed, such as the verification information of the header of each file to be written occupies 8 bytes. In this way, the verification information can be read in the first 8 bytes of each sub-file in the target file when reading the data in the target file; and when reading the data in the target file based on the offset from the starting point of the target file, the interference of the header of each sub-file can be excluded when determining the data position, so as to accurately locate the data to be read.
[0055] It should be understood that when the verification information is added to the header of the initial sub-file, the version number of the sub-file can also be added synchronously, in which case the number of bytes occupied by the header increases, such as 16 bytes.
[0056] For example, the starting position of the target data may be determined as the name of the subfile to be written into the subfile, and the length of the target data may be determined as the data length of the subfile to be written into the subfile.
[0057] In step S103, the sub-file to be written is written into the target file with the end position of the target file as the starting position.
[0058] It should be understood that after writing the sub-file to be written, the end position of the data stored in the target file may be updated.
[0059] By implementing continuous writing of data in the target file in the form of sub-files, each new data to be written will be stored in the target file in the form of a sub-file, so that the target file can include multiple sub-files, the data stored in adjacent sub-files are continuous, and a sub-file will not be modified or overwritten after it is written, which meets the requirements of the WORM mode.
[0060] In addition, it should be noted that, precisely because the data written each time will be stored in the target file in the form of sub-files, the target file can be organized into a virtual folder, which can include multiple sub-files, and each time data is written, a new sub-file will be generated in the virtual folder based on the data to be written this time.
[0061] The data management method provided in the embodiments of the present specification receives a write request sent by a client to instruct writing data to be written into a target file, determines target data in the data to be written based on the end position of the data stored in the target file and the requested write position in the write request; generates verification information based on the user identifier of the client and the data information of the target data, and generates a sub-file to be written based on the verification information and the target data; and writes the sub-file to be written into the target file with the end position of the target file as the starting position. The method writes the target data in the data to be written in the form of a sub-file, so there is no need to modify or overwrite the data already written in the target file, and an effect of adding data to the stored data can be achieved, that is, an effect of a data append write operation can be achieved; and the target data is written with the end position of the target file as the starting position, so there is no conflict between the target data and the data already written in the target file, thereby ensuring the consistency of the data; more importantly, the method adds verification information related to the user identification of the client and the target data in the file to be written, so that the validity of each sub-file written can be verified. For example, when querying, the verification information in the sub-file can be used to determine whether the sub-file is valid, thereby improving the security of data management and avoiding malicious sub-files from destroying data security.
[0062] In some embodiments of the present disclosure, determining the target data in the data to be written based on the end position of the stored data in the target file and the requested write position in the write request in step S101 may be performed in the following manner: First, determine the positional relationship between the offset of the requested write position and the offset of the end position of the target file. The positional relationship between the two can have three possible results, the first is that the offset of the requested write position is greater than the offset of the end position of the target file, the second is that the offset of the requested write position is consistent with the offset of the end position of the target file, and the third is that the offset of the requested write position is less than the offset of the end position of the target file.
[0063] Next, based on the positional relationship between the offset of the requested write position and the offset of the end position of the target file, the target data is determined in the data to be written. That is, when the positional relationship between the offset of the requested write position and the offset of the end position of the target file is one of the above three possible results, the operation is performed in the following different ways respectively.
[0064] In a possible implementation, if the offset of the requested write position is greater than the offset of the end position of the target file, there is no need to obtain the target data from the data to be written, that is, the data to be written does not contain the target data.
[0065] It should be noted that if the offset of the requested write position is greater than the offset of the end position of the stored data, it indicates that the data writing process is non-continuous writing, which is an illegal operation and there is no need to write data. Therefore, there is no need to obtain the target data to generate the file to be written.
[0066] For ease of understanding, a specific example is used below to illustrate the situation where the requested write position is located after the end position of the stored data.
[0067] As described in the above embodiments, the target file may include multiple sub-files, and thus the target file may be understood as a folder, and the multiple sub-files included in the target file may be organized together according to the directory structure. Figure 2 , Figure 2 is a schematic diagram of a directory structure of a target file provided by an exemplary embodiment, such as Figure 2 As shown, the file name of the target file can be start_offset, and the target file can include 5 sub-files, and the file names of these 5 sub-files are 0, 100, 500, 750 and 980 respectively. Assuming that the sub-files with file names 0, 100 and 500 have been written with data, the data length of the sub-file with file name 0 is 100, the data length of the sub-file with file name 100 is 400, and the data length of the sub-file with file name 500 is 250. Since the present specification can write data in the form of sub-files each time, the data length of each sub-file in the target file is also the length of the data stored in each sub-file, that is, in the example Figure 2 In the target file shown, the length of the stored data (denoted as written_len) is 750, or in other words, the end position of the stored data is 750 bytes.
[0068] If the offset of the requested write position of the data to be written is 800 and the length of the data to be written is 200, the requested write position is located after the end position of the stored data, which is an illegal non-continuous write operation. Figure 3 , Figure 3 is a schematic diagram of a data writing process provided by an exemplary embodiment. Figure 3 That is, it shows the data writing process when the offset of the requested write position is greater than the offset of the end position of the target file, such as Figure 3As shown, subfile 0 with a data length of 100, subfile 100 with a data length of 400, and subfile 500 with a data length of 250 are stored in the target file in sequence. The length of the stored data in the target file is written_len=750. According to the data writing rule, this data writing should start writing data from the position of offset=750, but the requested write position of the data to be written (that is, newappend) is offset=800, and the length of the data to be written is len=200, offset>written_len, and this data writing operation is obviously a non-continuous write operation, so an error can be directly reported without writing data.
[0069] In another possible implementation, if the offset of the requested write position is consistent with the offset of the end position of the target file, the data to be written is determined as the target data.
[0070] It should be noted that if the offset of the requested write position is consistent with the offset of the end position of the target file, it indicates that the data writing process is continuous writing and the position where the data is to be written is empty. Therefore, the full amount of data in the data to be written can be directly used as the target data, that is, the full amount of data in the data to be written can be used as the data actually to be written this time.
[0071] Still in Figure 2 Take the target file shown in the figure as an example, if the offset of the requested write position of the data to be written is not 750 and the length of the data to be written is 200, the offset of the requested write position is consistent with the offset of the end position of the stored data, and this is a continuous write operation starting from the end position of the stored data. Figure 4 , Figure 4 is another schematic diagram of a data writing process provided by an exemplary embodiment. Figure 4 That is, it shows the data writing process when the offset of the requested write position is consistent with the offset of the end position of the stored data, such as Figure 4As shown, subfile 0 with a data length of 100, subfile 100 with a data length of 400, and subfile 500 with a data length of 250 are stored in the target file in sequence. The length of the stored data in the target file is written_len=750. According to the data writing rule, this data writing should start writing data from the position of offset=750, and the requested write position of the data to be written (that is, newappend) is also offset=750, the length of the data to be written is len=200, offset=written_len, and this data writing operation is a continuous write operation starting from the data stored position. Then, the full amount of data in the data to be written can be directly used as the target data to be written, and the following can be used. Figure 2 In the directory structure shown, name the new subfolder 750.
[0072] In another possible implementation, if the offset of the requested write position is less than the offset of the end position of the data stored in the target file, the offset of the end position of the data to be written is determined based on the offset of the requested write position and the data length of the data to be written; further, if the offset of the end position of the data to be written is greater than the offset of the end position of the target file, the data in the data to be written whose offset is greater than the offset of the end position of the target file is determined as the target data. The write request also includes the data length (len) of the data to be written, and the data length can also be in bytes.
[0073] It should be noted that the above embodiments mainly introduce the situations where the requested write position is consistent with the end position of the stored data, and the requested write position is after the end position of the stored data. For these two situations, the end position of the data to be written must also be after the end position of the stored data. Therefore, there is no need to consider the positional relationship between the end position of the data to be written and the end position of the stored data. However, for the situation where the requested write position is before the end position of the stored data, the positional relationship between the end position of the data to be written and the end position of the stored data cannot be determined. Therefore, in the case where the requested write position is before the end position of the stored data, it is necessary to further subdivide multiple situations for consideration.
[0074] In one possible case, if the offset of the end position of the data to be written is less than or equal to the offset of the end position of the stored data, or, there is no need to obtain the target data from the data to be written.
[0075] It should be noted that in this case, if the end position of the data to be written is before the end position of the stored data or the end position of the data to be written is the same as the end position of the stored data, it indicates that all positions to be written with data this time have already stored data. And since it is the object storage WORM mode currently, modifying or overwriting the already written data is not allowed, so there is no need to obtain the target data.
[0076] Still taking the data writing in the target file as shown in Figure 2 as an example, if the offset of the requested write position of the data to be written is 750 and the length of the data to be written is 200, then the offset of the requested write position is the same as the offset of the end position of the stored data. At this time, it is a continuous write operation starting from the end position of the stored data. Refer to Figure 5 , Figure 5 which is another schematic diagram of the data writing process provided by an exemplary embodiment. Figure 5 That is, it shows the data writing process when the offset of the requested write position is less than the offset of the end position of the stored data and the offset of the end position of the data to be written is less than the offset of the end position of the stored data. As shown in Figure 5 , sub - file 0 with a data length of 100, sub - file 100 with a data length of 400, and sub - file 500 with a data length of 250 are stored in the target file in sequence. The length of the stored data in the target file is written_len = 750. According to the data writing rule, this data writing should start writing data from the position of offset = 750. However, the offset of the requested write position of the data to be written this time (that is, new append) is offset = 400, and the length of the data to be written is len = 200. Then it can be determined that the end write position is 600. offset < written_len and offset + len < written_len. This data writing operation needs to modify or overwrite the stored data, and this data writing method is not allowed in the object storage WORM mode, so there is no need to obtain the target data.
[0077] The above Figure 5 only takes the case where the end position of the data to be written is before the end position of the stored data as an example for explanation. The case where the end position of the data to be written is the same as the end position of the stored data is the same in principle and will not be elaborated here.
[0078] In another possible case, if the offset of the end position of the data to be written is greater than the offset of the end position of the stored data, then the data in the data to be written whose offset is greater than the offset of the end position of the target file can be determined as the target data.
[0079] For example, data in the data to be written whose offset is not greater than the offset of the end position of the target file is determined, and consistency verification is performed with the stored data corresponding to the same offset in the target file; if the result of the consistency verification is passed, the data in the data to be written whose offset is greater than the offset of the end position of the target file is determined as the target data.
[0080] Among them, the data in the data to be written whose offset is less than the offset of the end position of the target file has an offset from offset to written_len, and a length of written_len-offset, and the stored data corresponding to the same offset in the target file is the data in the stored data with an offset from offset to written_len, and a length is also written_len-offset; the data in the data to be written whose offset is greater than the end position of the target file is written_len to offset+len, and a length of offset+len-written_len.
[0081] Optionally, if the data in the data to be written whose offset is less than the end position of the target file is different from the stored data corresponding to the same offset in the target file, the result of the consistency verification is failed, and there is no need to obtain the target data. It should be noted that this situation indicates that if data needs to be written, the stored data needs to be modified or overwritten, and this data writing method is not allowed in the object storage WORM mode, so there is no need to obtain the target data.
[0082] Optionally, if the data in the data to be written whose offset is less than the end position of the target file is consistent with the stored data corresponding to the same offset in the target file, the result of the consistency verification is passed, and the data in the data to be written whose offset is greater than the end position of the target file is determined as the target data. It should be noted that in this case, there is no need to modify or overwrite the stored data, and there is no need to write the data in the data to be written that is located before the end position of the data stored in the target file. Instead, the data in the data to be written that is located after the end position of the data stored in the target file can be directly written to the target file, because the data in the data to be written that is located before the end position of the data stored in the target file has been stored in the target file.
[0083] Still in Figure 2Taking the example of writing data into the target file shown, if the offset of the requested write position of the data to be written is 700 and the length of the data to be written is 200, the offset of the requested write position is less than the offset of the end position of the stored data. At this time, data needs to be written before the end position of the stored data. Refer to Figure 6 , Figure 6 which is another schematic diagram of the data writing process provided by an exemplary embodiment. Figure 6 That is, it shows the data writing process in the case where the offset of the requested write position is less than the offset of the end position of the stored data and the offset of the end write position of the data to be written is greater than the offset of the end position of the stored data. As Figure 6 shown, sub-file 0 with a data length of 100, sub-file 100 with a data length of 400, and sub-file 500 with a data length of 250 are sequentially stored in the target file. The length of the stored data in the target file is written_len = 750. According to the data writing rule, this data writing should start writing data from the position of offset = 750. However, the offset of the requested write position of the data to be written this time (that is, new append) is offset = 700, and the length of the data to be written is len = 200. Then, it can be determined that the end write position is 900. offset < written_len and offset + len > written_len. The positions where data needs to be written during this data writing operation can be divided into two parts: 700 to 750 and 750 to 900. Among them, there is already stored data in the part from 700 to 750. If data needs to be written in this part of the position, the stored data needs to be modified or overwritten, and this data writing method is not allowed in the object storage WORM mode. Therefore, this part of the data does not need to be considered when generating the file to be written; for the part from 750 to 900, this part in the target database is empty and data can be directly written. Therefore, the part from 750 to 900 in the data to be written can be obtained as the target data.
[0084] In the above embodiment, the data to be written is split by the offset of the requested write position of the data to be written this time and the offset of the end position of the target file, so that the data part after the end position of the stored data in the data to be written is used as the target data. For the data part at the position where data has already been stored and needs to be written, there is no need to repeat the writing, thereby reducing the data processing volume during the data writing process and improving the data writing speed and data writing efficiency.
[0085] In some embodiments of the present disclosure, the method further includes the data reading process as Figure 7 shown, including steps S701 to S703.
[0086] In step S701, a read request sent by a client for instructing to read data to be read from a target file is received, and at least one target sub-file is determined in the target file based on a data location range in the read request.
[0087] The data position range in the read request includes a read start position and a read end position. The read start position can be represented by an offset in bytes, and the read end position can be represented by an offset in bytes.
[0088] It should be noted that the target file may include multiple sub-files, and the data in different sub-files are written in different batches, that is, each written data will be stored as a sub-file. For example, after receiving a read request, the sub-files in the target file can be listed, and the list results include the file names of all sub-files and the data length of each sub-file; then, according to the list results of all sub-files, the sub-file where the reading start point is located can be found first, and then the traversal can be continued backward until the sub-file where the reading end point is found. These two sub-files and all sub-files between them are the target sub-files.
[0089] The above process of determining the target subfile by traversal is only an exemplary implementation. In more possible implementations, other positioning methods may be used to determine the target subfile. For example, a binary search method may be used to determine the target subfile.
[0090] In step S702, for each target subfile, verification information is generated based on the user identification of the client and data information of the data in the target subfile, and whether the target subfile is valid is determined based on the generated verification information and the verification information in the target subfile.
[0091] Exemplarily, this step can generate verification information in the same manner as when writing data into the target file. For example, the user identification of the client and the data information of the target data can be input into a preset verification function, and the verification value output by the verification function is used as the verification information. For example, the user identification can be the identification of the user logged in on the client that sends the read request. After the user logs in based on the user identification, the client can generate a read request in response to the received operation instruction (i.e., the instruction input by the user through the operation of the client) and send it to the device that executes the method, so the read request carries the user identification. For example, the data information can be the data path, data length, file data (i.e., data content), etc. of the target data. For example, the verification function can be a hash function, such as the md5 function, etc. Then the verification value output by the verification function is a hash value.
[0092] If the verification information generated in this step is consistent with the verification information in the target subfile, it is determined whether the target subfile is valid; if the verification information generated in this step is inconsistent with the verification information in the target subfile, it is determined whether the target subfile is invalid.
[0093] In step S703, based on all valid target sub-files and the data location range in the read request, the data to be read is returned to the client.
[0094] By obtaining valid target sub-files, the required data portion is read from each valid target sub-file, and then the read data portions are combined to obtain the complete data to be read.
[0095] For ease of understanding, a specific example is used below to illustrate the data reading process provided in this specification.
[0096] See also Figure 8 , Figure 8 is a flowchart of a data reading process provided by an exemplary embodiment, such as Figure 8 As shown, the target data file includes 4 sub-files, among which the file name of the first sub-file is 0 and the data length (or file length) is 100, the file name of the second sub-file is 100 and the data length is 400, the file name of the third sub-file is 500 and the data length is 250, and the file name of the fourth sub-file is 750 and the data length is 150.
[0097] If the data reading instruction is read(offset=300,len=500), that is, to read data of length 500 starting from data storage position 300, the subfile corresponding to the position of offset=300 will be found first, that is, the subfile with file name 100, and then continue to traverse backward until the subfile corresponding to the position of offset+len=800, that is, the subfile with file name 750, then the target subfiles are the subfile with file name 100, the subfile with file name 500 and the subfile with file name 750.
[0098] Then, based on the user identification of the client sending the read request and the data information of each target subfile, verification information of each target subfile is generated; and then the generated verification information is compared with the verification information in the header of the target subfile to see whether they are consistent.
[0099] After comparison, the three target sub-files are all valid target sub-files, so they can be divided into the following three parts to read the data to be read from the three target sub-files respectively: File name: 100, read (offset = 200, len = 200), that is, from sub-file 100, get the data starting from data storage position 200 with a data length of 200; File name: 500, read (offset = 0, len = 250), that is, from sub-file 500, get the data starting from data storage position 0 and with a data length of 250; File name: 750, read (offset = 0, len = 50), that is, from sub-file 750, get the data starting from data storage position 0 and with a data length of 50; By combining these three parts of data, the complete data to be read can be obtained.
[0100] Compared with the patent application with publication number CN117130995A submitted by the applicant on August 4, 2023, the data management method provided in this specification sets verification information in the header of a sub-file each time it is written, so that when the user reads the data, the data can be avoided from being affected by erroneous data or even illegal malicious data, thereby improving data security, especially when appending to object storage files in WORM mode.
[0101] For ease of understanding, the above effect is described below with a specific example.
[0102] Please refer to the attached Fig. 9 , which exemplarily shows the result of writing data without adding verification information to the sub-file header, Fig. 9 Four subfiles have been written into the target file, namely, subfile with file name 0 and length 100, subfile with file name 100 and length 400, subfile with file name 500 and length 250, and subfile with file name 750 and length 150. The length of the target file is 900.
[0103] Please refer to the attached Fig.10 If an attacker forges a sub-file named 900 and writes it into the target file, the user will mistakenly believe that the length of the target file is 1000 when reading the target file, thereby reading the data in the sub-file forged by the attacker; and if the previous data is deleted or updated in the forged data, it will affect the data correctness and cause a data security incident.
[0104] Please refer to the attached Fig.11If the target file is written according to the data provided in this manual, verification information will be added to the headers of the four sub-files in the target file; if an attacker forges a sub-file named 900 and writes it into the target file, when the user reads the target file, the forged sub-file will be eliminated because it cannot pass the verification of the verification information, thereby preventing the user from reading incorrect data and improving data security.
[0105] Fig.12 is a schematic structural diagram of a device provided by an exemplary embodiment. Fig.12 At the hardware level, the device includes a processor 1202, an internal bus 1204, a network interface 1206, a memory 1208, and a non-volatile memory 1210, and may also include hardware required for other tasks. One or more embodiments of this specification may be implemented based on software, such as the processor 1202 reading the corresponding computer program from the non-volatile memory 1210 into the memory 1208 and then running it. Of course, in addition to the software implementation, one or more embodiments of this specification do not exclude other implementations, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0106] Please refer to Fig.13 , the data management device can be used for Fig.12 The data management device may include: The target data module 1301 is used to receive a write request sent by a client to instruct writing the data to be written into the target file, and determine the target data in the data to be written based on the end position of the data stored in the target file and the requested write position in the write request, wherein the start position of the target data is the end position of the data stored in the target file; The sub-file module 1302 is used to generate verification information based on the user identification of the client and the data information of the target data, and to generate a sub-file to be written based on the verification information and the target data; The writing module 1303 is used to write the sub-file to be written into the target file.
[0107] In a possible embodiment of the present specification, the target data module is used to, when determining the target data in the data to be written based on the end position of the stored data in the target file and the requested write position in the write request, to: If the requested write position is the end position of the data stored in the target file, the data to be written is determined as the target data.
[0108] In a possible embodiment of the present specification, the write request includes the data length of the data to be written; The target data module is used to determine the target data in the data to be written based on the end position of the stored data in the target file and the requested write position in the write request, and to: If the requested write position is before the end position of the data stored in the target file, determining the end position of the data to be written based on the requested write position and the data length of the data to be written; If the end position of the data to be written is after the end position of the data stored in the target file, the data in the data to be written that is after the end position of the data stored in the target file is determined as the target data.
[0109] In a possible embodiment of the present specification, when the target data module is used to determine the data in the data to be written that is located after the end position of the data stored in the target file as the target data, it is used to: If the data in the data to be written that is located before the end position of the data stored in the target file is the same as the data at the same position in the target file, the data in the data to be written that is located after the end position of the data stored in the target file is determined as the target data.
[0110] In a possible embodiment of the present specification, when the sub-file module is used to generate verification information based on the user identification of the client and the data information of the target data, it is used to: The user identification of the client and the data information of the target data are input into a preset verification function, and the verification value output by the verification function is used as verification information.
[0111] In a possible embodiment of the present specification, when the sub-file module is used to generate the sub-file to be written based on the verification information and the target data, it is used to: An initial sub-file is generated based on the target data, and the verification information is added to the header of the initial sub-file to obtain a sub-file to be written.
[0112] In a possible embodiment of the present specification, the device further includes a naming module, which is used to: The starting position of the target data is determined as the name of the subfile to be written into the subfile, and the length of the target data is determined as the data length of the subfile to be written into the subfile.
[0113] In a possible embodiment of the present specification, the device further includes a reading module, which is used to: receiving a read request sent by a client for instructing to read data to be read from a target file, and determining at least one target sub-file in the target file based on a data position range in the read request; For each target subfile, generating verification information based on the user identification of the client and data information of the data in the target subfile, and determining whether the target subfile is valid based on the generated verification information and the verification information in the target subfile; Based on all valid target sub-files and the data location range in the read request, the data to be read is returned to the client.
[0114] One or more embodiments of the present specification also propose a computer program product, including a computer program / instruction, which implements the steps of the method provided in the first aspect when the computer program / instruction is executed by a processor.
[0115] One or more embodiments of the present specification also propose a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0116] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, which may be in the form of a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email transceiver, a game console, a tablet computer, a wearable device or a combination of any of these devices.
[0117] In a typical configuration, a computer includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0118] Memory may include non-permanent storage in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0119] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include temporary computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0120] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0121] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0122] The terms used in one or more embodiments of this specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of this specification. The singular forms of "a", "said" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0123] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0124] It should be understood that although the terms first, second, third, etc. may be used to describe various information in one or more embodiments of this specification, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0125] The above description is merely a preferred embodiment of one or more embodiments of the present specification and is not intended to limit one or more embodiments of the present specification. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of the present specification shall be included in the scope of protection of one or more embodiments of the present specification.
Claims
1. A data management method, the method include: Receiving a write request sent by a client for instructing to write data to be written into a target file, and determining target data within the data to be written based on an end position of the target file and a requested write position in the write request, wherein the target data is part or all of the data to be written; Generating verification information based on the user identification of the client and the data information of the target data, and generating a sub-file to be written based on the verification information and the target data; The end position of the target file is taken as the starting position, and the sub-file to be written is written into the target file.
2. The data management method according to claim 1, wherein the target data is determined within the data to be written based on the end position of the target file and the requested write position in the write request, include: If the offset of the requested write position is consistent with the offset of the end position of the target file, the data to be written is determined as the target data, wherein the offset is used to represent the offset relative to the start position of the target file.
3. The data management method according to claim 1, wherein the write request includes the data length of the data to be written; determining the target data within the data to be written based on the end position of the target file and the requested write position in the write request, include: If the offset of the requested write position is less than the offset of the end position of the target file, then based on the offset of the requested write position and the data length of the data to be written, determining the offset of the end position of the data to be written, wherein the offset is used to represent the offset relative to the start position of the target file; If the offset of the end position of the data to be written is greater than the offset of the end position of the target file, the data in the data to be written whose offset is greater than the offset of the end position of the target file is determined as the target data.
4. The data management method according to claim 3, wherein the offset data in the data to be written whose offset is greater than the end position of the target file is determined as the target data, include: Determine data in the data to be written whose offset is not greater than the offset of the end position of the target file, and perform consistency verification with the stored data corresponding to the same offset in the target file; If the result of the consistency verification is that the verification is passed, the data whose offset in the data to be written is greater than the offset of the end position of the target file is determined as the target data.
5. The data management method according to claim 1, wherein the verification information is generated based on the user identification of the client and the data information of the target data. include: The user identification of the client and the data information of the target data are input into a preset verification function, and the verification value output by the verification function is used as verification information.
6. The data management method according to claim 1, wherein the sub-file to be written is generated based on the verification information and the target data. include: An initial sub-file is generated based on the target data, and the verification information is added to the header of the initial sub-file to obtain a sub-file to be written.
7. The data management method according to claim 6, further comprising: include: The starting position of the target data is determined as the name of the subfile to be written into the subfile, and the length of the target data is determined as the data length of the subfile to be written into the subfile.
8. The data management method according to claim 1, further comprising: include: receiving a read request sent by a client for instructing to read data to be read from a target file, and determining at least one target sub-file in the target file based on a data position range in the read request; For each target subfile, generating verification information based on the user identification of the client and data information of the data in the target subfile, and determining whether the target subfile is valid based on the generated verification information and the verification information in the target subfile; Based on all valid target sub-files and the data location range in the read request, the data to be read is returned to the client.
9. A data management device, the device include: a target data module, configured to receive a write request sent by a client for instructing to write data to be written into a target file, and determine target data in the data to be written based on an end position of the data stored in the target file and a requested write position in the write request, wherein the target data is part or all of the data to be written; A sub-file module, used to generate verification information based on the user identification of the client and the data information of the target data, and to generate a sub-file to be written based on the verification information and the target data; The writing module is used to write the sub-file to be written into the target file with the end position of the target file as the starting position.
10. The data management device according to claim 9, further comprising a reading module, configured to: receiving a read request sent by a client for instructing to read data to be read from a target file, and determining at least one target sub-file in the target file based on a data position range in the read request; For each target subfile, generating verification information based on the user identification of the client and data information of the data in the target subfile, and determining whether the target subfile is valid based on the generated verification information and the verification information in the target subfile; Based on all valid target sub-files and the data location range in the read request, the data to be read is returned to the client.
11. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 8.
12. An electronic device, include: processor; a memory for storing processor-executable instructions; The processor implements the method according to any one of claims 1 to 8 by running the executable instructions.
13. A computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Video file index information construction method and apparatus, and video file index information query method and apparatus
CN106933974A
Data processing method and device, storage medium and equipment
CN115168304A
File processing method and device, electronic equipment and readable storage medium
CN115495020A
Data management method and related device
CN115658622A
Data processing method and device, equipment and medium
CN117130995A