Structured data storage method, electronic equipment and computer program product

Through memory mapping technology, structured data is converted into multiple entities in the virtual memory space and written to disk files, solving the time delay and performance overhead problems during data storage in the prior art, and achieving more efficient data persistence.

CN120216501APending Publication Date: 2025-06-27HANGZHOU QULIAN TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510192099.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When storing structured data into a relational database, the prior art needs to be converted into a data structure inside the database engine and serialized, resulting in high time delay and system performance overhead.

Method used

Through memory mapping technology, structured data is converted into multiple entities in the virtual memory space and written to disk files, avoiding the limitations of the internal mechanism of the database system and directly implementing memory updates to disk storage.

Benefits of technology

Reduces the time delay and system performance overhead of data persistence, and improves the efficiency and performance of data storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216501A_ABST
    Figure CN120216501A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data storage, and provides a structured data storage method, electronic equipment and a computer program product. According to the method, to-be-stored structured data is converted into a plurality of entities described in a first virtual memory space according to a defined data structure, the data structure comprises a fixed-length field and a non-fixed-length field, and the content of each entity comprises data of the fixed-length field and index information of the non-fixed-length field. The index information is the offset of the storage position of the data of the corresponding non-fixed-length field relative to the initial address of the second virtual memory space; the first virtual memory space is a virtual space corresponding to the first file in a memory; the second virtual memory space is a virtual space corresponding to the second file in the memory; and writing the contents of the plurality of entities into a first file, and writing the data of the non-fixed-length fields of the plurality of entities into a second file. By adopting the method, the time delay of data persistence and the system performance overhead can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data storage, and in particular, to a method for storing structured data, an electronic device, and a computer program product. Background Art

[0002] Currently, structured data generated by business systems is generally stored using various relational databases such as MySQL. However, when structured data enters the engine management space of the database system, it usually needs to be converted into the internal data structure of the engine and serialized before being persisted, resulting in high time latency and additional system performance overhead. Summary of the Invention

[0003] In view of this, embodiments of this application provide a method for storing structured data, an electronic device, and a computer program product, which can reduce the time latency of data persistence and system performance overhead.

[0004] The first aspect of the embodiments of this application provides a method for storing structured data, including:

[0005] Obtain the structured data to be stored;

[0006] Convert the structured data into multiple entities described in the first virtual memory space according to the defined data structure; wherein, the data structure includes fixed-length fields and non-fixed-length fields, the content of each entity includes the data of the fixed-length fields and the index information of the non-fixed-length fields, and the index information is the offset of the storage position of the data of the corresponding non-fixed-length field relative to the initial address of the second virtual memory space; the first virtual memory space is a virtual space corresponding to the first file in the disk in the memory constructed by memory mapping; the second virtual memory space is a virtual space corresponding to the second file in the disk in the memory constructed by memory mapping;

[0007] Write the content of the multiple entities into the first file, and write the data of the non-fixed-length fields of the multiple entities into the second file.

[0008] The technical solution of the embodiment of the present application creates a first file and a second file in a disk, and respectively constructs a first virtual memory space corresponding to the first file in the memory and a second virtual memory space corresponding to the second file in the memory through memory mapping; after obtaining the structured data to be stored, according to the defined data structure, the structured data is converted into multiple entities described in the first virtual memory space, the data structure includes fixed-length fields and non-fixed-length fields, the content of each entity includes the data of the fixed-length field and the index information of the non-fixed-length field, and the index information is the offset of the storage position of the data of the corresponding non-fixed-length field relative to the initial address of the second virtual memory space; then, the content of multiple entities is written into the first file, and the data of the non-fixed-length fields of multiple entities is written into the second file. The above process obtains the virtual memory space corresponding to the disk file through memory mapping, and the virtual address of the virtual memory space can be directly mapped to obtain the structured data stored in the disk file, so that the limitation of the internal mechanism of the database system can be broken away, and the time delay of data persistence and the system performance overhead can be reduced to a certain extent.

[0009] In an implementation manner of the embodiment of the present application, the second file includes multiple storage blocks; writing the data of the non-fixed-length fields of multiple entities into the second file includes:

[0010] Obtaining a bitmap file corresponding to the second file; wherein, each bit of the bitmap file is respectively used to represent the occupancy situation of each storage block of the second file;

[0011] According to the bitmap file, determining target storage blocks that can write data from multiple storage blocks, and writing the data of the non-fixed-length fields of multiple entities into the target storage blocks.

[0012] In an implementation manner of the embodiment of the present application, obtaining the bitmap file corresponding to the second file includes:

[0013] According to the virtual address in the third virtual memory space, obtaining the bitmap file from the disk; wherein, the third virtual memory space is constructed through memory mapping, and is the virtual space corresponding to the bitmap file in the memory.

[0014] In an implementation manner of the embodiment of the present application, the method further includes:

[0015] Determining storage blocks to be sorted out from each storage block of the second file;

[0016] Determining a target element object corresponding to the storage block to be sorted out from each element object included in the invalid character array; wherein, each element object corresponds to each storage block of the second file one by one, and each element object records the number of invalid characters of all the data written in the corresponding storage block;

[0017] If the number of invalid characters recorded by the target element object exceeds a preset threshold, a data reorganization operation is performed on the storage block to be reorganized.

[0018] In one implementation manner of the embodiment of the present application, the method further includes:

[0019] When any storage block of the second file is written with data, a first byte character is added in front of the written data;

[0020] When the length of any written data in the any storage block changes, the first byte character in front of the any data is modified to a second byte character, and the number of invalid characters recorded in the element object corresponding to the any storage block is updated.

[0021] In one implementation manner of the embodiment of the present application, each element object also records the respective entity identifiers and field identifiers corresponding to the data of each non-fixed-length field written in the corresponding storage block; performing a data reorganization operation on the storage block to be reorganized includes:

[0022] Determine the valid data with the first byte character in front in the storage block to be reorganized, and write the valid data into a new storage block;

[0023] Update the element object corresponding to the new storage block according to the entity identifier and field identifier corresponding to the valid data.

[0024] In one implementation manner of the embodiment of the present application, after performing a data reorganization operation on the storage block to be reorganized, it further includes:

[0025] Determine the storage blocks that do not need to be reorganized from the respective storage blocks of the second file;

[0026] Move the storage blocks that do not need to be reorganized to the initial position of the storage blocks to be reorganized in the second file;

[0027] Determine the entity to be updated in the first file according to the entity identifier recorded in the element object corresponding to the storage block that does not need to be reorganized;

[0028] Read the content of the entity to be updated from the first file, and update the index information of the non-fixed-length fields in the content of the entity to be updated according to the field identifier and the initial position recorded in the element object corresponding to the storage block that does not need to be reorganized.

[0029] In one implementation manner of the embodiment of the present application, after converting the structured data into multiple entities described in the first virtual memory space according to the defined data structure, it further includes:

[0030] Take each entity as a node, and implement the sorting of multiple entities in the first virtual memory space based on the large top heap algorithm or the small top heap algorithm; wherein, the value of the node is the value of the specified field of the corresponding entity.

[0031] In a second aspect of the embodiments of the present application, a structured data storage device is provided, including:

[0032] A data acquisition module, configured to acquire structured data to be stored;

[0033] A data conversion module, configured to convert the structured data into multiple entities described in a first virtual memory space according to a defined data structure; wherein the data structure includes fixed-length fields and non-fixed-length fields, and the content of each entity includes the data of the fixed-length fields and the index information of the non-fixed-length fields. The index information is the offset of the storage position of the data of the corresponding non-fixed-length field relative to the initial address of the second virtual memory space; the first virtual memory space is a virtual space corresponding to a first file in the disk constructed by memory mapping; the second virtual memory space is a virtual space corresponding to a second file in the disk constructed by memory mapping;

[0034] A data writing module, configured to write the content of the multiple entities into the first file, and write the data of the non-fixed-length fields of the multiple entities into the second file.

[0035] In a third aspect of the embodiments of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the structured data storage method provided in the first aspect of the embodiments of the present application is implemented.

[0036] In a fourth aspect of the embodiments of the present application, a computer program product is provided. When the computer program product runs on an electronic device, the electronic device is enabled to execute the structured data storage method provided in the first aspect of the embodiments of the present application.

[0037] In a fifth aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the structured data storage method provided in the first aspect of the embodiments of the present application is implemented.

[0038] It can be understood that the beneficial effects of the above second aspect to fifth aspect can refer to the relevant descriptions in the above first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is a flowchart of a structured data storage method provided by an embodiment of the present application;

[0040] Figure 2 is a schematic structural diagram of a structured data storage device provided by an embodiment of the present application;

[0041] Figure 3 It is a schematic diagram of an electronic device provided by an embodiment of the present application. Specific embodiments

[0042] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from hindering the description of the present application. Additionally, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0043] Currently, structured data generated by business systems such as banks and government affairs are generally stored using various relational databases such as MySQL. The update and query statements of the business system must go through internal control processes such as SQL parsing and transaction management in the database system execution framework, which will cause a certain execution delay. Moreover, when structured data enters the engine management space of the database system, it needs to be converted into the internal data structure of the engine and serialized before completing persistence, which will result in a relatively high time delay and additional system performance overhead.

[0044] In view of the above technical problems existing in storing structured data using relational databases, the embodiments of the present application propose a structured data storage method, an electronic device, and a computer program product. By means of the memory mapping technology based on the operating system, the effect of directly mapping memory updates to disk storage is achieved, which can bypass the cumbersome mechanism of the database system and reduce the time delay of data persistence and system performance overhead. For more specific technical implementation details of the embodiments of the present application, please refer to the various method embodiments described below.

[0045] It should be understood that the execution subject of each method embodiment proposed by the present application can be various types of electronic devices. For example, it can be a mobile phone, a tablet computer, a wearable device, a desktop computer, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a large-screen TV, etc. The embodiments of the present application do not impose any restrictions on the specific type of the electronic device.

[0046] Please refer toFigure 1 , which shows a structured data storage method provided by an embodiment of the present application, including:

[0047] 101. Obtain the structured data to be stored;

[0048] First, obtain the structured data to be stored, and this structured data needs to be persisted to the disk. Structured data can also be called row data, which is data logically expressed and implemented by a two-dimensional table structure, and strictly follows the data format and length specifications. The embodiments of the present application do not impose any restrictions on the attributes of the structured data to be stored, such as type, size, number of fields, and specific fields.

[0049] 102. Convert the structured data into multiple entities described in the first virtual memory space according to the defined data structure;

[0050] Since the structured data strictly follows the data format and length specifications, the data can be described in memory by defining a suitable data structure, that is, obtaining the memory description of the structured data, and these memory descriptions can be called entities to be persisted. For example, assuming that the structured data is a two-dimensional table with M rows and N columns, then M entities described in memory can be obtained, and each entity has N fields.

[0051] The technical solution of the embodiments of the present application pre-creates two files on the disk, denoted as the first file and the second file respectively, and maps these two files into memory through the memory mapping method (i.e., mmap). In this way, the first file will obtain a continuous virtual memory space denoted as the first virtual memory space, and the second file will also obtain a continuous virtual memory space denoted as the second virtual memory space. Among them, the first file is used to store all data in the entity except for the data of non-fixed-length fields, and the second file is used to store the data of non-fixed-length fields in the entity.

[0052] After obtaining the structured data to be stored, convert the structured data into multiple entities described in the first virtual memory space according to the defined data structure. Among them, the data structure includes fixed-length fields and non-fixed-length fields, and the content of each entity includes the data of the fixed-length fields and the index information of the non-fixed-length fields. The index information is the offset of the storage position of the data of the corresponding non-fixed-length field relative to the initial address of the second virtual memory space.

[0053] As an example, for the Go language, the fixed-length fields in the entity can be represented by basic types (such as int32, int64, float64, [N]byte, etc.), and specific data instances can include: statistical data collected regularly, price information, etc. For such data, the following entity data structure can be defined in memory:

[0054]

[0055] However, for non-fixed-length fields in an entity, such as some string-type data, it cannot be directly represented as fixed-length content. These data can be defined in memory as the following entity data structure:

[0056]

[0057] The non-fixed-length field field6 consists of 8 bytes, where the first 4 bytes represent the actual length of the data, and the last 4 bytes represent the logical offset of the data in the actual storage file, that is, the index information of the above non-fixed-length field. Using this logical offset, the storage location of the data of this non-fixed-length field in the second file can be located.

[0058] For the technical solution of the embodiment of the present application, if most or all of the data in the structured data to be stored are fixed-length field data, that is, few or no data are non-fixed-length field data, better data storage performance and efficiency can be obtained. Also taking the Go language as an example, the entity in the memory representation can directly obtain the memory size occupied by each entity through functions built into the language such as sizeof. Due to the characteristics of the Go language, the structure in memory can be directly cast to a fixed-size byte-type data, and the memory space size occupied in this memory can be directly mapped to the disk space size when stored on the disk, thereby saving the complex process of entity serialization and deserialization and effectively reducing the latency of data persistence and the system performance overhead.

[0059] In addition, considering that there are multiple entities in the structured data, and the fixed-length field data and non-fixed-length field data in the entity are stored in different files, in order to distinguish the entities so that each non-fixed-length field data can find the corresponding entity, a unique entity identifier (that is, entity ID) can be assigned to each entity in the memory representation. This entity identifier can be of types such as int64 or a fixed-length byte array, etc., and is used to uniquely represent the corresponding entity.

[0060] 103. Write the content of multiple entities into the first file, and write the data of the non-fixed-length fields of multiple entities into the second file.

[0061] After converting structured data into multiple entities described in the first virtual memory space, the content of the multiple entities, that is, the content such as fixed-length fields, data of fixed-length fields, non-fixed-length fields, and index information of non-fixed-length fields, etc., is written into the first file on the disk, and the data of the non-fixed-length fields of the multiple entities is written into the second file on the disk, thus completing the persistent operation of the structured data. In this process, the operating system can automatically map the virtual address in the first virtual memory space to the physical address of the first file on the disk, so as to write the entity content into the first file; similarly, the operating system can automatically map the virtual address in the second virtual memory space to the physical address of the second file on the disk, so as to write the data of the fixed-length fields into the second file. Moreover, from the implementation perspective, since the execution process of memory mapping only creates virtual memory identifiers and maps them to file descriptors in the system kernel, the return latency of the system call interface will not be affected by the preset parameter "file mapping capacity size".

[0062] Specifically, when writing the content of multiple entities into the first file, each entity in the memory can directly ignore the type conversion and be transformed into a continuous multi-byte memory space, and this memory space can be directly persisted into the first file. Based on this operation process, when the entity content is stored in the first file, it has the following two characteristics:

[0063] (1) The storage space occupied by the content of each entity must be continuous and fixed-length;

[0064] (2) The content of each entity can be obtained by adding an offset to the starting position of the virtual address obtained after performing memory mapping on the first file.

[0065] Regarding the above characteristic (1), it can be seen from the entity data structure defined above that the content of each entity includes fixed-length fields, data of fixed-length fields, non-fixed-length fields, and index information of non-fixed-length fields. Among them, the fields of different entities are the same, the data lengths of fixed-length fields are the same, and the index information is also of a fixed length (for example, 8 bytes). Therefore, the storage space occupied by the content of each entity must be of a fixed length. In addition, since the first virtual memory space includes continuous virtual addresses, based on the memory mapping mechanism, the content of each entity will be written into consecutive physical addresses in the first file, so the storage space occupied by the content of each entity must be continuous.

[0066] For the above feature (2), assuming that the memory space occupied by the content of each entity is M, the corresponding offset is M, and the starting position of the virtual address obtained after the first file performs memory mapping is addr. Then, in the first virtual memory space: the data within the range of addr to addr + M is the content of the first entity, the data within the range of addr + M to addr + 2M is the content of the second entity... and so on. The data within the range of addr + (N - 1)M to addr + NM is the content of the Nth entity.

[0067] The data storage process of the first file is relatively simple, while that of the second file is relatively complex. The second file is used to store the data of non-fixed-length fields in the entity. Since the data of each non-fixed-length field belongs to a certain entity stored in the first file, the data stored in the second file needs to be mapped and linked with the entities stored in the first file. When operations such as insertion, update, or deletion occur to the entity data stored in the first file, the lengths of some of the data stored in the second file are likely to change. For example, when the upper-layer business logic changes or an entity is deleted, it may cause the lengths of the data of some non-fixed-length fields stored in the second file to change. The following introduces a data management method for the second file. Since the data written into the second file may change in length, as the system runs over time, the second file may have file fragments of different sizes. For ease of management, the second file can be divided into blocks, so that the second file will include multiple storage blocks, and the sizes of each storage block can be the same or different. To conform to the management habits of the operating system and reduce the operation difficulty, the sizes of each storage block can be made the same, and each storage block can be 4KB or a multiple of 4KB. After dividing the second file into multiple storage blocks, a bitmap file (i.e., bitmap) can be introduced to manage each storage block of the second file. Each bit of the bitmap file is used to represent the occupancy of each storage block of the second file. For example, when a certain bit is 0, it means that the corresponding storage block can still write data, and when a certain bit is 1, it means that the corresponding storage block is full and cannot write new data. The bitmap file can maintain an identifier of the currently writing block. Before the storage block corresponding to this block identifier is full, all the written data will be stored in this storage block. This block identifier needs to be stored separately. A feasible way is to use 8 bytes at the beginning of the bitmap file to record the number of used storage blocks in the second file, and this number is equal to the number of bits of the bitmap file. For example, assuming the number of bits is 100, it means that 100 storage blocks in the second file have been used. At this time, write the value 100 into the 8 bytes at the beginning of the bitmap file.

[0068] In an implementation of an embodiment of the present application, the second file includes a plurality of storage blocks; writing data of non-fixed length fields of a plurality of entities into the second file includes:

[0069] (1) Obtain the bitmap file corresponding to the second file; wherein, each bit of the bitmap file is respectively used to represent the occupancy situation of each storage block of the second file;

[0070] (2) According to the bitmap file, determine the target storage blocks that can write data from the plurality of storage blocks, and write the data of the non-fixed length fields of the plurality of entities into the target storage blocks.

[0071] When writing the data of the non-fixed length fields of a plurality of entities into the second file, first obtain the bitmap file corresponding to the second file. According to the specific values of each bit in the bitmap file, the occupancy situation of each storage block included in the second file can be determined, so as to find one or more storage blocks that can write data, denoted as target storage blocks, and then write the data of the non-fixed length fields of the plurality of entities into the target storage blocks. Here, the data can be written into each storage block in sequence. After the previous storage block is filled with data, continue to write data to the next storage block until all the data of the non-fixed length fields of the entities are written into the second file.

[0072] In an implementation of an embodiment of the present application, obtaining the bitmap file corresponding to the second file includes:

[0073] Obtain the bitmap file from the disk according to the virtual address in the third virtual memory space; wherein, the third virtual memory space is constructed by means of memory mapping, and the virtual space corresponding to the bitmap file in the memory.

[0074] The bitmap file is created on the disk. To facilitate reading the bitmap file, the bitmap file can be mapped into the memory in advance by means of memory mapping. In this way, the bitmap file will obtain a continuous virtual memory space denoted as the third virtual memory space. When the system starts each time, the operating system can automatically map the virtual address in the third virtual memory space to the physical address of the bitmap file on the disk, so as to obtain the content of the bitmap file. This process does not need to traverse the second file and is very convenient to operate.

[0075] Referring to the previous description, the data written in the second file may change in length. As the system runs over time, different sizes of data fragments, that is, invalid data, may appear in each storage block included in the second file. To save storage space, it is necessary to regularly perform data sorting operations on each storage block of the second file to clear the invalid data and release the storage space.

[0076] To facilitate recording the invalid data contained in each storage block, an array of invalid characters can be created and maintained in memory. The length of the array of invalid characters is the same as the number of bits in the bitmap file, that is, the number of element objects in the array of invalid characters is the same as the number of storage blocks in the second file. For example, assume that the bitmap file has 100 bits, corresponding to 100 storage blocks in the second file. Then the array of invalid characters also has 100 element objects, and these 100 element objects, 100 storage blocks, and 100 bits are all in one-to-one correspondence.

[0077] Each element object can record the number of invalid characters of all the data written in the corresponding storage block, as well as record the entity identifier and field identifier corresponding to each non-fixed-length field data written in the corresponding storage block.

[0078] As an example, the data structure definition of the element object InvalidMark is as follows:

[0079] type InvalidMark struct{

[0080] invalidLen int16 / / The number of invalid characters of all the data written in the corresponding storage block.

[0081] keys[]*FieldMark

[0082] }

[0083] type FieldMark struct{

[0084] keys[]int64 / / Represents the entity identifier corresponding to the data

[0085] fieldIdx[]int / / Represents the field identifier corresponding to the data, used to mark which non-fixed-length field in the entity the data belongs to.

[0086] }

[0087] Considering that an entity may have multiple non-fixed-length fields, the data of each non-fixed-length field needs to be uniquely determined by "entity identifier" + "field identifier". The entity identifier is used to locate the entity to which the data belongs, and the field identifier is used to locate which non-fixed-length field in the corresponding entity the data is. invalidLen represents the number of invalid characters of all the data written in the corresponding storage block, and can be used to characterize the proportion of invalid data in the corresponding storage block. Therefore, when the number of invalid characters invalidLen is too large, it can be considered that the corresponding storage block needs to perform data reorganization operations.

[0088] In one implementation manner of the embodiment of the present application, the method further includes:

[0089] (1) When any storage block of the second file is written with data, a first byte character is added in front of the written data;

[0090] (2) When the length of any written data in the any storage block changes, the first byte character in front of the any data is modified to a second byte character, and the number of invalid characters recorded in the element object corresponding to the any storage block is updated.

[0091] When any storage block of the second file is written with data, a first byte character is added in front of the written data. The first byte character is used to indicate that the data is valid and to quickly split the data content in the second file; conversely, when the length of any written data in the any storage block changes, it means that the any data may be deleted or replaced by other new values, so the any data becomes invalid data. At this time, the first byte character in front of the any data is modified to a second byte character, and the second byte character is used to indicate that the data is invalid. As an example, the first byte character can be a null byte <null>, its binary representation is 00000000; the second byte character can be <1>, and its binary representation is 00000001. Additionally, since new invalid data appears in the arbitrary storage block, the number of invalid characters invalidLen recorded in the element object corresponding to the arbitrary storage block also needs to be updated accordingly. Assuming the length of the arbitrary data is 7 bytes, it is necessary to increase the number of invalid characters invalidLen recorded in the element object corresponding to the arbitrary storage block by 7 bytes, and so on.

[0092] The system background can run a timed data cleaning task to perform data cleaning operations on each storage block of the second file, that is, to periodically clean up the fragments of the second file.

[0093] In one implementation manner of the embodiment of the present application, the method further includes:

[0094] (1) Determine the storage blocks to be cleaned from each storage block of the second file;

[0095] (2) Determine the target element object corresponding to the storage block to be cleaned from each element object included in the invalid character array; wherein, each element object corresponds to each storage block of the second file one by one, and each element object records the number of invalid characters of all the data written in the corresponding storage block;

[0096] (3) If the number of invalid characters recorded in the target element object exceeds a preset threshold, perform data cleaning operations on the storage block to be cleaned.

[0097] When the data sorting task runs, first determine the storage blocks to be sorted from each storage block of the second file. Here, a storage block can be randomly selected or selected in sequence as the storage block to be sorted. In actual operation, a bit can be randomly selected from the bitmap file, and the storage block corresponding to this bit is used as the storage block to be sorted. After selecting the storage block to be sorted, search for the element object corresponding to the storage block to be sorted from each element object included in the invalid character array, and denote it as the target element object. Obtain the number of invalid characters invalidLen recorded by the target element object. If the number of invalid characters invalidLen exceeds a certain preset threshold (for example, 50% of the total number of characters in the storage block), it means that there is too much invalid data in the storage block to be sorted, so a data sorting operation is performed on the storage block to be sorted; otherwise, if the number of invalid characters invalidLen does not exceed the preset threshold, it means that there is less invalid data in the storage block to be sorted and there is no need to sort it temporarily. At this time, the next storage block can be continuously obtained as the new storage block to be sorted, and the same processing process is repeated. In addition, in order to avoid excessive data volume involved in a single data sorting task resulting in long-term occupation of storage blocks in the background, an upper limit on the number of storage blocks sorted in a single data sorting task can be set. Assuming it is 20, when 20 storage blocks are selected as the storage blocks to be sorted and the sorting is completed, the current data sorting task ends, and wait for the next data sorting task to be executed.

[0098] In an implementation manner of the embodiment of the present application, each element object also records the entity identifier and field identifier corresponding to each non-fixed-length field data written in the corresponding storage block; performing a data sorting operation on the storage block to be sorted includes:

[0099] (1) Determine the data with the first byte character in front in the storage block to be sorted as valid data, and write the valid data into a new storage block;

[0100] (2) Update the element object corresponding to the new storage block according to the entity identifier and field identifier corresponding to the valid data.

[0101] When performing data reorganization operations on the storage block to be reorganized, traverse the data of each variable-length field in the storage block to be reorganized. If the first byte of a certain data is, for example, 00000000, then this data is recorded as valid data and needs to be retained. It can be written into a new storage block, and the corresponding element object of this new storage block also needs to be updated accordingly. For example, the entity identifier and field identifier corresponding to the valid data need to be added. If the first byte of a certain data is, for example, 00000001, then this data is recorded as invalid data and is deleted or ignored. The invalid data will not be persisted to the new storage block. When the reorganization of all data in the storage block to be reorganized is completed and all valid data in the storage block to be reorganized is written into the new storage block, the original data in the storage block to be reorganized can be cleared.

[0102] As an example, assume that traversal starts from the 100th bit of a bitmap file. When traversing to the 101st bit, it is found that the number of invalid characters recorded in the element object corresponding to storage block A is 3000 (assuming the size of the storage block is 4KB). Since the threshold of more than 50% of the character number is not exceeded, storage block A needs to perform the above data reorganization operation. Assume that the specific data content of storage block A is: <null>hello <null>world<1>greeting <null>palyground<1>prefect<1>standa rd<1>happy birthday<1>create <null>books, when sorting out data <null>Retain the data starting with <null>hello <null>world <null>palyground <null>books, the sorted data content is written into a new storage block B.

[0103] The content of the element object InvalidMark A in storage block A is as follows:

[0104] {

[0105] invalidLen: 48(9 + 8 + 9 + 15 + 7),

[0106] keys: [1, 2, 6, 9, 10, 14, 17, 39, 41], / / The key here is a randomly set string representing data <null>The entity identifier corresponding to "hello" is 1, data <null>The entity identifier corresponding to "world" is 2... and so on.

[0107] }

[0108] The content of the element object InvalidMark B of the new storage block B is as follows:

[0109] {

[0110] invalidLen: 0,

[0111] keys: [1, 2, 9, 41],

[0112] }

[0113] Generally speaking, since the proportion of invalid data in the storage block to be sorted is relatively high, the number of valid data written into the new storage block is small, so there will be a lot of remaining space in the new storage block. To avoid wasting storage space, after finishing sorting a storage block and writing the valid data into the new storage block, the new storage block is not persisted temporarily, but continues to traverse to find the next storage block to be sorted, so that the valid data of the subsequent storage blocks can be concatenated after the valid data of the previous storage block until the data in the new storage block is full or all the data sorting operations of all storage blocks are completed.

[0114] The above invalid character array is a data structure existing in memory. Assuming that there are non-fixed-length fields in the entity, the above invalid character array needs to be rebuilt after the system crashes and restarts. This process involves traversing all entities and constructing the element objects corresponding to all storage blocks. In actual operation, the construction operation of the invalid character array can be run in the background, so that the system can provide services normally during the construction of the invalid character array. It should be noted that to avoid data logic errors, the data sorting operation of the storage block needs to be paused during this process until the construction of the invalid character array is completed and then the data sorting operation of the storage block is resumed. Moreover, when the second file needs to write data into the new storage block, if the corresponding old storage block has been traversed during the construction of the invalid character array (that is, the element object corresponding to the old storage block has been correctly constructed), the element object corresponding to the old storage block can be directly updated; if the corresponding old storage block has not been traversed during the construction of the invalid character array, the data to be updated of the element object corresponding to the old storage block is temporarily cached until the element object corresponding to the old storage block is correctly constructed, and then the cached data to be updated is written into the element object corresponding to the old storage block.

[0115] On the other hand, since the size of each storage block is fixed, for example, 4KB, it is very likely that a certain piece of data is split and written into two storage blocks. In response to this situation, specified byte characters can be added before and after the cross-block data as cross-block identifiers, which can facilitate the identification and restoration of cross-block data. For example, assuming that the cross-block identifier is <2>, there are still 3 bytes of space left at the end of storage block A, but the data to be stored is "empty" which occupies 5 bytes. At this time, "empty" can be split and written into the end of storage block A and the beginning of the next storage block B respectively. Specifically, the end of storage block A is <null>e<2>, the beginning of storage block B is <2>mpty. In particular, if only 2 bytes remain at the end of storage block A, since 2 bytes are only enough to write <null>With respect to <2>, since the actual valid byte e cannot be written, these two bytes can be regarded as invalid bytes. At this time, empty can only be written to the next storage block B.

[0116] Although the mechanism of cross-block data storage can maximize the utilization of storage space, such an operation will greatly increase the complexity of data management. In another processing method, the storage space at the end of the storage block that is not sufficient to accommodate the complete data can be marked as invalid bytes (for example, all are set to <1>) to simplify the data management operation.

[0117] After completing the data sorting operation of the storage block to be sorted, in order to prevent the content of the second file from being quickly exhausted, other storage blocks that do not need to be sorted can be moved to the position of the storage block to be sorted. Such a design is called "block movement". The following describes the specific operation method of "block movement".

[0118] In one implementation manner of the embodiment of the present application, after performing the data sorting operation on the storage block to be sorted, it further includes:

[0119] (1) Determine the storage blocks that do not need to be sorted from each storage block of the second file;

[0120] (2) Move the storage blocks that do not need to be sorted to the initial position of the storage block to be sorted in the second file;

[0121] (3) Determine the entity to be updated in the first file according to the entity identifier recorded in the element object corresponding to the storage block that does not need to be sorted;

[0122] (4) Read the content of the entity to be updated from the first file, and update the index information of the non-fixed-length fields in the content of the entity to be updated according to the field identifier and the initial position recorded in the element object corresponding to the storage block that does not need to be sorted.

[0123] After completing the data sorting operation of the storage block to be sorted, start the "block movement" operation. First, determine a storage block that does not need to be sorted from each storage block of the second file, denoted as the storage block that does not need to be sorted, which is usually the storage block at the end of the second file. Then, move the storage block that does not need to be sorted to the initial position of the storage block to be sorted in the second file. Such an operation will cause the storage positions of all valid data in the storage block that does not need to be sorted to be moved. Therefore, the corresponding entity content recorded in the first file also needs to be updated. Specifically, according to the entity identifier recorded in the element object corresponding to the storage block that does not need to be sorted, the entity to be updated in the first file can be determined. Here, the skip list technology can be used to index the storage position of the corresponding entity in the first file based on the entity identifier; then, read the content of the entity to be updated from the first file, and update the index information of the non-fixed-length fields in the content of the entity to be updated according to the field identifier and the initial position recorded in the element object corresponding to the storage block that does not need to be sorted.

[0124] For example, assume that an unorganized storage block is moved from the position of 32 * 1024 to the position of 4 * 1024, and the Keys field of the corresponding element object InvalidMark is {{101, 3} and {120, 3}}. First, use the skip list technology to index the offset of the corresponding entity in the first file based on the entity identifier 101, assume it is 1000; then, read the data in the range of 1000 to 1000 + M (entity size) in the first file, convert it into the corresponding entity content in memory, and modify the last 4 bytes (i.e., the index information) of the third field in the entity content to 4 * 1024; similarly, modify the last 4 bytes (i.e., the index information) of the third field in the entity content of entity 120 to 4 * 1024 as well.

[0125] It should be noted that when using the skip list technology, since the skip list is also a data structure existing in memory, similar to the above-mentioned invalid character array, when the system crashes and restarts after data has been written to the first file and the second file, the skip list needs to be rebuilt. The reconstruction of the skip list also involves traversing all entities. In actual operation, the reconstruction process of the skip list can be combined with the reconstruction process of the above-mentioned invalid character array and run in the background, that is, after the system crashes and restarts, the invalid character array and the skip list are rebuilt together in the background. During the reconstruction of the skip list, if the required index content cannot be found using the skip list, then it can only be queried by traversing the first file. In some scenarios, the structured data generated by the business system also has a clear basis for sorting, and it is necessary to sort each data according to the value of the specified field, such as sorting by cost from low to high, sorting by profit from high to low, and so on. For such scenarios, the technical solution of the embodiment of the present application can introduce the large top heap algorithm or the small top heap algorithm to sort each entity in the memory representation to meet business requirements such as obtaining the top-N data instances. The following describes the specific operation method.

[0126] In one implementation manner of the embodiment of the present application, after converting the structured data into multiple entities described in the first virtual memory space according to the defined data structure, it further includes:

[0127] Regarding each entity as a node, implementing the sorting of multiple entities in the first virtual memory space based on the large top heap algorithm or the small top heap algorithm; where the value of the node is the value of the specified field of the corresponding entity.

[0128] After introducing the max-heap algorithm or min-heap algorithm, since the data structures of max-heap or min-heap inherently have sorting capabilities, it is convenient to achieve efficient real-time sorting of each entity and solve the problem of the top-N entities. Specifically, the memory management scheme of max-heap or min-heap can use a byte array. Since the virtual addresses of byte arrays in memory must be continuous, the pointer of this byte array can directly use the virtual memory address obtained after memory mapping of the first file. Taking each entity as a node, the value of the node is the value of the specified field of the corresponding entity (such as cost value or profit value, etc.). Sort multiple entities based on the max-heap algorithm or min-heap algorithm in the first virtual memory space, so that the sorted entity data is also mapped to the first file.

[0129] Since one node in the max-heap or min-heap stores one entity, and the memory size occupied by each entity data is exactly the same, the size of each actual node in the max-heap or min-heap is determined. Taking the max-heap as an example, according to the shape of the max-heap data structure, it can be determined that if the node index of the current node in the heap is x, then the node indexes of the left and right child nodes of this node are 2x and 2x + 1 respectively. Assuming that the size of each node is 64 bytes, the starting position of each entity in memory is the node index * 64, and this starting position in memory is also the offset of the first file. Therefore, data changes in the max-heap or min-heap can be directly reflected in the disk file through memory mapping without performing additional serialization operations and persistence operations.

[0130] In addition, as the system continues to run, the number of nodes in the max-heap or min-heap will continue to increase, and the entity content in the first file will also continue to increase. At this time, the first file needs to have the function of determining the total number of current nodes. The first file can reserve the first 8 bytes to store the total number of nodes N in the max-heap or min-heap, and N * M (entity size) is the current legal content length of the first file.

[0131] When the max-heap algorithm or min-heap algorithm is running, it usually performs operations such as "sifting up" and "sifting down", which will cause the situation of node position swapping. Taking the max-heap as an example, when any node in the max-heap is swapped, if it is the node at the beginning or end, the starting memory offset of the node is 0 or the current legal content length N * M of the first file; if it is the node in the middle, the skip list data structure can be used to quickly index the position offset of the node in the first file according to the entity identifier, that is, the offset relative to the starting position of the max-heap memory.

[0132] When a node needs to update the data of a non-fixed-length field, it first obtains the index information and data length of the data of the non-fixed-length field in the second file through the memory structure. According to the index information and data length, the corresponding storage block A can be found in the second file, and the element object InvalidMark A corresponding to the storage block A is read; then, according to the bitmap file, the storage block B that can write data in the second file is determined, and the new data is written into the storage block B. If the storage space of the storage block B is not enough, the next storage block C is continuously obtained until all the new data is written. Empty bytes are added in front of the new data <null>; Next, find the element object InvalidMark A corresponding to storage block A, and regard all the space occupied by the old data as invalid bytes, that is, the empty bytes in front of the old data <null>Replace it with the byte character <1>, and then adjust the information such as the number of invalid characters, entity identifier, and field identifier recorded in the element object InvalidMark A accordingly.

[0133] The technical solution of the embodiment of the present application creates a first file and a second file on the disk, and respectively constructs a first virtual memory space corresponding to the first file in the memory and a second virtual memory space corresponding to the second file in the memory through the memory mapping method; after obtaining the structured data to be stored, according to the defined data structure, the structured data is converted into multiple entities described in the first virtual memory space. The data structure includes fixed-length fields and non-fixed-length fields. The content of each entity includes the data of the fixed-length field and the index information of the non-fixed-length field. The index information is the offset of the storage position of the data of the corresponding non-fixed-length field relative to the initial address of the second virtual memory space; then, the content of multiple entities is written into the first file, and the data of the non-fixed-length fields of multiple entities is written into the second file. The above process obtains the virtual memory space corresponding to the disk file through the memory mapping method, and the structured data stored in the disk file can be directly mapped and obtained by using the virtual address of the virtual memory space, which can break away from the limitations of the internal mechanism of the database system and reduce the time delay of data persistence and the system performance overhead to a certain extent.

[0134] In summary, the structured data storage method proposed in the embodiment of the present application does not depend on the database and can avoid various limitations of the internal mechanism of the database system; by adopting the memory mapping method based on the operating system, the effect of directly mapping the memory update to the disk storage is realized, and the delay of data persistence can be reduced to the greatest extent; moreover, the embodiment of the present application can also introduce data structures such as a max heap or a min heap, and sort each entity in the memory representation in combination with an algorithm to meet the requirements of various service scenarios such as obtaining the top-N data instances, and has a broad market application prospect.

[0135] It should be understood that the magnitudes of the sequence numbers of the steps in the above various embodiments do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiment of the present application.

[0136] The above mainly describes a structured data storage method. Next, a structured data storage device will be described.

[0137] Please refer to Figure 2 , which shows a structured data storage device provided by an embodiment of the present application, including:

[0138] A data acquisition module 201, configured to acquire the structured data to be stored;

[0139] A data conversion module 202 is configured to convert structured data into multiple entities described in a first virtual memory space according to a defined data structure. The data structure includes fixed-length fields and non-fixed-length fields. The content of each entity includes the data of the fixed-length fields and the index information of the non-fixed-length fields. The index information is the offset of the storage location of the data of the corresponding non-fixed-length field relative to the initial address of the second virtual memory space. The first virtual memory space is a virtual space corresponding to a first file in the disk constructed by memory mapping. The second virtual memory space is a virtual space corresponding to a second file in the disk constructed by memory mapping.

[0140] A data writing module 203 is configured to write the content of multiple entities into the first file and write the data of the non-fixed-length fields of multiple entities into the second file.

[0141] In an implementation manner of the embodiment of the present application, the second file includes multiple storage blocks. The data writing module includes:

[0142] A bitmap file acquisition unit is configured to acquire a bitmap file corresponding to the second file. Each bit of the bitmap file is respectively used to represent the occupancy of each storage block of the second file.

[0143] A data writing unit is configured to determine a target storage block for writing data from multiple storage blocks according to the bitmap file, and write the data of the non-fixed-length fields of multiple entities into the target storage block.

[0144] In an implementation manner of the embodiment of the present application, the bitmap file acquisition unit includes:

[0145] A file acquisition subunit is configured to acquire a bitmap file from the disk according to a virtual address in a third virtual memory space. The third virtual memory space is a virtual space corresponding to the bitmap file in the disk constructed by memory mapping.

[0146] In an implementation manner of the embodiment of the present application, the structured data storage device further includes:

[0147] A storage block to be sorted out determination module is configured to determine a storage block to be sorted out from each storage block of the second file.

[0148] An element object determination module is configured to determine a target element object corresponding to the storage block to be sorted out from each element object included in the invalid character array. Each element object corresponds to each storage block of the second file, and each element object records the number of invalid characters of all data written in the corresponding storage block.

[0149] A data sorting module, which is used to perform data sorting operations on the storage block to be sorted if the number of invalid characters recorded in the target element object exceeds a preset threshold.

[0150] In an implementation manner of the embodiment of the present application, the structured data storage device further includes:

[0151] A byte character adding module, which is used to add a first byte character in front of the written data when any storage block of the second file is written with data;

[0152] A byte character modifying module, which is used to modify the first byte character in front of the arbitrary data to a second byte character when the length of any written data in the arbitrary storage block changes, and update the number of invalid characters recorded in the element object corresponding to the arbitrary storage block.

[0153] In an implementation manner of the embodiment of the present application, each element object also records the respective entity identifiers and field identifiers corresponding to the data of each non-fixed length field written in the corresponding storage block; the data sorting module includes:

[0154] A data transfer and storage unit, which is used to determine the data with the first byte character in front in the storage block to be sorted as valid data, and write the valid data into a new storage block;

[0155] An element object updating unit, which is used to update the element object corresponding to the new storage block according to the entity identifier and field identifier corresponding to the valid data.

[0156] In an implementation manner of the embodiment of the present application, the structured data storage device further includes:

[0157] A storage block without sorting determination module, which is used to determine the storage blocks without sorting from each storage block of the second file;

[0158] A storage block moving module, which is used to move the storage blocks without sorting to the initial position of the storage blocks to be sorted in the second file;

[0159] A to-be-updated entity determination module, which is used to determine the to-be-updated entity in the first file according to the entity identifier recorded in the element object corresponding to the storage block without sorting;

[0160] An entity content updating module, which is used to read the content of the to-be-updated entity from the first file, and update the index information of the non-fixed length fields in the content of the to-be-updated entity according to the field identifier and the initial position recorded in the element object corresponding to the storage block without sorting.

[0161] In an implementation manner of the embodiment of the present application, the structured data storage device further includes:

[0162] An entity sorting module, which is used to take each entity as a node and implement the sorting of multiple entities in the first virtual memory space based on the max-heap algorithm or the min-heap algorithm; wherein, the value of the node is the value of the specified field of the corresponding entity.

[0163] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the structured data storage method described in any of the above embodiments.

[0164] An embodiment of the present application further provides a computer program product. When the computer program product runs on an electronic device, it causes the electronic device to execute the structured data storage method described in any of the above embodiments.

[0165] Figure 3 is a schematic diagram of an electronic device provided by an embodiment of the present application. As Figure 3 shown, the electronic device 3 of this embodiment includes: a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30. When the processor 30 executes the computer program 32, it implements the steps in the embodiments of the above various structured data storage methods, such as Figure 1 the steps 101 - step 103 shown. Alternatively, when the processor 30 executes the computer program 32, it implements the functions of each module / unit in the above device embodiments, such as implementing Figure 2 the functions of module 201 - module 203 of the device shown.

[0166] The computer program 32 can be divided into one or more modules / units. The one or more modules / units are stored in the memory 31 and executed by the processor 30 to complete the present application. The one or more modules / units can be a series of computer program instruction segments capable of completing specific functions, and these instruction segments are used to describe the execution process of the computer program 32 in the electronic device 3.

[0167] The so-called processor 30 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0168] The memory 31 may be an internal storage unit of the electronic device 3, such as the hard disk or memory of the electronic device 3. The memory 31 may also be an external storage device of the electronic device 3, such as a plug-in hard disk equipped on the electronic device 3, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 31 may also include both the internal storage unit of the electronic device 3 and the external storage device. The memory 31 is used to store the computer program and other programs and data required by the electronic device. The memory 31 may also be used to temporarily store data that has been output or will be output.

[0169] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be assigned to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0170] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, the specific working processes of the system, device, and unit described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0171] In the above embodiments, the descriptions of the various embodiments each have their own emphasis. For parts not described in detail or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0172] Those of ordinary skill in the art will appreciate that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician may use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0173] In the embodiments provided in this application, it should be understood that the disclosed apparatus and method can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings, direct couplings, or communication connections shown or discussed with each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0174] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of this application.

[0175] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0176] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, it can also be completed by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0177] The above-described embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of this application, and should all be included within the protection scope of this application.< / null> < / null> < / null> < / null> < / null> < / null> < / null> < / null> < / null> < / null> < / null> < / null> < / null> < / null> < / null> < / null>

Claims

1. A structured data storage method, characterized in that: include: Get structured data to be stored; According to the defined data structure, the structured data is converted into a plurality of entities described in the first virtual memory space; wherein the data structure includes a fixed-length field and a non-fixed-length field, and the content of each entity includes data of the fixed-length field and index information of the non-fixed-length field, and the index information is the offset of the storage location of the data of the corresponding non-fixed-length field relative to the initial address of the second virtual memory space; the first virtual memory space is constructed by memory mapping, and is a virtual space corresponding to the first file in the disk in the memory; the second virtual memory space is constructed by memory mapping, and is a virtual space corresponding to the second file in the disk in the memory; The contents of the plurality of entities are written into the first file, and the data of the non-fixed length fields of the plurality of entities are written into the second file.

2. The method according to claim 1, characterized in that The second file includes a plurality of storage blocks; and writing the data of the non-fixed length fields of the plurality of entities into the second file includes: Obtaining a bitmap file corresponding to the second file; wherein each bit of the bitmap file is used to represent the occupancy of each storage block of the second file; According to the bitmap file, a target storage block into which data can be written is determined from the plurality of storage blocks, and the data of the non-fixed-length fields of the plurality of entities are written into the target storage block.

3. The method according to claim 2, characterized in that The obtaining the bitmap file corresponding to the second file includes: The bitmap file is obtained from the disk according to the virtual address in the third virtual memory space; wherein the third virtual memory space is constructed by memory mapping, and the bitmap file corresponds to the virtual space in the memory.

4. The method according to claim 2, characterized in that Also includes: Determining storage blocks to be sorted from each storage block of the second file; Determine a target element object corresponding to the storage block to be sorted from each element object included in the invalid character array; wherein each element object corresponds to each storage block of the second file one by one, and each element object records the number of invalid characters of all data written in the corresponding storage block; If the number of invalid characters recorded in the target element object exceeds a preset threshold, a data sorting operation is performed on the storage block to be sorted.

5. The method according to claim 4, characterized in that Also includes: When data is written into any storage block of the second file, a first byte symbol is added in front of the written data; When the length of any data written in the arbitrary storage block changes, the first byte character in front of the arbitrary data is modified to a second byte character, and the number of invalid characters recorded in the element object corresponding to the arbitrary storage block is updated.

6. The method according to claim 5, characterized in that Each of the element objects also records the entity identifier and field identifier corresponding to each of the data of the non-fixed length fields written in the corresponding storage block; performing the data sorting operation on the storage block to be sorted includes: Determine the data in the storage block to be sorted that is preceded by the first byte symbol as valid data, and write the valid data into a new storage block; The element object corresponding to the new storage block is updated according to the entity identifier and the field identifier corresponding to the valid data.

7. The method according to claim 6, characterized in that After performing the data sorting operation on the storage block to be sorted, the method further includes: Determine a storage block that does not need to be sorted from each storage block of the second file; Moving the storage block not requiring arrangement to the initial position of the storage block to be arranged in the second file; Determine the entity to be updated in the first file according to the entity identifier of the element object record corresponding to the non-arrangement storage block; The content of the entity to be updated is read from the first file, and the index information of the non-fixed length field in the content of the entity to be updated is updated according to the field identifier and the initial position of the element object record corresponding to the non-arrangement storage block.

8. The method according to any one of claims 1 to 7, characterized in that After converting the structured data into a plurality of entities described in the first virtual memory space according to the defined data structure, the method further includes: Each of the entities is taken as a node, and the multiple entities are sorted in the first virtual memory space based on a max-heap algorithm or a min-heap algorithm; wherein the value of the node is the value of a designated field of the corresponding entity.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the structured data storage method according to any one of claims 1 to 8 is implemented.

10. A computer program product, characterized in that When the computer program product is run on an electronic device, the electronic device is enabled to execute the structured data storage method according to any one of claims 1 to 8.