Data management method and device, equipment, storage medium and computer program

By storing the off-line storage address of large field data in the database, and directly linking the data main table and off-line data table, the time-consuming and rewinding problems of Chunk ID application in the prior art are solved, and data storage and management efficiency is improved.

CN119938664APending Publication Date: 2025-05-06HUAWEI TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311464440.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In high concurrency data storage scenarios, the existing technology relies on globally incremented Chunk ID for off-line storage of large-field data, which makes it a long time to apply for Chunk ID and supports up to 232 identifier serial numbers, which is prone to rewinding problems, which seriously affects the performance of data storage.

Method used

By storing the off-row storage address of large field data in the database, instead of Chunk ID, directly associate the main data table with the off-row data table, to achieve efficient management of the data stored in the off-row database.

Benefits of technology

In high concurrency data storage scenarios, data storage efficiency is improved, performance bottlenecks in Chunk ID application and verification are avoided, and data management is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938664A_ABST
    Figure CN119938664A_ABST
Patent Text Reader

Abstract

The invention discloses a data management method and device, equipment, a storage medium and a computer program, and belongs to the technical field of databases. The method comprises the following steps: receiving a data storage request of a first tuple, wherein the first tuple comprises data of a plurality of fields; if the first data meets the out-of-row storage condition, the first data is stored in an out-of-row data table, and the first data is data of any field in the multiple fields; acquiring a first storage address of the first data in the out-of-row data table; and storing the first storage address in the first tuple of the data main table. Namely, the data main table and the out-of-row data table are directly associated based on the out-of-row storage address of the data, in the data management process, for a certain field in the tuple, the data in the out-of-row data table can be directly managed based on the out-of-row storage address recorded in the data main table, and in a high-concurrency data management scene, the data management efficiency is greatly improved. And the data management efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of database technology, and in particular to a data management method, device, equipment, storage medium and computer program. Background Art

[0002] In a database, a tuple refers to a record or a row of data in the database's data master table. A tuple is divided into a tuple header and attribute data. The tuple header is used to store some metadata, such as visibility information, bitmaps, etc., and the attribute data includes data of multiple attributes. Each attribute can also be understood as a field. Among the multiple attributes, there may be an attribute whose data length exceeds a threshold. This attribute can be called a large field, and the data of this attribute can be called large field data, but large field data cannot be directly stored in the data master table.

[0003] In the related art, openGauss (an open source relational database) provides a TOAST mechanism to support the database to store large field data. When a user stores data exceeding 2KB bytes, the data can be used as large field data, and the TOAST mechanism will try to store the large field data out of the row, that is, store the large field data in the TOAST table. When storing large field data out of the row, first apply for a 32-bit unique identifier for the large field data as a chunk ID (Chunk ID), which is the index value of the index out-of-row storage position, and store the Chunk ID in the tuple to which the large field data belongs in the data master table; at the same time, record the Chunk ID in the TOAST index table, as well as the out-of-row storage address indicating the large field data (i.e., the physical location in the TOAST table). When accessing the large field data stored out of the row, first obtain the Chunk ID corresponding to the large field data from the data master table, then determine the out-of-row storage address corresponding to the Chunk ID from the TOAST index table, and then use the out-of-row storage address to read the large field data from the TOAST table.

[0004] However, the above-mentioned off-line storage method relies on globally incremented Chunk IDs. However, in scenarios with high concurrency of data storage requests, it takes a long time to apply for Chunk IDs of large field data stored off-line. Moreover, the TOAST mechanism supports a maximum of 2 32 The identifier serial number is used as the Chunk ID. When applying for a Chunk ID, 2 32 There is a possibility of rollback of the identifier sequence number, that is, the second 32After the identifier sequence number is rolled back, it is necessary to re-apply from the first identifier sequence number to use the applied identifier sequence number as the Chunk ID. Therefore, once the identifier sequence number rolls back, it will take a long time to apply for the Chunk ID, seriously affecting the data storage performance. Summary of the invention

[0005] The present application provides a data management method, device, equipment, storage medium and computer program, which can improve the efficiency of data management. The technical solution is as follows:

[0006] In a first aspect, a data management method is provided, the method comprising:

[0007] Receive a data storage request for a first tuple, where the first tuple includes data of multiple fields; if the first data meets an out-of-row storage condition, store the first data in an out-of-row data table, where the first data is data of any field of the multiple fields; obtain a first storage address of the first data in the out-of-row data table; and store the first storage address in the first tuple of the data master table.

[0008] Optionally, storing the first data in an out-of-row data table includes: slicing the first data to obtain at least one slice data; and storing the at least one slice data in the out-of-row data table.

[0009] Optionally, the at least one slice data includes first slice data and second slice data connected in sequence; storing the at least one slice data in the out-of-row data table includes: storing the second slice data in a first physical area in the out-of-row data table; storing the first slice data and a physical address of the first physical area in a second physical area in the out-of-row data table;

[0010] The storing the first storage address in the first tuple of the data main table includes: storing the physical address of the second physical area in the first tuple of the data main table.

[0011] It can be seen that when the embodiment of the present application stores the first tuple in the database, if the first data in the first tuple meets the off-row storage condition, the first data is determined to be large field data, the first data is stored in the off-row data table, and the data head address of the first data in the off-row data table is stored in the data master table. In this way, based on the off-row storage address of the data, the data master table and the off-row data table are directly associated. In the data management process, the data stored in the off-row database can be directly managed based on the off-row storage address recorded in the data master table, which improves the data storage efficiency in high-concurrency data storage scenarios.

[0012] Optionally, the method further comprises:

[0013] Receive a data update request for the first tuple, the data update request is used to request to update the first data through the second data; if the second data meets the out-of-row storage condition, store the second data in the out-of-row data table, and obtain the second storage address of the second data in the out-of-row data table; update the first storage address in the first tuple of the data master table to the second storage address.

[0014] Optionally, after receiving the data update request for the first tuple, the method further includes: if the second data does not satisfy the out-of-row storage condition, updating the first tuple in the data master table based on the second data.

[0015] Optionally, after updating the first tuple in the data main table, the method further includes: deleting the first data in the out-of-row data table based on the first storage address.

[0016] It can be seen that in the embodiment of the present application, when the second data is used in the database to update the first data in the first tuple, if the second data meets the out-of-row storage condition, the second data is stored in the out-of-row data table, and the out-of-row storage address of the second data is used in the data master table to update the first storage address of the first data. In this way, based on the out-of-row storage address of the data, the data master table and the out-of-row data table are directly associated. In the data management process, the data stored in the out-of-row database can be directly managed based on the out-of-row storage address recorded in the data master table, which improves the data storage efficiency in high-concurrency data storage scenarios.

[0017] Optionally, the method further comprises:

[0018] Receive a data query request for the first tuple, wherein the data query request is used to query the data of a target field among the multiple fields; if the off-row storage address of the target field is stored in the data master table, read the data of the target field from the off-row data table based on the off-row storage address; and output the data query result of the first tuple based on the data of the target field.

[0019] Optionally, the data of the target field includes multiple slice data; and outputting the data query result of the first tuple based on the data of the target field includes: splicing the multiple slice data in the reading order of the multiple slice data; and outputting the query result of the first tuple based on the spliced ​​data.

[0020] It can be seen that when the embodiment of the present application queries the data of the first tuple in the database, if the data of the target field is stored in the off-row data table, after obtaining the off-row storage address of the target field in the data master table, based on the off-row storage address, the data of the target field is read from the off-row data table, and the data query result of the first tuple is output. In this way, based on the off-row storage address of the data, the data master table and the off-row data table are directly associated. In the data management process, the data stored in the off-row database can be directly managed based on the off-row storage address recorded in the data master table, which improves the data storage efficiency in high-concurrency data storage scenarios.

[0021] In a second aspect, a data management device is provided, wherein the data management device has the function of implementing the data management method in the first aspect. The data management device includes at least one module, and the at least one module is used to implement the data management method provided in the first aspect.

[0022] In a third aspect, a computer device is provided, the computer device comprising a processor and a memory, the memory being used to store a computer program for executing the data management method provided in the first aspect. The processor is configured to execute the computer program stored in the memory to implement the data management method described in the first aspect.

[0023] Optionally, the computer device may further include a communication bus, which is used to establish a connection between the processor and the memory.

[0024] In a fourth aspect, a computer-readable storage medium is provided, wherein the storage medium stores instructions, and when the instructions are executed on a computer, the computer executes the data management method described in the first aspect.

[0025] In a fifth aspect, a computer program product comprising instructions is provided, and when the instructions are executed on a computer, the computer is caused to execute the data management method described in the first aspect. In other words, a computer program is provided, and when the computer program is executed on a computer, the computer is caused to execute the steps of the data management method described in the first aspect.

[0026] The technical effects obtained by the above-mentioned second, third, fourth and fifth aspects are similar to the technical effects obtained by the corresponding technical means in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is a schematic diagram of a data table structure under a TOAST mechanism provided in an embodiment of the present application;

[0028] Figure 2This is a schematic diagram of a data storage process under a TOAST mechanism provided in an embodiment of the present application;

[0029] Figure 3 is a schematic diagram of a data management scenario provided in an embodiment of the present application;

[0030] Figure 4 It is a schematic diagram of the architecture of a data management system provided in an embodiment of the present application;

[0031] Figure 5 is a schematic diagram of the structure of a computing node provided in an embodiment of the present application;

[0032] Figure 6 It is a flowchart of a data management method provided in an embodiment of the present application;

[0033] Figure 7 It is a flowchart of another data management method provided in an embodiment of the present application;

[0034] Figure 8 It is a flowchart of another data management method provided in an embodiment of the present application;

[0035] Fig. 9 It is a schematic diagram of a data table structure under an ENTOAST mechanism provided in an embodiment of the present application;

[0036] Fig.10 This is a data storage flow chart under an ENTOAST mechanism provided in an embodiment of the present application;

[0037] Fig.11 This is a data update flow chart under an ENTOAST mechanism provided in an embodiment of the present application;

[0038] Fig.12 This is a data query flow chart under an ENTOAST mechanism provided in an embodiment of the present application;

[0039] Fig.13 It is a structural diagram of a data management device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0040] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the implementation methods of the present application will be further described in detail below in conjunction with the accompanying drawings.

[0041] For ease of understanding, before explaining in detail the memory page table caching method provided in the embodiment of the present application, the terminology, relevant background and implementation environment involved in the embodiment of the present application are first introduced.

[0042] First, the terms involved in the embodiments of the present application are introduced.

[0043] 1. Database management system (DBMS)

[0044] DBMS is a large-scale software for manipulating and managing databases. It is used to establish, use and maintain databases. It manages and controls the database in a unified manner to ensure the security and integrity of the database. Users access the data in the database through DBMS, and database administrators also perform database maintenance work through DBMS. It can support multiple applications and users to establish, modify and query the database in different ways at the same time or at different times. Most DBMS provide data definition language (DDL), data manipulation language (DML), and data query language (DQL) for users to define the database model structure and permission constraints, and implement operations such as appending and deleting data.

[0045] 2. Tuple

[0046] Simply put, in a database, a tuple refers to a record or a row of data, which is the most basic data unit in the database.

[0047] A tuple is actually a series of bytes. How to interpret these bytes depends on the database schema. The DBMS translates these bytes into values ​​of the corresponding type according to the table definition. The tuple is divided into two parts. The first part is the tuple header and the second part is the attribute data. The tuple header stores some metadata, such as visibility information, bitmaps, etc. The tuple data part generally stores the specific data under each attribute in the order of table definition.

[0048] 3. Page

[0049] In a database, a page is a continuous piece of physical storage space, usually 4KB or 8KB in size. Database management systems use pages to store data and indexes, and each page has a unique identifier called a page ID.

[0050] A page is the most basic storage unit in a database. All data and indexes are stored in a page. When the database needs to read or write data, it loads the page into memory for operation.

[0051] 4. Row identifier (tuple ID, TID)

[0052] In the openGauss database, TID is a special data type that indicates the physical location of a row. It consists of two parts: page ID and tuple ID.

[0053] 5. Large fields

[0054] In a database, a large field refers to a column (a column can also be understood as an attribute or field) that stores a large amount of data, usually binary data such as text, images, audio or video. In a database, since the size of a large field exceeds that of ordinary data types, large fields usually require special processing and storage methods.

[0055] 6. Out-of-row storage

[0056] Off-row storage refers to storing certain row data (i.e. tuple data) in a database in an independent storage medium instead of in the main table of the main database. This storage method can reduce the database load, reduce the latency of data access, and improve the performance and scalability of the database.

[0057] Among them, out-of-row storage is usually used to store large binary objects (BLOB) or large character objects (CLOB), because the data size of these objects may exceed the limit of the database management system.

[0058] Row chaining means that in a relational database, the data of a tuple is divided into multiple slices of data according to a fixed size, and they are connected using physical or logical location information, so that the complete data of a tuple can be obtained at one time when reading.

[0059] 8. Data manipulation language (DML)

[0060] DML is used to operate data in the database, including adding, deleting, modifying, etc. Common DML statements include insert (INSERT) statements, update (UPDATE) statements, delete (DELETE) statements, etc.

[0061] 9. Data query language (DQL)

[0062] DQL is used to query data in the database and does not modify the data. The most common DQL statement is the SELECT statement.

[0063] Secondly, the relevant background of the embodiments of the present application is introduced.

[0064] To address the storage problem of large field data in the database, openGauss provides an oversized-attribute storage technique (TOAST) to support the database in storing large fields.

[0065] When the user stores large field data exceeding 2KB, the TOAST mechanism will try to store the large field data out of line, that is, store the large field data in the TOAST table. When storing large field data out of line, first apply for a 32-bit unique identifier for the large field data as a chunk ID (Chunk ID). The Chunk ID is the index value of the out-of-line storage location. The Chunk ID is stored in the tuple to which the large field data belongs in the data master table; at the same time, the Chunk ID is recorded in the TOAST index table, as well as the out-of-line storage address indicating the large field data (i.e., the physical location in the TOAST table). When accessing large field data stored out of line, first obtain the Chunk ID corresponding to the large field data from the data master table, then determine the out-of-line storage address corresponding to the Chunk ID from the TOAST index table, and then use the out-of-line storage address to read the large field data from the TOAST table.

[0066] See also Figure 1 For large field data with a data length exceeding 2KB, assuming that the value of the location index (i.e., Chunk ID) stored in the data master table is 16630, then according to the Chunk ID in the data master table, the storage location of the large field data in the TOAST table (denoted as ctid) is determined from the TOAST index table as (0, 1) and (0, 2). Based on the ctid value, read the slice data No. 0 corresponding to the large field data in the first row of page 0 in the TOAST table, and read the slice data No. 1 corresponding to the large field data in the second row of page 0 to obtain the complete large field data.

[0067] Among them, in the TOAST mechanism, the data master table is used to store the visibility information of the tuple, which is used for visibility judgment in the data DML and DQL processes. At the same time, the data master table also stores the ChunkID corresponding to at least one large field data in the tuple; the TOAST index table is used to store the correspondence between the Chunk ID and the out-of-row storage address of the large field data; the TOAST table is used to store large field data.

[0068] See also Figure 2 , the out-of-row storage of large field data is realized through the TOAST mechanism, including the following steps:

[0069] (1) Determine whether the length of the tuple exceeds the threshold.

[0070] If it exceeds the threshold (for example, 2KB), the following step (2) is executed; if it does not exceed the threshold, the tuple can be inserted into the data main table.

[0071] (2) Try to perform data compression processing on the target field in the tuple, and again determine whether the length of the tuple after the data compression processing exceeds a threshold (for example, 2KB).

[0072] If it exceeds the threshold, execute the following step (3); if it does not exceed the threshold, insert the compressed tuple into the main data table.

[0073] (3) If the data length of the target field still exceeds 2KB after data compression, it is determined as large field data, and a unique identifier of type uint32 is applied for each large field data. That is, the identifier serves as the Chunk ID of the large field data, which is used to index the out-of-row storage location.

[0074] Specifically, the implementation process of applying for a Chunk ID for large field data is as follows: based on the TOAST index table, first add a global OID generation lock (referred to as OID lock) to the identifier serial number, and scan the latest identifier as the Chunk ID of the large field data, then add 1 to the identifier serial number and release the OID lock. Next, to ensure the unique validity of the Chunk ID, it is necessary to check whether the currently applied Chunk ID is repeated in the identifier serial number. If the Chunk ID is repeated, add 1 to the Chunk ID value and apply again. If there is an integer overflow when applying for the Chunk ID, continue to apply from the first valid identifier. If a valid Chunk ID cannot be applied for, continue to wait; if a valid Chunk ID is applied for, execute the following step (4).

[0075] (4) The large field data is sliced ​​according to a fixed size to obtain at least one slice data, and then traversed in sequence to insert at least one slice data into the TOAST table, and based on the Chunk ID of the large field, the physical address of each slice data is stored in the TOAST index table to update the TOAST index table.

[0076] (5) Based on the updated TOAST index table, the Chunk ID of the large field data is inserted into the corresponding tuple of the data master table, thus ending the out-of-row storage process of the large field data in the tuple.

[0077] However, when storing large field data in a database based on the TOAST mechanism, there is at least one of the following problems.

[0078] (1) The off-line storage of large field data relies on a globally incremented identifier sequence number (i.e., Chunk ID), and uses Chunk ID as the primary key index of the off-line storage address. As a result, in high-concurrency scenarios, operations such as applying for Chunk ID, checking the validity of Chunk ID (i.e., checking whether the applied Chunk ID is repeated in the identifier sequence number), inserting Chunk ID into the TOAST index table in increasing order, and checking the uniqueness of index values ​​in the TOAST index table (i.e., checking whether there are duplicate Chunk IDs in the TOAST index table) may introduce a large number of conflicting operations, which seriously degrades the concurrent performance of the off-line storage of large field data.

[0079] (2) As the number of chunk IDs in the TOAST index table increases, the time spent on chunk ID uniqueness verification will increase, resulting in a significant performance drop.

[0080] (3) Since the total number of globally incremented identifier serial numbers is 4B, at most 2 can exist at the same time. 32 An identifier serial number, that is, when applying for a Chunk ID, 2 32 There is a possibility of rollback of the identifier sequence number. 32 After the identifier sequence number is rolled over, it is necessary to reapply from the first identifier sequence number to use the applied identifier as the Chunk ID. Therefore, once the identifier sequence number rolls back, it is necessary to wait when applying for the Chunk ID, which seriously affects user services.

[0081] (4) The pointer used to indicate the out-of-row storage location in the current data master table occupies 18B. A single row (i.e., a single tuple) in the data master table cannot exceed the page size (usually 8KB). Therefore, the number of columns of a single tuple in the data master table (i.e., the number of fields or attributes contained in the tuple) is limited to (page size - page metadata) / 18B≈440 columns. In other words, the upper limit of the number of fields or attributes contained in a tuple in the data master table is less than 450, resulting in poor scalability of single tuple attributes.

[0082] Based on at least one of the above problems, an embodiment of the present application provides a data management method, which implements off-row storage, update, query and other operations of large field data through an enhanced over-size attribute storage mechanism (enhanced TOAST, abbreviated as ENTOAST), and directly stores the off-row storage address of the large field data in the data master table of the database, so that based on the off-row storage address, the large field data is directly read from the off-row data table, overcoming the problems caused by referencing Chunk ID for large field data management and improving data management efficiency.

[0083] Next, the application scenario, system architecture, and implementation environment of the embodiments of the present application are introduced.

[0084] The database can be implemented by one or more computing nodes, see Figure 3 The schematic diagram of the data management scenario based on the database is shown, in which the database client may include multiple clients, and the multiple clients can initiate data management requests for the database. Taking the database deployed in a single data node as an example, the database includes software modules such as the port listening daemon in the data kernel, the SQL engine and the storage engine, as well as hardware modules such as shared memory and disk, and the storage engine is directly connected to the shared memory.

[0085] Among them, the storage engine is an important component of the database management system, and its main function is to manage the storage and retrieval of data. In the embodiment of the present application, the storage engine can provide large field storage services and large field reading services, and the large field storage services and large field reading services use the method of off-line storage combined with row chaining to replace the traditional TOAST index table and TOAST table storage method.

[0086] See also Figure 4 When implementing the data management method provided in the embodiment of the present application, the SQL engine in the data kernel can process the natural language-based data request sent by the client into a machine language that can be recognized and executed by the machine. For example, in response to a data operation request initiated by the client, the SQL engine sends a data operation instruction (i.e., a DML statement) to the storage engine, and the storage engine responds to the data storage instruction and processes the requested large field data in the shared memory; in response to a data query request initiated by the client, the SQL engine sends a data query instruction (i.e., a DQL statement) to the storage engine, and the storage engine responds to the data query instruction and reads the requested large field data from the shared memory.

[0087] It should be understood that the shared memory can transfer the stored data to the disk, and the embodiment of the present application does not limit the timing of transferring data between the shared memory and the disk.

[0088] It should be noted that Figure 3 and Figure 4 The example of deploying a database on a single data node is taken as an example. In actual applications, the database can also be deployed on multiple data nodes. The software modules and hardware modules in each data node are deployed in the same way. The embodiment of the present application does not limit the number of data nodes deployed on the database.

[0089] The data management scenario described in the embodiment of the present application is intended to more clearly illustrate the technical solution of the embodiment of the present application, and does not constitute a limitation on the technical solution provided in the embodiment of the present application. Ordinary technicians in this field can know that with the emergence of new databases, the technical solution provided in the embodiment of the present application is also applicable to similar technical problems.

[0090] Please refer to Figure 5 , Figure 5 501, a communication bus 502, a memory 503, and at least one communication interface 504.

[0091] The processor 501 may be a general-purpose central processing unit (CPU), a network processor (NP), a microprocessor, or may be one or more integrated circuits for implementing the solution of the present application, such as an application-specific integrated circuit (ASIC), a programmable logic device (PLD) or a combination thereof. The above-mentioned PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL) or any combination thereof.

[0092] The communication bus 502 is used to transmit information between the above components. The communication bus 502 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0093] The memory 503 may be a read-only memory (ROM), a random access memory (RAM), an electrically erasable programmable read-only memory (EEPROM), an optical disc (including a compact disc read-only memory (CD-ROM), a compressed optical disc, a laser disc, a digital versatile disc, a Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 503 may exist independently and be connected to the processor 501 via the communication bus 502. The memory 503 may also be integrated with the processor 501.

[0094] The communication interface 504 uses any transceiver-like device for communicating with other devices or communication networks. The communication interface 504 includes a wired communication interface and may also include a wireless communication interface. Among them, the wired communication interface may be, for example, an Ethernet interface. The Ethernet interface may be an optical interface, an electrical interface, or a combination thereof. The wireless communication interface may be a wireless local area network (WLAN) interface, a cellular network communication interface, or a combination thereof, etc.

[0095] As an example, the processor 501 may include one or more CPUs, such as Figure 5 CPU0 and CPU1 are shown in the figure.

[0096] As an example, a data node may include multiple processors, such as Figure 5 501 and processor 505 are shown in FIG. Each of these processors may be a single-core processor or a multi-core processor. A processor herein may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0097] In some embodiments, the data node may further include an output device and an input device. The output device communicates with the processor 501 and may display information in a variety of ways. For example, the output device may be a liquid crystal display (LCD), a light emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector. The input device communicates with the processor 501 and may receive user input in a variety of ways. For example, the input device may be a mouse, a keyboard, a touch screen device, or a sensor device.

[0098] In some embodiments, the memory 503 is used to store the program code 510 for executing the solution of the present application, and the processor 501 can execute the program code 510 stored in the memory 503. The program code 510 may include one or more software modules, and the data node may implement the following by the processor 501 and the program code 510 in the memory 503. Figure 6 The data management method provided by the embodiment.

[0099] Next, the data management method provided in the embodiment of the present application is described in detail.

[0100] Figure 6 This is a flow chart of a data management method provided in an embodiment of the present application, which can be applied to the above data node. Figure 6 , the method comprises the following steps.

[0101] Step 601: Receive a data storage request for a first tuple, where the first tuple includes data of multiple fields.

[0102] Among them, the first tuple is any tuple requested to be stored in the database. Among the multiple fields included in the first tuple, the data length of some fields may exceed the data length threshold, and the data length of some fields may not exceed the data length threshold. For data exceeding the data length threshold, it is large field data and cannot be directly stored in the main data table of the database. It needs to be compressed or stored out of row.

[0103] Among them, the data length threshold can be set based on the page size of the data main table and / or the out-of-row data table. Since tuples cannot be stored across pages, when the page size is 8KB and one page stores 4 tuples, the data length threshold can be set to 2KB. Of course, in actual application, the data length threshold can also be set to other values, such as 4KB, etc., and the embodiments of the present application do not limit this.

[0104] In some embodiments, after receiving the data request for the first tuple, first determine whether the first tuple meets the out-of-row storage condition based on the tuple length of the first tuple. If the out-of-row storage condition is met, consider storing some or all fields of the first tuple out-of-row; if the out-of-row storage condition is not met, directly store the first tuple in the data master table.

[0105] The tuple length refers to the total length of the data in each field in the tuple, that is, the total number of bytes; the out-of-row storage condition is that the data length is greater than the data length threshold.

[0106] Optionally, when the tuple length of the first tuple is greater than the data length threshold, that is, the out-of-row storage condition is met, data compression processing can be performed on the large field data in the first tuple whose data length exceeds the data length threshold, and based on the data compression result, the tuple length of the first tuple is recalculated. If the tuple length of the first tuple after data compression processing is greater than the data length threshold, the large field data in the first tuple is stored out-of-row; if the tuple length of the first tuple after data compression processing is less than or equal to the data length threshold, the compressed large field data in the first tuple is stored in the data master table.

[0107] Among the multiple fields of the first tuple, some fields support compression. Therefore, if the first tuple meets the out-of-row storage condition, the data in the fields that support compression in the first tuple can be compressed first. If the first tuple still meets the out-of-row storage condition after the data compression process, the large field data in the first tuple will be stored out-of-row.

[0108] It should be noted that, in the process of determining whether to store the first tuple out of line based on the tuple length of the first tuple, if there is large field data in the first tuple, the large field data can be directly stored out of line. In order to reduce the physical area occupied when the large field data is stored out of line, the large field data can be compressed first, and based on the data compression result, the compressed data can also be stored in the out-of-line data table.

[0109] That is, the first tuple after compression meets the out-of-row storage condition, indicating that the total length of the data of the compressed field and the uncompressed field is greater than 2KB. At this time, for the compressed field, its data length may exceed 2KB or may not exceed 2KB. However, since the tuple length is greater than the data length threshold, the large field data after compression in the first tuple can be stored out-of-row.

[0110] As an example, if the first tuple includes field 1, field 2, field 3, field 4 and field 5, if field 2, field 3 and field 5, and the tuple length of the first tuple exceeds the data length threshold. At this time, the data of field 2, field 3 and field 5 that support compression processing are first compressed. If after the data compression processing, the data length of field 2 is greater than the data length threshold, the data length of field 3 and field 5 is less than the data length threshold, and the tuple length of the first tuple is greater than the data length threshold, then the compressed data of field 2, field 3 and field 5 are stored out of line. If after the data compression processing, the data lengths of field 2, field 3 and field 5 are all less than the data length threshold, at this time, the tuple length of the first tuple may be less than the data length threshold, so the first tuple can be stored in the data master table, wherein field 1 and field 4 store the original data, and field 2, field 3 and field 5 store the compressed data; of course, the compressed data of field 2, field 3 and field 5 can also be directly stored out of line.

[0111] Step 602: If the first data meets the out-of-row storage condition, the first data is stored in the out-of-row data table, and the first data is data of any field in the plurality of fields.

[0112] In an embodiment of the present application, the off-row data table is an ENTOAST table, which is used to store large field data that needs to be stored off-row in the data main table, as well as the correspondence between the large field data and the physical address.

[0113] In other words, if the first data is original data whose data length is greater than the data length threshold, or compressed data whose data length after data compression processing is greater than the data length threshold, or large field data in the first tuple when the tuple length of the first tuple after data compression processing is greater than the data length threshold.

[0114] In a possible implementation, the implementation process of storing the first data in the out-of-row data table is: slicing the first data to obtain at least one slice data; and storing the at least one slice data in the out-of-row data table.

[0115] The first data may be sliced ​​according to a slice length, which is less than or equal to the data length threshold, for example, the slice length may be 1KB, 2KB, etc., which is not limited in the present embodiment.

[0116] If at least one slice data includes first slice data and second slice data connected sequentially, the implementation process of storing at least one slice data in the out-of-row data table may be: storing the second slice data in a first physical area in the out-of-row data table; storing the first slice data and the physical address of the first physical area in a second physical area in the out-of-row data table.

[0117] It should be noted that "sequential connection" means that the first slice data is located before the second slice data in the first data. However, when storing data in the out-of-row data table, the implementation of this application is based on the storage order from "back to front", storing the second slice data first and then the first slice data. In other words, the connection order of the first slice data and the second slice data is opposite to the storage order in the out-of-row data table.

[0118] In the out-of-row data table, each physical area corresponds to a physical address, and the physical address of the slice data stored first will be stored in the physical area where the next slice data is stored. In this way, the physical addresses of multiple slice data are connected in series, and based on the physical address stored in the physical area of ​​a slice data, the next slice data can continue to be indexed.

[0119] Optionally, when storing the first data in an out-of-row data table, not only the first data itself can be stored, but also data association information corresponding to the first data can be stored in the second physical area where the first slice data is located, such as the original length of the field data (rawsize), the compressed length of the field data (textsize), etc. The embodiment of the present application does not impose any restrictions on this.

[0120] Step 603: Obtain a first storage address of the first data in the out-of-row data table.

[0121] Since the physical address of the first physical area is recorded in the second physical area, after the second physical area where the first slice data is located is determined in the out-of-row data table, the physical address of the first physical area can be read from the physical area, thereby continuing to read the second slice data from the first physical area.

[0122] Therefore, if it is necessary to obtain at least one slice data stored in the out-of-row data table of the first data, it is only necessary to determine the storage position of the first slice data in the out-of-row data table of the first data.

[0123] In a possible implementation, step 603 may be implemented by obtaining a physical address of a first slice of the first data in the out-of-row data table to obtain a first storage address. In other words, the first storage address is the first address of the first data in the out-of-row data table.

[0124] Step 604: Store the first storage address in the first tuple of the data main table.

[0125] That is, the data head address of the first data in the off-row data table is stored in the first tuple of the data main table.

[0126] Taking the example that at least one slice data of the first data includes first slice data and second slice data connected in sequence, the implementation process of step 604 is: storing the physical address of the second physical area storing the first data slice in the first tuple of the data master table.

[0127] It can be seen that when the embodiment of the present application stores the first tuple in the database, if the first data in the first tuple meets the off-row storage condition, the first data is determined to be large field data, the first data is stored in the off-row data table, and the data head address of the first data in the off-row data table is stored in the data master table. In this way, based on the off-row storage address of the data, the data master table and the off-row data table are directly associated. In the data management process, the data stored in the off-row database can be directly managed based on the off-row storage address recorded in the data master table, which improves the data storage efficiency in high-concurrency data storage scenarios.

[0128] Based on the above steps 601 to 604, after the data storage operation of the first tuple is completed in the data main table and the off-row data table, at least one large field data in the first tuple can also be updated based on the database. Figure 7 The data management method provided in the embodiment of the present application also includes the following steps.

[0129] Step 605: Receive a data update request for the first tuple, where the data update request is used to request to update the first data using the second data.

[0130] In a possible implementation, the data update request for the first tuple carries the tuple ID of the first tuple and the update data of each field in the first tuple to indicate the update data of each field, and updates the first tuple in the data master table, that is, requests to update the first tuple as a whole.

[0131] In this case, the first tuple can be updated based on the updated data of each field; or the above steps 601-604 can be used in the data main table to regenerate a row of data to store the updated first tuple and mark the first tuple before the update for deletion.

[0132] In another possible implementation, the data update request carries the tuple ID of the first tuple, the field identifier of the first field, and the second data, to indicate that the data of the first field in the first tuple is updated with the second data, that is, a request is made to update the data of a single field in the first tuple.

[0133] In this case, the first data in the first tuple may be updated based on the second data, for example, the first storage address of the first data in the data master table may be updated.

[0134] The embodiment of the present application does not limit the information carried in the data update request, and may include data of all fields after the first tuple is updated, or may only include data of at least one field that needs to be updated.

[0135] Step 606: If the second data meets the out-of-row storage condition, the second data is stored in the out-of-row data table, and a second storage address of the second data in the out-of-row data table is obtained.

[0136] Since the data length of the second data may be greater than the data length threshold, or may be less than or equal to the data length threshold, before using the second data to update the first data, it is necessary to determine whether the second data can be directly stored in the data master table.

[0137] As described above, the off-row storage condition is that the data length is greater than the data length threshold. Therefore, when the data length of the second data is greater than the data length threshold, it means that the second data is large field data and cannot be directly stored in the data master table. In this case, the second data needs to be stored off-row. When the data length of the second data is less than or equal to the data length threshold, it means that the second data is data that can be normally stored in the database. In this case, the second data is directly stored in the data master table.

[0138] Similarly, the second storage address of the second data in the out-of-row data table is the data head address of the second data in the out-of-row data table. In other words, the second storage address is the physical address of the physical area corresponding to the first slice data (that is, the last slice data to perform the out-of-row storage operation) corresponding to the second data in the out-of-row data table.

[0139] Step 607: Update the first storage address in the first tuple of the data main table to the second storage address.

[0140] When the second data meets the out-of-row storage condition, after determining the second storage address of the second data in the out-of-row data table, delete the first storage address of the first data at the field corresponding to the first data in the data main table and write the second storage address.

[0141] In some embodiments, if the second data does not satisfy the out-of-row storage condition, the first tuple in the data master table is updated based on the second data.

[0142] That is, when the data length of the second data is less than the data length threshold, the second data can be directly stored in the data master table without the need for out-of-row storage.

[0143] In one possible implementation, when the first tuple in the data main table is updated based on the second data, a second tuple can be added to the data main table based on the first tuple and the second data, and the second tuple includes the second data and is the same as the data of other fields in the first tuple; the tuple including the first data is deleted from the data main table to implement the update of the first tuple.

[0144] In another possible implementation, when the first tuple in the data main table is updated based on the second data, the first storage address of the first data may be deleted, and the second data may be written to the field corresponding to the first data.

[0145] In some embodiments, when the first data in the above-mentioned first tuple does not meet the out-of-row storage condition and the first data is stored in the data master table, the process of using the second data to update the first data is as follows: if the data length of the second data is greater than the data length of the first data, the second data cannot directly update the first data at the original field position. At this time, based on the first tuple and the second data, a second tuple can be added to the data master table, the second tuple includes the second data and is the same as the data of other fields in the first tuple; the first tuple including the first data is deleted from the data master table to update the first tuple. If the data length of the second data is less than or equal to the data length of the first data, the second data is used to directly update the first data at the original field position (data overwriting) to update the first tuple.

[0146] In addition, when the first data is stored in the off-row data table, after the first tuple is updated in the data main table using the second data, the first data needs to be deleted in the off-row data table based on the first storage address of the first data stored in the data main table.

[0147] It can be seen that in the embodiment of the present application, when the second data is used in the database to update the first data in the first tuple, if the second data meets the out-of-row storage condition, the second data is stored in the out-of-row data table, and the out-of-row storage address of the second data is used in the data master table to update the first storage address of the first data. In this way, based on the out-of-row storage address of the data, the data master table and the out-of-row data table are directly associated. In the data management process, the data stored in the out-of-row database can be directly managed based on the out-of-row storage address recorded in the data master table, which improves the data storage efficiency in high-concurrency data storage scenarios.

[0148] Based on the above steps 601 to 604, after the first tuple is stored in the data master table, a data query request for at least one field in the first tuple can be initiated based on the database to obtain data of at least one field in the first tuple from the database. Figure 8 The data management method provided in the embodiment of the present application also includes the following steps.

[0149] Step 608: Receive a data query request for the first tuple, where the data query request is used to query data of a target field among multiple fields.

[0150] The target fields may be some fields in the first tuple, or may be all fields in the first tuple. The embodiment of the present application does not limit the number of target fields.

[0151] Step 609: If the out-of-row storage address of the target field is stored in the data main table, the data of the target field is read from the out-of-row data table based on the out-of-row storage address.

[0152] As mentioned above, when the data of a field is stored in an off-line data table, only the off-line storage address of the data is stored at the location of the field in the data master table. Therefore, when the off-line storage address of the target field is stored in the data master table, it means that the data of the target field is stored in the off-line data table, and the data of the target field needs to be obtained from the off-line data table based on the off-line storage address.

[0153] In a possible implementation, the implementation process of reading the data of the target field from the off-row data table based on the off-row storage address is as follows: based on the off-row storage address, the first slice data is obtained from the corresponding physical area in the off-row data table. If the physical address is still stored in the physical area indicated by the off-row storage address, the next slice data is read based on the physical address, and so on, until the physical address is not stored in the physical area of ​​the slice data, and multiple slice data corresponding to the target field are obtained. If the physical address is not stored in the physical area indicated by the off-row storage address, it means that the data of the target field is not sliced ​​when it is stored out of the row, and the first slice data is the complete data of the target field.

[0154] Step 610: Based on the data of the target field, output the data query result of the first tuple.

[0155] When the data of the target field includes multiple slice data, in a possible implementation, the implementation process of step 610 is: splicing the multiple slice data according to the reading order of the multiple slice data; and outputting the query result of the first tuple based on the spliced ​​data.

[0156] After splicing multiple slice data, the complete data of the target field can be obtained. At this time, the query result of the first tuple can be output based on the complete data of the target field.

[0157] Optionally, the query result of the first tuple may include all the data of the first tuple, or may only include the data of the target field in the first tuple, which is not limited in this embodiment of the present application.

[0158] It can be seen that when the embodiment of the present application queries the data of the first tuple in the database, if the data of the target field is stored in the off-row data table, after obtaining the off-row storage address of the target field in the data master table, based on the off-row storage address, the data of the target field is read from the off-row data table, and the data query result of the first tuple is output. In this way, based on the off-row storage address of the data, the data master table and the off-row data table are directly associated. In the data management process, the data stored in the off-row database can be directly managed based on the off-row storage address recorded in the data master table, which improves the data storage efficiency in high-concurrency data storage scenarios.

[0159] In some embodiments, after receiving a data deletion request for the first tuple, the data of each field in the first tuple can be determined in the data main table and the off-row data table according to the above step 609, and then the first tuple can be deleted in the data main table and the row data table.

[0160] That is, based on the above embodiment, the tuple deletion operation can be decomposed into a data query operation and a data deletion operation. After determining the specific position of the first tuple in the main data table and the off-row data table, the deletion operation can be performed on the data of the corresponding first tuple. The embodiment of the present application will not be elaborated here.

[0161] Based on the above method embodiments, next, in combination with the attached Figure 9-12 , the specific process of data management based on the ENTOAST mechanism in the embodiment of the present application is described in detail.

[0162] The data management method of the embodiment of the present application involves two data tables, namely, a data master table in the database and an out-of-row data table (i.e., an ENTOAST table). In the data master table, based on the structure of the target field, if the field data of the target field is large field data larger than 2KB, when storing the large field data in the ENTOAST table, the large field data is first sliced ​​to obtain at least one slice data, and then the slice data is stored in the ENTOAST table in sequence according to the storage order of "from back to front". After the last slice data is stored, the row identifier (TID) of the last slice data in the ENTOAST table is stored in the corresponding tuple in the data master table.

[0163] That is, the off-row storage address recorded in the data master table in the embodiment of the present application is: the TID (page ID, row ID) of the field data of the target field in the off-row data table.

[0164] As an example, see Fig. 9 , assuming that the large field data of the target field is sliced ​​to obtain the first slice data and the second slice data connected in sequence, and it is determined that the first row and the second row of page 0 of the ENTOAST table are free positions, the second slice data is first stored in the second row, and then the physical addresses (0, 2) of the first slice data and the second row are stored in the first row, and the first address of the large field data in the off-row data table is returned to the data main table, that is, the physical address (0, 1) of the first slice data in the off-row data table, so that the off-row storage address stored at the target field position in the data main table is (0, 1).

[0165] Optionally, when storing the large field data of the target field in the ENTOAST table, data association information of the large field data can also be stored in the storage space of the first slice data, such as the original length of the field data (rawsize), the compressed length of the field data (textsize), etc. The embodiment of the present application does not impose any restrictions on this.

[0166] See also Fig. 9 In the storage area of ​​the first slice data, the data association information of the large field data of the target field is stored in the meta area. In this way, after obtaining the last first slice data, the data association information of the target field can be determined.

[0167] During the data reading process, based on the off-row storage address TID=(0,1) stored in the data master table, the first slice data is obtained in the first row of page 0 in the ENTOAST table, and then based on the physical address (0,2) stored in the first row, the second slice data is read in the second row of page 0. Since the physical address stored in the second row is (0,0), the second slice data stored in the second row is the last slice data. In this way, based on the off-row storage address stored in the data master table, the two slice data corresponding to the target field can be read from the ENTOAST table, and after splicing the two slice data, the complete data of the target field can be obtained.

[0168] See also Fig.10The large field data storage process based on the ENTOAST mechanism provided in the embodiment of the present application is: listen to the data insertion message, if it is listened that a tuple needs to be stored in the database, first determine whether the tuple length of the tuple is greater than the data length threshold, if the tuple length is less than or equal to the data length threshold, then directly insert the tuple into the data master table; if the tuple length is greater than the data length threshold, then perform data compression processing on the target field that supports compression in the tuple, and determine whether the tuple length of the tuple is greater than the data length threshold after the data compression processing, if the tuple length after the data compression processing is less than or equal to the data length threshold, then directly insert the tuple after the data compression processing into the data master table; if the tuple length after the data compression processing is greater than the data length threshold, then the large field data of the target field greater than the data length threshold after the data compression processing is stored out of row.

[0169] When performing off-row storage, the large field data is sliced, and at least one slice data obtained after the slicing process is stored in the ENTOAST table, and the physical addresses of the physical areas of each slice data are used for concatenation. When all slice data of the large field data are stored in the ENTOAST table, the data head address of the large field data in the off-row data table (that is, the physical address of the last stored slice data) is returned, and the data head address is stored in the corresponding tuple of the data main table, ending the data storage process.

[0170] Optionally, when inserting slice data of the target data into the ENTOAST table in sequence, in the process of traversing the storage location, the tuple data marked as deleted in the traversed page can be deleted, and the physical space occupied by it can be released.

[0171] See also Fig.11 The large field data update process based on the ENTOAST mechanism provided in the embodiment of the present application is: monitor the data update message. If it is monitored that the new version tuple needs to be used to update the old version tuple in the database, first perform visibility judgment on the old version tuple in the data master table. If the old version tuple is visible, the data in the old version tuple is updated based on the new version tuple; if the old version tuple is not visible, it means that the old version tuple is not stored in the data master table, and the data update process is terminated directly.

[0172] If the old version tuple exists in the data master table, determine whether the tuple length of the new version tuple exceeds the data length threshold. If the tuple length is less than or equal to the data length threshold, directly use the new version tuple to update the old version tuple; if the tuple length of the new version tuple is greater than the data length threshold, slice the new version tuple and store it out of row. Then, based on the out-of-row storage address, update the TID stored in the old version tuple in the data master table.

[0173] Furthermore, after storing the out-of-row storage address of the new version tuple in the data master table, it is necessary to mark the data corresponding to the TID in the out-of-row data table as invalid based on the TID of the old version tuple, thereby ending the data update process.

[0174] It should be understood that only updating the data of the entire tuple is used as an example here. As described above, when updating a tuple, the data of all fields may be updated, or only the data of at least one field may be updated. The update logic is similar.

[0175] See also Fig.12 The large field data query process based on the ENTOAST mechanism provided in the embodiment of the present application is: listen to the data query request. If it is listened to that the data of a tuple needs to be queried in the database, first determine whether the tuple is visible in the main table of the database. If the tuple is visible, then based on the main table of the data and the off-row data table, obtain the data of the tuple; if the tuple is not visible, it means that the tuple is not stored in the main table of the data, and then the data query process ends.

[0176] In the case where the tuple exists in the data master table, if the target field in the tuple has an out-of-row storage address TID, then based on the out-of-row storage address TID, at least one slice data is obtained in the out-of-row data table, and then the at least one slice data is spliced. Based on the spliced ​​data, the data query result of the tuple is output, and the data query process ends.

[0177] Optionally, when obtaining slice data of the target field in the ENTOAST table, tuple data marked as deleted in the traversed pages may be deleted, and the physical space occupied by the tuple data may be released.

[0178] It should be understood that here only the example of querying the data of the entire tuple is used. As described above, when querying the tuple, you can query to obtain the data of all fields, or you can query only the data of at least one field. The query logic is similar.

[0179] To summarize, the data management method provided in the embodiment of the present application directly associates the data main table and the off-row data table based on the off-row storage address of the data. During the data management process, for a certain field in the tuple, the data stored in the off-row database can be directly managed based on the off-row storage address recorded in the data main table, thereby improving data management efficiency in high-concurrency data management scenarios.

[0180] Fig.13 is a structural diagram of a data management device provided in an embodiment of the present application. The data management device can be implemented as part or all of a data node by software, hardware, or a combination of both. Fig.13The data management device 1300 includes: a request receiving module 1301, a data storage module 1302, an address obtaining module 1303 and an address storage module 1304.

[0181] A request receiving module 1301 is used to receive a data storage request for a first tuple, where the first tuple includes data of multiple fields;

[0182] A data storage module 1302, configured to store the first data in an out-of-row data table if the first data meets an out-of-row storage condition, wherein the first data is data of any field in a plurality of fields;

[0183] An address acquisition module 1303 is used to acquire a first storage address of the first data in the out-of-row data table;

[0184] The address storage module 1304 is used to store the first storage address in the first tuple of the data main table.

[0185] Optionally, the data storage module 1302 includes:

[0186] A data slicing unit, configured to perform slicing processing on the first data to obtain at least one slice data;

[0187] The data storage unit is used to store at least one slice data in an out-of-row data table.

[0188] Optionally, the at least one slice data includes first slice data and second slice data connected in sequence; the data storage unit is specifically used to:

[0189] storing the second slice data in a first physical area in the out-of-row data table;

[0190] storing the first slice data and the physical address of the first physical area in the second physical area of ​​the out-of-row data table;

[0191] The address storage module 1303 is specifically used for:

[0192] The physical address of the second physical area is stored in the first tuple of the data master table.

[0193] Optionally, in the data management device 1300:

[0194] The request receiving module 1301 is further used to receive a data update request of the first tuple, where the data update request is used to request to update the first data by using the second data;

[0195] The data storage module 1302 is further configured to store the second data in the out-of-row data table if the second data meets the out-of-row storage condition, and obtain a second storage address of the second data in the out-of-row data table;

[0196] The address storage module 1303 is further used to update the first storage address in the first tuple of the data main table to the second storage address.

[0197] Optionally, after receiving the data update request of the first tuple, the data management device 1300 further includes:

[0198] The data updating module is used to update the first tuple in the data main table based on the second data if the second data does not meet the out-of-row storage condition.

[0199] Optionally, after updating the first tuple in the data master table, the data management device 1300 further includes:

[0200] The data updating module is further used to delete the first data in the out-of-row data table based on the first storage address.

[0201] Optionally, the data management device 1300 further includes:

[0202] The request receiving module is further used to receive a data query request of the first tuple, where the data query request is used to query data of a target field in the multiple fields;

[0203] A data reading module is used to read the data of the target field from the off-row data table based on the off-row storage address if the off-row storage address of the target field is stored in the data master table;

[0204] The result output module is also used to output the data query result of the first tuple based on the data of the target field.

[0205] Optionally, the data of the target field includes a plurality of slice data; the result output module includes:

[0206] A data splicing unit, used for splicing a plurality of slice data according to the reading order of the plurality of slice data;

[0207] The result output unit is used to output the query result of the first tuple based on the concatenated data.

[0208] In an embodiment of the present application, the data management device directly associates the data main table and the off-row data table based on the storage address of the data in the off-row data table. During the data management process, the data in the off-row database can be directly managed based on the storage address recorded in the data main table, thereby improving data management efficiency in high-concurrency data management scenarios.

[0209] It should be noted that: when the data management device provided in the above embodiment manages the data stored in the database, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the data management device provided in the above embodiment and the data management method embodiment belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0210] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a digital versatile disc (DVD)), or a semiconductor medium (e.g., a solid state disk (SSD)). It is worth noting that the computer-readable storage medium mentioned in the embodiment of the present application may be a non-volatile storage medium, in other words, a non-transient storage medium.

[0211] It should be understood that the "multiple" mentioned herein refers to two or more. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in this article is only a description of the association relationship of the associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in order to facilitate the clear description of the technical solution of the embodiments of the present application, in the embodiments of the present application, the words "first", "second" and the like are used to distinguish between the same items or similar items with basically the same functions and effects. Those skilled in the art can understand that the words "first", "second" and the like do not limit the quantity and execution order, and the words "first", "second" and the like do not limit them to be necessarily different.

[0212] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in the embodiments of the present application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions.

[0213] The above-mentioned embodiments are provided for the present application and are not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A data management method, characterized in that: The method comprises: receiving a data storage request for a first tuple, the first tuple including data of a plurality of fields; If the first data meets the out-of-row storage condition, the first data is stored in the out-of-row data table, and the first data is data of any field in the multiple fields; Obtain a first storage address of the first data in the out-of-row data table; The first storage address is stored in the first tuple of the data master table.

2. The method according to claim 1, characterized in that The storing the first data into an out-of-row data table includes: Performing slicing processing on the first data to obtain at least one slice data; The at least one slice data is stored in the out-of-row data table.

3. The method according to claim 2, characterized in that The at least one slice data includes first slice data and second slice data connected in sequence; The storing the at least one slice data into the out-of-row data table comprises: storing the second slice data in a first physical area in the out-of-row data table; storing the first slice data and a physical address of the first physical area in a second physical area in the out-of-row data table; The storing the first storage address in the first tuple of the data main table includes: The physical address of the second physical area is stored in the first tuple of the data master table.

4. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: receiving a data update request for the first tuple, wherein the data update request is used to request to update the first data by using second data; If the second data meets the out-of-row storage condition, the second data is stored in the out-of-row data table, and a second storage address of the second data in the out-of-row data table is obtained; The first storage address in the first tuple of the data master table is updated to the second storage address.

5. The method according to any one of claims 1 to 3, characterized in that: After receiving the data update request of the first tuple, the method further includes: If the second data does not satisfy the out-of-row storage condition, the first tuple in the data master table is updated based on the second data.

6. The method according to claim 4 or 5, characterized in that After updating the first tuple in the data master table, the method further includes: Based on the first storage address, the first data in the out-of-row data table is deleted.

7. The method according to any one of claims 1 to 6, characterized in that: The method further comprises: receiving a data query request for the first tuple, wherein the data query request is used to query data of a target field among the multiple fields; If the out-of-row storage address of the target field is stored in the data master table, then based on the out-of-row storage address, the data of the target field is read from the out-of-row data table; Based on the data of the target field, a data query result of the first tuple is output.

8. The method according to claim 7, characterized in that The data of the target field includes a plurality of slice data; and outputting the data query result of the first tuple based on the data of the target field includes: splicing the plurality of slice data according to the reading order of the plurality of slice data; Based on the concatenated data, the query result of the first tuple is output.

9. A data management device, characterized in that: The device comprises: A request receiving module, configured to receive a data storage request for a first tuple, wherein the first tuple includes data of a plurality of fields; A data storage module, configured to store the first data in an out-of-row data table if the first data meets an out-of-row storage condition, wherein the first data is data of any field in the multiple fields; An address acquisition module, used for acquiring a first storage address of the first data in the out-of-row data table; The address storage module is used to store the first storage address in the first tuple of the data main table.

10. The device according to claim 9, characterized in that The data storage module comprises: A data slicing unit, configured to perform slicing processing on the first data to obtain at least one slice data; A data storage unit is used to store the at least one slice data in the out-of-row data table.

11. The device according to claim 10, characterized in that The at least one slice data includes first slice data and second slice data connected in sequence; the data storage unit is specifically used for: storing the second slice data in a first physical area in the out-of-row data table; storing the first slice data and a physical address of the first physical area in a second physical area in the out-of-row data table; The address storage module is specifically used for: The physical address of the second physical area is stored in the first tuple of the data master table.

12. The device according to any one of claims 9 to 11, characterized in that: The device comprises: The request receiving module is further used to receive a data update request for the first tuple, wherein the data update request is used to request to update the first data through second data; The data storage module is further configured to store the second data in the out-of-row data table if the second data meets the out-of-row storage condition, and obtain a second storage address of the second data in the out-of-row data table; The address storage module is further used to update the first storage address in the first tuple of the data main table to the second storage address.

13. The device according to any one of claims 9 to 11, characterized in that: After receiving the data update request of the first tuple, the device further includes: A data updating module is used to update the first tuple in the data main table based on the second data if the second data does not meet the out-of-row storage condition.

14. The device according to claim 12 or 13, characterized in that After the first tuple in the data master table is updated, the device further includes: The data updating module is further used to delete the first data in the out-of-row data table based on the first storage address.

15. The device according to any one of claims 9 to 14, characterized in that: The device also includes: The request receiving module is further used to receive a data query request for the first tuple, wherein the data query request is used to query data of a target field among the multiple fields; A data reading module, configured to read the data of the target field from the off-row data table based on the off-row storage address if the off-row storage address of the target field is stored in the data master table; The result output module is further used to output the data query result of the first tuple based on the data of the target field.

16. The device according to claim 15, characterized in that The data of the target field includes a plurality of slice data; the result output module includes: A data splicing unit, configured to splice the plurality of slice data according to a reading order of the plurality of slice data; A result output unit is used to output the query result of the first tuple based on the concatenated data.

17. A computer device, characterized in that: The computer device includes a memory and a processor; The memory is used to store computer programs; The processor is used to execute the computer program to implement the data management method according to any one of claims 1 to 8.

18. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the data management method according to any one of claims 1 to 8 is implemented.

19. A computer program product, characterized in that The computer program product stores computer instructions, and when the computer instructions are executed by a processor, the data management method according to any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • Data management method and apparatus, and device, storage medium and computer program

    EP4797112A1

  • Data management method and apparatus, and device, storage medium and computer program

    WO2025091965A1