Data management method and apparatus, and device, storage medium and computer program
By storing the off-row storage address of large field data in the database, directly linking the main data table and off-row data table, the performance problem of the TOAST mechanism in high concurrency scenarios is solved, and efficient data management and storage is achieved.
Patent Information
- Application Number
- PCT/CN2024/101437
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-03
- Filing Date
- 2024-06-25
- Publication Date
- 2025-05-08
AI Technical Summary
In the high concurrency data storage scenario, the existing TOAST mechanism relies on globally incremented Chunk ID, which makes it take a long time to apply for Chunk ID, and the identifier serial number rewinding may lead to serious performance degradation.
By storing the off-row storage address of large field data in the database, instead of relying on Chunk ID, the main data table and off-row data table are directly associated to achieve high concurrent data storage and management.
Improve the efficiency of data storage and management, avoid the performance bottlenecks caused by Chunk ID application and verification, and enhance the processing capabilities of the database in high concurrency scenarios.
Smart Images

Figure CN2024101437_08052025_PF_FP_ABST
Abstract
Description
Data management method, device, equipment, storage medium and computer program
[0001] This application claims priority to Chinese patent application number 202311464440.8, filed on November 3, 2023, entitled “Data management method, device, equipment, storage medium and computer program”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of database technology, and in particular to a data management method, apparatus, device, storage medium, and computer program. Background Art
[0003] In a database, a tuple refers to a record or row of data in the database's main data table. A tuple consists of a tuple header and attribute data. The tuple header stores metadata, such as visibility information and bitmaps, while the attribute data includes data for multiple attributes. Each attribute can also be considered a field. Among these multiple attributes, the data length of one attribute may exceed a threshold. This attribute is called a large field, and the data for this attribute is called large field data. However, large field data cannot be stored directly in the main data table.
[0004] In the related art, openGauss (an open source relational database) provides a TOAST mechanism to support the database to store large field data. When a user stores data exceeding 2KB bytes, the data can be used as large field data, and the TOAST mechanism will attempt to store the large field data off-line, that is, store the large field data in a TOAST table. When storing large field data off-line, first apply for a 32-bit unique identifier for the large field data as a chunk ID (Chunk ID). The Chunk ID is the index value of the off-line storage location of the index, and the Chunk ID is stored in the tuple to which the large field data belongs in the data master table; at the same time, the Chunk ID is recorded in the TOAST index table, as well as the off-line storage address indicating the large field data (i.e., the physical location in the TOAST table). When accessing the large field data stored off-line, first obtain the Chunk ID corresponding to the large field data from the data master table, then determine the off-line storage address corresponding to the Chunk ID from the TOAST index table, and then use the off-line storage address to read the large field data from the TOAST table.
[0005] However, the above-mentioned off-line storage method relies on globally incremented Chunk IDs. However, in scenarios with high concurrency of data storage requests, it takes a long time to apply for Chunk IDs for large field data stored off-line. Moreover, the TOAST mechanism supports a maximum of 2 32The identifier serial number is used as the Chunk ID. When applying for the Chunk ID, 2 32 There is a possibility of rollback of the identifier serial number, that is, the second 32 After the identifier sequence number is rolled over, it is necessary to re-apply from the first identifier sequence number to use the applied identifier sequence number as the Chunk ID. Therefore, once the identifier sequence number rolls over, it will take a long time to apply for the Chunk ID, seriously affecting data storage performance.
[0006] Summary of the Invention
[0007] This application provides a data management method, apparatus, device, storage medium, and computer program that can improve data management efficiency. The technical solution is as follows:
[0008] In a first aspect, a data management method is provided, the method comprising:
[0009] Receive a data storage request for a first tuple, where the first tuple includes data of multiple fields; if the first data meets an out-of-row storage condition, store the first data in an out-of-row data table, where the first data is data of any field of the multiple fields; obtain a first storage address of the first data in the out-of-row data table; and store the first storage address in the first tuple of the data master table.
[0010] It can be seen from this that when the present application stores the first tuple in the database, if the first data in the first tuple meets the off-row storage condition, the first data is determined to be large field data, the first data is stored in the off-row data table, and the data first address of the first data in the off-row data table is stored in the data master table. In this way, based on the off-row storage address of the data, the data master table and the off-row data table are directly associated. During the data management process, the data stored in the off-row database can be directly managed based on the off-row storage address recorded in the data master table, thereby improving data storage efficiency in high-concurrency data storage scenarios.
[0011] Optionally, storing the first data in an out-of-row data table includes: slicing the first data to obtain at least one slice data; and storing the at least one slice data in the out-of-row data table.
[0012] Optionally, the at least one slice data includes first slice data and second slice data connected sequentially; storing the at least one slice data in the out-of-row data table includes: storing the second slice data in a first physical area in the out-of-row data table; storing the first slice data and a physical address of the first physical area in a second physical area in the out-of-row data table; storing the first storage address in the first tuple of the data master table includes: storing the physical address of the second physical area in the first tuple of the data master table.
[0013] Optionally, the method also includes: receiving a data update request for the first tuple, the data update request being used to request an update of the first data through second data; if the second data meets the out-of-row storage condition, storing the second data in the out-of-row data table, and obtaining a second storage address of the second data in the out-of-row data table; updating the first storage address in the first tuple of the data main table to the second storage address.
[0014] Optionally, after receiving the data update request for the first tuple, the method further includes: if the second data does not meet the out-of-row storage condition, updating the first tuple in the data master table based on the second data.
[0015] Optionally, after updating the first tuple in the data main table, the method further includes: deleting the first data in the out-of-row data table based on the first storage address.
[0016] It can be seen from this that when the present application uses the second data in the database to update the first data in the first tuple, if the second data meets the off-row storage condition, the second data is stored in the off-row data table, and the off-row storage address of the second data is used in the data master table to update the first storage address of the first data. In this way, based on the off-row storage address of the data, the data master table and the off-row data table are directly associated. During the data management process, the data stored in the off-row database can be directly managed based on the off-row storage address recorded in the data master table, thereby improving data storage efficiency in high-concurrency data storage scenarios.
[0017] Optionally, the method also includes: receiving a data query request for the first tuple, the data query request being used to query the data of a target field among the multiple fields; if the off-row storage address of the target field is stored in the data master table, then based on the off-row storage address, reading the data of the target field from the off-row data table; and outputting the data query result of the first tuple based on the data of the target field.
[0018] Optionally, the data of the target field includes multiple slice data; and outputting the data query result of the first tuple based on the data of the target field includes: splicing the multiple slice data in the reading order of the multiple slice data; and outputting the query result of the first tuple based on the spliced data.
[0019] It can be seen from this that when the present application queries the data of the first tuple in the database, if the data of the target field is stored in the off-row data table, then after obtaining the off-row storage address of the target field in the data master table, based on the off-row storage address, the data of the target field is read from the off-row data table, and the data query result of the first tuple is output. In this way, based on the off-row storage address of the data, the data master table and the off-row data table are directly associated. In the data management process, the data stored in the off-row database can be directly managed based on the off-row storage address recorded in the data master table, which improves the data storage efficiency in high-concurrency data storage scenarios.
[0020] In a second aspect, a data management device is provided, wherein the data management device has the function of implementing the data management method in the first aspect. The data management device includes at least one module, and the at least one module is used to implement the data management method provided in the first aspect.
[0021] In a third aspect, a computer device is provided, comprising a processor and a memory, wherein the memory is configured to store a computer program for executing the data management method provided in the first aspect. The processor is configured to execute the computer program stored in the memory to implement the data management method described in the first aspect.
[0022] Optionally, the computer device may further include a communication bus, which is used to establish a connection between the processor and the memory.
[0023] In a fourth aspect, a computer-readable storage medium is provided, wherein the storage medium stores instructions, and when the instructions are executed on a computer, the computer executes the data management method described in the first aspect.
[0024] In a fifth aspect, a computer program product comprising instructions is provided. When the instructions are executed on a computer, the computer is caused to perform the data management method described in the first aspect. Alternatively, a computer program is provided. When the computer program is executed on a computer, the computer is caused to perform the steps of the data management method described in the first aspect.
[0025] The technical effects obtained in the above-mentioned second, third, fourth and fifth aspects are similar to those obtained by the corresponding technical means in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] FIG1 is a schematic diagram of a data table structure under a TOAST mechanism provided in an embodiment of the present application;
[0027] FIG2 is a schematic diagram of a data storage process under a TOAST mechanism provided in an embodiment of the present application;
[0028] FIG3 is a schematic diagram of a data management scenario provided in an embodiment of the present application;
[0029] FIG4 is a schematic diagram of the architecture of a data management system provided in an embodiment of the present application;
[0030] FIG5 is a schematic diagram of the structure of a computing node provided in an embodiment of the present application;
[0031] FIG6 is a flow chart of a data management method provided in an embodiment of the present application;
[0032] FIG7 is a flow chart of another data management method provided in an embodiment of the present application;
[0033] FIG8 is a flow chart of another data management method provided in an embodiment of the present application;
[0034] FIG9 is a schematic diagram of a data table structure under an ENTOAST mechanism provided in an embodiment of the present application;
[0035] FIG10 is a flowchart of data storage under an ENTOAST mechanism provided in an embodiment of the present application;
[0036] FIG11 is a data update flow chart under an ENTOAST mechanism provided in an embodiment of the present application;
[0037] FIG12 is a data query flow chart under an ENTOAST mechanism provided in an embodiment of the present application;
[0038] FIG13 is a schematic structural diagram of a data management device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0039] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0040] For ease of understanding, before explaining in detail the data management method provided in the embodiment of the present application, the terminology, relevant background and implementation environment involved in the embodiment of the present application are first introduced.
[0041] First, the terms involved in the embodiments of the present application are introduced.
[0042] 1. Database management system (DBMS)
[0043] A DBMS is a large-scale software program that manipulates and manages databases, used to create, use, and maintain them. It provides unified management and control of databases to ensure their security and integrity. Users access data in the database through the DBMS, and database administrators also use it to maintain the database. It can support multiple applications and users using different methods to create, modify, and access the database simultaneously or at different times. Most DBMSs provide data definition language (DDL), data manipulation language (DML), and data query language (DQL) for users to define the database's schema structure and permission constraints, and to implement operations such as appending and deleting data.
[0044] 2. Tuple
[0045] Simply put, in a database, a tuple refers to a record or a row of data, which is the most basic data unit in the database.
[0046] A tuple is essentially a series of bytes. How these bytes are interpreted depends on the database schema. The DBMS translates these bytes into values of the corresponding type based on the table definition. A tuple is divided into two parts: the tuple header and the attribute data. The tuple header contains metadata, such as visibility information and bitmaps. The tuple data generally stores the specific data for each attribute in the order specified in the table definition.
[0047] 3. Page
[0048] In a database, a page is a contiguous block of physical storage space, typically 4KB or 8KB in size. Database management systems use pages to store data and indexes, and each page has a unique identifier called a page ID.
[0049] A page is the most basic storage unit in a database. All data and indexes are stored in a page. When the database needs to read or write data, it loads the page into memory for operation.
[0050] 4. Row identifier (tuple ID, TID)
[0051] In the openGauss database, TID is a special data type that represents the physical location of a row. It consists of two parts: page ID and tuple ID.
[0052] 5. Large fields
[0053] In a database, a large field is a column (also known as an attribute or field) that stores large amounts of data, typically binary data such as text, images, audio, or video. Because large fields exceed the size of common data types, they often require special processing and storage methods.
[0054] 6. Off-row storage
[0055] Off-row storage refers to storing certain rows of data (i.e., tuple data) in a database on a separate storage medium rather than in the main database table. This storage method can reduce database load, reduce data access latency, and improve database performance and scalability.
[0056] Among them, out-of-row storage is usually used to store large binary objects (BLOBs) or large character objects (CLOBs), because the data size of these objects may exceed the limit of the database management system.
[0057] 7. Row Links
[0058] Row chaining refers to splitting the data of a tuple into multiple slices of fixed size in a relational database and connecting them using physical or logical location information, so that the complete data of a tuple can be obtained at one time when reading.
[0059] 8. Data manipulation language (DML)
[0060] DML is used to operate on data in the database, including adding, deleting, and modifying. Common DML statements include insert (INSERT) statements, update (UPDATE) statements, and delete (DELETE) statements.
[0061] 9. Data Query Language (DQL)
[0062] DQL is used to query data in the database and does not modify the data. The most common DQL statement is the SELECT statement.
[0063] Secondly, the relevant background of the embodiments of the present application is introduced.
[0064] To address the storage problem of large field data in the database, openGauss provides an oversized-attribute storage technique (TOAST) to support the storage of large fields in the database.
[0065] When a user stores large field data exceeding 2KB, the TOAST mechanism will attempt to store the large field data off-line, that is, store the large field data in the TOAST table. When storing large field data off-line, a 32-bit unique identifier is first applied for the large field data as a chunk ID (Chunk ID). The Chunk ID is the index value for the off-line storage location. The Chunk ID is stored in the tuple to which the large field data belongs in the data master table. At the same time, the Chunk ID and the off-line storage address indicating the large field data (i.e., the physical location in the TOAST table) are recorded in the TOAST index table. When accessing large field data stored off-line, the Chunk ID corresponding to the large field data is first obtained from the data master table, and then the off-line storage address corresponding to the Chunk ID is determined from the TOAST index table. The large field data is then read from the TOAST table using the off-line storage address.
[0066] Referring to Figure 1, for large field data with a data length exceeding 2KB, assuming that the value of the location index (i.e., Chunk ID) stored in the data master table is 16630, then according to the Chunk ID in the data master table, the storage location (denoted as ctid) of the large field data in the TOAST table is determined from the TOAST index table as (0, 1) and (0, 2). Based on the ctid value, the slice data No. 0 corresponding to the large field data is read from the first row of page 0 in the TOAST table, and the slice data No. 1 corresponding to the large field data is read from the second row of page 0 to obtain the complete large field data.
[0067] In the TOAST mechanism, the data master table is used to store tuple visibility information for visibility judgment during data DML and DQL processes. The data master table also stores the chunk ID corresponding to at least one large field data in the tuple; the TOAST index table is used to store the correspondence between the chunk ID and the off-row storage address of the large field data; and the TOAST table is used to store large field data.
[0068] As shown in Figure 2, the TOAST mechanism is used to implement off-row storage of large field data, which includes the following steps:
[0069] (1) Determine whether the length of the tuple exceeds the threshold.
[0070] If it exceeds a threshold (for example, 2KB), the following step (2) is executed; if it does not exceed the threshold, the tuple can be inserted into the data master table.
[0071] (2) Try to perform data compression processing on the target field in the tuple, and again determine whether the length of the tuple after the data compression processing exceeds a threshold (for example, 2KB).
[0072] If the threshold is exceeded, the following step (3) is executed; if the threshold is not exceeded, the compressed tuple is inserted into the data master table.
[0073] (3) If the target field still has a data length of more than 2KB after data compression, it is determined to be large field data. A unique identifier of type uint32 is applied for each large field data. That is, the identifier serves as the Chunk ID of the large field data and is used to index the off-row storage location.
[0074] Specifically, the implementation process of applying for a Chunk ID for large field data is as follows: based on the TOAST index table, first add a global OID generation lock (referred to as an OID lock) to the identifier serial number, and scan the latest identifier as the Chunk ID of the large field data, then add 1 to the identifier serial number and release the OID lock. Next, to ensure the unique validity of the Chunk ID, it is necessary to check whether the currently applied Chunk ID is repeated in the identifier serial number. If the Chunk ID is repeated, add 1 to the Chunk ID value and apply again. If there is an integer overflow when applying for the Chunk ID, continue to apply from the first valid identifier. If no valid Chunk ID is applied for, continue to wait; if a valid Chunk ID is applied for, execute the following step (4).
[0075] (4) The large field data is sliced according to a fixed size to obtain at least one slice data, and then traversed sequentially to insert at least one slice data into the TOAST table, and based on the Chunk ID of the large field, the physical address of each slice data is stored in the TOAST index table to update the TOAST index table.
[0076] (5) Based on the updated TOAST index table, the Chunk ID of the large field data is inserted into the corresponding tuple of the data master table, ending the out-of-row storage process of the large field data in the tuple.
[0077] However, when storing large field data in a database based on the TOAST mechanism, there is at least one of the following problems.
[0078] (1) The off-row storage of large field data relies on a globally incremented identifier serial number (i.e., Chunk ID), and uses the Chunk ID as the primary key index of the off-row storage address. As a result, in high-concurrency scenarios, operations such as applying for Chunk IDs, checking the validity of Chunk IDs (i.e., checking whether the applied Chunk IDs are repeated in the identifier serial number), concurrently inserting Chunk IDs into the TOAST index table in increasing order, and checking the uniqueness of index values in the TOAST index table (i.e., checking whether there are duplicate Chunk IDs in the TOAST index table) may introduce a large number of conflicting operations, which seriously degrades the concurrent performance of the off-row storage of large field data.
[0079] (2) As the number of chunk IDs in the TOAST index table increases, the time consumed in checking the uniqueness of the chunk ID becomes longer and longer, resulting in a significant performance drop.
[0080] (3) Since the total number of globally incremented identifier serial numbers is 4B, there can only be a maximum of 2 32 Identifier serial number, that is, when applying for Chunk ID, 2 32 There is a possibility of rollback of the identifier serial number. 32 After the identifier sequence number is rolled over, it is necessary to re-apply from the first identifier sequence number to use the applied identifier as the Chunk ID. Therefore, once the identifier sequence number rolls over, it is necessary to wait when applying for the Chunk ID, which seriously affects user services.
[0081] (4) The pointer used to indicate the off-row storage location in the current data master table occupies 18 bytes. A single row (i.e., a single tuple) in the data master table cannot exceed the page size (usually 8KB). Therefore, the number of columns in a single tuple in the data master table (i.e., the number of fields or attributes contained in the tuple) is limited to (page size - page metadata) / 18 bytes ≈ 440 columns. In other words, the upper limit of the number of fields or attributes contained in a tuple in the data master table is less than 450, resulting in poor scalability of single tuple attributes.
[0082] Based on at least one of the above problems, an embodiment of the present application provides a data management method, which implements off-row storage, update, query and other operations of large field data through an enhanced over-size attribute storage mechanism (enhanced TOAST, abbreviated as ENTOAST), and directly stores the off-row storage address of the large field data in the data master table of the database. Based on the off-row storage address, the large field data is directly read from the off-row data table, overcoming the problems caused by referencing Chunk ID for large field data management and improving data management efficiency.
[0083] Next, the application scenarios, system architecture, and implementation environment of the embodiments of the present application are introduced.
[0084] A database can be implemented using one or more compute nodes. Figure 3 shows a schematic diagram of a database-based data management scenario. The database client can include multiple clients, each of which can initiate data management requests against the database. For example, if the database is deployed on a single data node, the database includes software modules such as the port listening daemon in the data kernel, the SQL engine, and the storage engine, as well as hardware modules such as shared memory and disks. The storage engine directly interfaces with the shared memory.
[0085] Among them, the storage engine is an important component of the database management system, and its main function is to manage the storage and retrieval of data. In the embodiment of the present application, the storage engine can provide large field storage service and large field read service, and the large field storage service and large field read service use off-line storage combined with row chaining to replace the traditional TOAST index table and TOAST table storage method.
[0086] Referring to FIG4 , when implementing the data management method provided in an embodiment of the present application, the SQL engine in the data kernel can process the natural language-based data request sent by the client into machine language that can be recognized and executed by the machine. For example, in response to a data operation request initiated by the client, the SQL engine sends a data operation instruction (i.e., a DML statement) to the storage engine, and the storage engine processes the requested large field data in the shared memory in response to the data storage instruction; in response to a data query request initiated by the client, the SQL engine sends a data query instruction (i.e., a DQL statement) to the storage engine, and the storage engine reads the requested large field data from the shared memory in response to the data query instruction.
[0087] It should be understood that the shared memory can transfer the stored data to the disk, and the embodiment of the present application does not limit the timing of data transfer between the shared memory and the disk.
[0088] It should be noted that Figures 3 and 4 only take the database deployment on a single data node as an example. In actual applications, the database can also be deployed on multiple data nodes. The software modules and hardware modules in each data node are deployed in the same way. The embodiment of the present application does not limit the number of data nodes deployed on the database.
[0089] The data management scenario described in the embodiment of the present application is intended to more clearly illustrate the technical solution of the embodiment of the present application, and does not constitute a limitation on the technical solution provided in the embodiment of the present application. Ordinary technicians in this field can know that with the emergence of new databases, the technical solution provided in the embodiment of the present application is also applicable to similar technical problems.
[0090] Please refer to Figure 5, which is a schematic diagram illustrating the structure of a data node according to an embodiment of the present application. The data node can be any computer device with data read and write functions, such as a server. The data node includes at least one processor 501, a communication bus 502, a memory 503, and at least one communication interface 504.
[0091] The processor 501 may be a general-purpose central processing unit (CPU), a network processor (NP), a microprocessor, or one or more integrated circuits for implementing the solution of the present application, such as an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0092] Communication bus 502 is used to transmit information between the above components. Communication bus 502 can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used in Figure 5, but this does not mean that there is only one bus or one type of bus.
[0093] The memory 503 may be a read-only memory (ROM), a random access memory (RAM), an electrically erasable programmable read-only memory (EEPROM), an optical disc (including a compact disc read-only memory (CD-ROM), a compact disc, a laser disc, a digital versatile disc, a Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 503 may exist independently and be connected to the processor 501 via the communication bus 502. The memory 503 may also be integrated with the processor 501.
[0094] The communication interface 504 uses any transceiver-like device for communicating with other devices or communication networks. The communication interface 504 includes a wired communication interface and may also include a wireless communication interface. The wired communication interface may be, for example, an Ethernet interface. The Ethernet interface may be an optical interface, an electrical interface, or a combination thereof. The wireless communication interface may be a wireless local area network (WLAN) interface, a cellular network communication interface, or a combination thereof.
[0095] As an example, the processor 501 may include one or more CPUs, such as CPU0 and CPU1 shown in FIG. 5 .
[0096] As an example, a data node may include multiple processors, such as processor 501 and processor 505 shown in Figure 5. Each of these processors may be a single-core processor or a multi-core processor. A processor herein may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0097] In some embodiments, the data node may also include an output device and an input device. The output device communicates with the processor 501 and can display information in a variety of ways. For example, the output device may be a liquid crystal display (LCD), a light emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector. The input device communicates with the processor 501 and can receive user input in a variety of ways. For example, the input device may be a mouse, keyboard, touch screen device, or sensor device.
[0098] In some embodiments, the memory 503 is used to store program code 510 for executing the solution of the present application, and the processor 501 can execute the program code 510 stored in the memory 503. The program code 510 may include one or more software modules, and the data node can implement the data management method provided in the embodiment of Figure 6 below through the processor 501 and the program code 510 in the memory 503.
[0099] Next, the data management method provided in the embodiment of the present application is described in detail.
[0100] FIG6 is a flow chart of a data management method provided in an embodiment of the present application, which can be applied to the above-mentioned data node. Referring to FIG6 , the method includes the following steps.
[0101] Step 601: Receive a data storage request for a first tuple, where the first tuple includes data of multiple fields.
[0102] Among them, the first tuple is any tuple requested to be stored in the database. Among the multiple fields included in the first tuple, the data length of some fields may exceed the data length threshold, and the data length of some fields may not exceed the data length threshold. For data that exceeds the data length threshold, it is large field data and cannot be directly stored in the main data table of the database. It needs to be compressed or stored out of row.
[0103] The data length threshold can be set based on the page size of the data master table and / or the off-row data table. Since tuples cannot be stored across pages, when the page size is 8KB and one page stores 4 tuples, the data length threshold can be set to 2KB. Of course, in actual applications, the data length threshold can also be set to other values, such as 4KB, etc., and this embodiment of the present application does not limit this.
[0104] In some embodiments, after receiving a data request for a first tuple, a determination is first made based on the tuple length of the first tuple to determine whether the first tuple meets the off-row storage condition. If the off-row storage condition is met, consideration is given to storing some or all of the fields of the first tuple off-row. If the off-row storage condition is not met, the first tuple is directly stored in the data master table.
[0105] The tuple length refers to the total length of the data in each field in the tuple, that is, the total number of bytes; the out-of-row storage condition is that the data length is greater than the data length threshold.
[0106] Optionally, when the tuple length of the first tuple is greater than the data length threshold, that is, the off-row storage condition is met, data compression processing can be performed on the large field data in the first tuple whose data length exceeds the data length threshold, and based on the data compression result, the tuple length of the first tuple is recalculated. If the tuple length of the first tuple after data compression processing is greater than the data length threshold, the large field data in the first tuple is stored off-row; if the tuple length of the first tuple after data compression processing is less than or equal to the data length threshold, the compressed large field data in the first tuple is stored in the data master table.
[0107] Among the multiple fields of the first tuple, some fields support compression. Therefore, if the first tuple meets the out-of-row storage condition, the data in the fields that support compression in the first tuple can be compressed first. If the first tuple still meets the out-of-row storage condition after the data compression process, the large field data in the first tuple will be stored out-of-row.
[0108] It should be noted that when determining whether to store the first tuple out-of-row based on its tuple length, if the first tuple contains large field data, the large field data can be directly stored out-of-row. To reduce the physical area occupied by the large field data when stored out-of-row, the large field data can be compressed first. Based on the compression result, the compressed data can also be stored in the out-of-row data table.
[0109] That is, the first tuple after compression meets the off-line storage condition, indicating that the total length of the data of the compressed field and the uncompressed field is greater than 2KB. At this time, for the compressed field, its data length may exceed 2KB or may not exceed 2KB, but because the tuple length is greater than the data length threshold, the large field data after compression in the first tuple can be stored off-line.
[0110] As an example, if the first tuple includes field 1, field 2, field 3, field 4, and field 5, and if field 2, field 3, and field 5, and the tuple length of the first tuple exceeds the data length threshold. At this time, the data of field 2, field 3, and field 5 that support compression processing are first compressed. If, after the data compression processing, the data length of field 2 is greater than the data length threshold, the data length of field 3 and field 5 is less than the data length threshold, and the tuple length of the first tuple is greater than the data length threshold, then the compressed data of field 2, field 3, and field 5 are stored out-of-row. If, after the data compression processing, the data lengths of field 2, field 3, and field 5 are all less than the data length threshold, at this time, the tuple length of the first tuple may be less than the data length threshold. Therefore, the first tuple can be stored in the data master table, where field 1 and field 4 store the original data, and field 2, field 3, and field 5 store the compressed data; of course, the compressed data of field 2, field 3, and field 5 can also be directly stored out-of-row.
[0111] Step 602: If the first data meets the out-of-row storage condition, the first data is stored in the out-of-row data table, where the first data is data of any field in the plurality of fields.
[0112] In an embodiment of the present application, the off-row data table is an ENTOAST table, which is used to store large field data that needs to be stored off-row in the data main table, as well as the correspondence between the large field data and the physical address.
[0113] In other words, if the first data is original data whose data length is greater than the data length threshold, or compressed data whose data length after data compression processing is greater than the data length threshold, or large field data in the first tuple when the tuple length of the first tuple after data compression processing is greater than the data length threshold.
[0114] In a possible implementation, the implementation process of storing the first data in the out-of-row data table is: slicing the first data to obtain at least one slice data; and storing the at least one slice data in the out-of-row data table.
[0115] The first data may be sliced according to a slice length that is less than or equal to the data length threshold, for example, the slice length may be 1KB, 2KB, etc., which is not limited in this embodiment of the present application.
[0116] If at least one slice data includes first slice data and second slice data connected sequentially, the implementation process of storing at least one slice data in the out-of-row data table can be: storing the second slice data in the first physical area of the out-of-row data table; storing the first slice data and the physical address of the first physical area in the second physical area of the out-of-row data table.
[0117] It should be noted that "sequential connection" means that the first slice data is located before the second slice data in the first data. However, when storing data in the out-of-row data table, this embodiment of the application stores the data in a "back-to-front" order, storing the second slice data first and the first slice data later. In other words, the order in which the first and second slice data are connected is the opposite of their storage order in the out-of-row data table.
[0118] In the out-of-row data table, each physical area corresponds to a physical address. The physical address of the first slice data is stored in the physical area where the next slice data is stored. In this way, the physical addresses of multiple slice data are concatenated, and the physical address stored in the physical area of one slice data can be used to index the next slice data.
[0119] Optionally, when storing the first data in the out-of-row data table, not only the first data itself can be stored, but also the data association information corresponding to the first data can be stored in the second physical area where the first slice data is located, such as the original length of the field data (rawsize), the compressed length of the field data (textsize), etc. The embodiment of the present application does not impose any restrictions on this.
[0120] Step 603: Obtain a first storage address of the first data in the out-of-row data table.
[0121] Since the physical address of the first physical area is recorded in the second physical area, after determining the second physical area where the first slice data is located in the out-of-row data table, the physical address of the first physical area can be read from the physical area, thereby continuing to read the second slice data from the first physical area.
[0122] Therefore, if it is necessary to obtain at least one slice data stored in the out-of-row data table of the first data, it is only necessary to determine the storage position of the first slice data in the out-of-row data table of the first data.
[0123] In one possible implementation, step 603 may be implemented by obtaining the physical address of the first slice of the first data in the out-of-row data table to obtain the first storage address. In other words, the first storage address is the first address of the first data in the out-of-row data table.
[0124] Step 604: Store the first storage address in the first tuple of the data main table.
[0125] That is, the data starting address of the first data in the off-row data table is stored in the first tuple of the data main table.
[0126] Taking the example that at least one slice data of the first data includes first slice data and second slice data connected sequentially, the implementation process of step 604 is: storing the physical address of the second physical area storing the first data slice in the first tuple of the data master table.
[0127] Thus, it can be seen that when the embodiment of the present application stores the first tuple in the database, if the first data in the first tuple meets the off-row storage condition, the first data is determined to be large field data, the first data is stored in the off-row data table, and the data starting address of the first data in the off-row data table is stored in the data master table. In this way, based on the off-row storage address of the data, the data master table and the off-row data table are directly associated. During the data management process, the data stored in the off-row database can be directly managed based on the off-row storage address recorded in the data master table, thereby improving data storage efficiency in high-concurrency data storage scenarios.
[0128] After completing the data storage operation for the first tuple in the data master table and the off-row data table based on steps 601 to 604 above, at least one large field data in the first tuple can also be updated based on the database. Based on this, referring to FIG. 7 , the data management method provided in this embodiment of the present application further includes the following steps.
[0129] Step 605: Receive a data update request for the first tuple, where the data update request is used to request to update the first data using the second data.
[0130] In one possible implementation, a data update request for the first tuple carries the tuple ID of the first tuple and the updated data for each field in the first tuple, indicating the updated data for each field, and updates the first tuple in the data master table. In other words, a request is made to update the entire first tuple.
[0131] In this case, the first tuple can be updated based on the updated data of each field; or the above steps 601-604 can be used in the data master table to regenerate a row of data to store the updated first tuple and mark the first tuple before the update for deletion.
[0132] In another possible implementation, the data update request carries the tuple ID of the first tuple, the field identifier of the first field, and the second data, indicating that the data of the first field in the first tuple should be updated using the second data. That is, the request is to update the data of a single field in the first tuple.
[0133] In this case, the first data in the first tuple may be updated based on the second data, for example, the first storage address of the first data in the data master table may be updated.
[0134] The embodiment of the present application does not limit the information carried in the data update request, which may include data of all fields after the first tuple is updated, or may only include data of at least one field that needs to be updated.
[0135] Step 606: If the second data meets the out-of-row storage condition, the second data is stored in the out-of-row data table, and a second storage address of the second data in the out-of-row data table is obtained.
[0136] Since the data length of the second data may be greater than the data length threshold, or may be less than or equal to the data length threshold, before using the second data to update the first data, it is necessary to determine whether the second data can be directly stored in the data master table.
[0137] As described above, the condition for off-row storage is that the data length is greater than the data length threshold. Therefore, when the data length of the second data is greater than the data length threshold, it indicates that the second data is large field data and cannot be stored directly in the data master table. In this case, the second data needs to be stored off-row. When the data length of the second data is less than or equal to the data length threshold, it indicates that the second data can be normally stored in the database. In this case, the second data is stored directly in the data master table.
[0138] Similarly, the second storage address of the second data in the out-of-row data table is the first data address of the second data in the out-of-row data table. In other words, the second storage address is the physical address of the physical area corresponding to the first slice data corresponding to the second data (that is, the last slice data to perform an out-of-row storage operation) in the out-of-row data table.
[0139] Step 607: Update the first storage address in the first tuple of the data master table to the second storage address.
[0140] When the second data meets the out-of-row storage condition, after determining the second storage address of the second data in the out-of-row data table, delete the first storage address of the first data in the field corresponding to the first data in the data main table and write the second storage address.
[0141] In some embodiments, if the second data does not meet the out-of-row storage condition, the first tuple in the data master table is updated based on the second data.
[0142] That is, when the data length of the second data is less than the data length threshold, the second data can be directly stored in the data master table without the need for out-of-row storage.
[0143] In one possible implementation, when the first tuple in the data master table is updated based on the second data, a second tuple can be added to the data master table based on the first tuple and the second data. The second tuple includes the second data and is the same as the data of other fields in the first tuple. The tuple including the first data is deleted from the data master table to implement the update of the first tuple.
[0144] In another possible implementation, when the first tuple in the data master table is updated based on the second data, the first storage address of the first data may be deleted, and the second data may be written to the field corresponding to the first data.
[0145] In some embodiments, when the first data in the above-mentioned first tuple does not meet the off-row storage condition and the first data is stored in the data master table, the process of using the second data to update the first data is as follows: if the data length of the second data is greater than the data length of the first data, the second data cannot directly update the first data at the original field position. At this time, based on the first tuple and the second data, a second tuple can be added to the data master table. The second tuple includes the second data and is the same as the data of other fields in the first tuple; the first tuple including the first data is deleted from the data master table to update the first tuple. If the data length of the second data is less than or equal to the data length of the first data, the second data is used to directly update the first data at the original field position (data overwriting) to update the first tuple.
[0146] In addition, when the first data is stored in the off-row data table, after the first tuple is updated in the data main table using the second data, the first data needs to be deleted in the off-row data table based on the first storage address of the first data stored in the data main table.
[0147] Thus, it can be seen that in the embodiment of the present application, when the second data is used in the database to update the first data in the first tuple, if the second data meets the off-row storage condition, the second data is stored in the off-row data table, and the off-row storage address of the second data is used in the data master table to update the first storage address of the first data. In this way, based on the off-row storage address of the data, the data master table and the off-row data table are directly associated. During the data management process, the data stored in the off-row database can be directly managed based on the off-row storage address recorded in the data master table, thereby improving data storage efficiency in high-concurrency data storage scenarios.
[0148] After the first tuple is stored in the data master table based on steps 601-604, a data query request for at least one field in the first tuple can be initiated based on the database to obtain data for at least one field in the first tuple from the database. Based on this, referring to FIG8 , the data management method provided in this embodiment of the application further includes the following steps.
[0149] Step 608: Receive a data query request for the first tuple, where the data query request is used to query data of a target field among multiple fields.
[0150] The target fields may be part of the fields in the first tuple, or may be all of the fields in the first tuple. The embodiment of the present application does not limit the number of target fields.
[0151] Step 609: If the off-row storage address of the target field is stored in the data master table, the data of the target field is read from the off-row data table based on the off-row storage address.
[0152] As explained above, when a field's data is stored in an off-row data table, the master table only stores the data's off-row storage address at that field's location. Therefore, if the master table stores the target field's off-row storage address, it indicates that the target field's data is stored in the off-row data table. You need to retrieve the target field's data from the off-row data table based on that off-row storage address.
[0153] In one possible implementation, the implementation process of reading the data of the target field from the off-row data table based on the off-row storage address is as follows: based on the off-row storage address, obtain the first slice data from the corresponding physical area in the off-row data table. If a physical address is still stored in the physical area indicated by the off-row storage address, continue to read the next slice data based on the physical address, and so on, until no physical address is stored in the physical area of the slice data, then obtain multiple slice data corresponding to the target field. If no physical address is stored in the physical area indicated by the off-row storage address, it means that the data of the target field was not sliced when stored off-row, and the first slice data is the complete data of the target field.
[0154] Step 610: Based on the data of the target field, output the data query result of the first tuple.
[0155] When the data of the target field includes multiple slice data, in a possible implementation, the implementation process of step 610 is: splicing the multiple slice data according to the reading order of the multiple slice data; and outputting the query result of the first tuple based on the spliced data.
[0156] After splicing multiple slice data, the complete data of the target field can be obtained. At this time, the query result of the first tuple can be output based on the complete data of the target field.
[0157] Optionally, the query result of the first tuple may include all the data of the first tuple, or may only include the data of the target field in the first tuple, which is not limited in this embodiment of the present application.
[0158] Thus, it can be seen that when the embodiment of the present application queries the data of the first tuple in the database, if the data of the target field is stored in the off-row data table, after obtaining the off-row storage address of the target field in the data master table, the data of the target field is read from the off-row data table based on the off-row storage address, and the data query result of the first tuple is output. In this way, based on the off-row storage address of the data, the data master table and the off-row data table are directly associated. During the data management process, the data stored in the off-row database can be directly managed based on the off-row storage address recorded in the data master table, thereby improving data storage efficiency in high-concurrency data storage scenarios.
[0159] In some embodiments, after receiving a data deletion request for the first tuple, the data of each field in the first tuple can be determined in the data main table and the off-row data table according to the above step 609, and then the first tuple can be deleted in the data main table and the row data table.
[0160] That is, based on the above embodiment, the tuple deletion operation can be decomposed into a data query operation and a data deletion operation. After determining the specific position of the first tuple in the main data table and the off-row data table, the deletion operation can be performed on the data of the corresponding first tuple. The embodiment of this application will not be elaborated here.
[0161] Based on the above method embodiments, next, in conjunction with Figures 9-12, the specific process of data management based on the ENTOAST mechanism in the embodiment of the present application is described in detail.
[0162] The data management method of the embodiment of the present application involves two data tables, namely the data master table in the database and the off-row data table (i.e., ENTOAST table). In the data master table, based on the structure of the target field, if the field data of the target field is large field data larger than 2KB, then when storing the large field data in the ENTOAST table, the large field data is first sliced to obtain at least one slice data, and then the slice data is stored in the ENTOAST table in a "from back to front" storage order. After the last slice data is stored, the row identifier (TID) of the last slice data in the ENTOAST table is stored in the corresponding tuple in the data master table.
[0163] That is, the off-row storage address recorded in the data master table in the embodiment of the present application is: the TID (page ID, row ID) of the field data of the target field in the off-row data table.
[0164] As an example, referring to FIG9 , it is assumed that the large field data of the target field is sliced to obtain the first slice data and the second slice data connected in sequence, and it is determined that the first row and the second row of page 0 of the ENTOAST table are free positions, then the second slice data is first stored in the second row, and then the physical addresses (0, 2) of the first slice data and the second row are stored in the first row, and the first address of the large field data in the off-row data table, that is, the physical address (0, 1) of the first slice data in the off-row data table, is returned to the data main table, and the off-row storage address stored at the target field position in the data main table is (0, 1).
[0165] Optionally, when storing the large field data of the target field in the ENTOAST table, data association information of the large field data can also be stored in the storage space of the first slice data, such as the original length of the field data (rawsize), the compressed length of the field data (textsize), etc. The embodiment of the present application does not impose any restrictions on this.
[0166] 9 , in the storage area of the first slice data, the meta area stores the data association information of the target field's large field data. Thus, after acquiring the last first slice data, the data association information of the target field can be determined.
[0167] During the data reading process, based on the off-row storage address TID=(0,1) stored in the data master table, the first slice data is obtained from the first row of page 0 in the ENTOAST table, and then based on the physical address (0,2) stored in the first row, the second slice data is read from the second row of page 0. Since the physical address stored in the second row is (0,0), the second slice data stored in the second row is the last slice data. In this way, based on the off-row storage address stored in the data master table, the two slice data corresponding to the target field can be read from the ENTOAST table. After splicing these two slice data, the complete data of the target field can be obtained.
[0168] Referring to Figure 10, the large field data storage process based on the ENTOAST mechanism provided by the embodiment of the present application is: listen to the data insertion message. If it is listened that a tuple needs to be stored in the database, first determine whether the tuple length of the tuple is greater than the data length threshold. If the tuple length is less than or equal to the data length threshold, the tuple is directly inserted into the data main table; if the tuple length is greater than the data length threshold, the target field that supports compression in the tuple is compressed, and it is determined whether the tuple length of the tuple is greater than the data length threshold after the data compression processing. If the tuple length after the data compression processing is less than or equal to the data length threshold, the tuple after the data compression processing is directly inserted into the data main table; if the tuple length after the data compression processing is greater than the data length threshold, the large field data of the target field that is greater than the data length threshold after the data compression processing is stored out of row.
[0169] When performing off-row storage, the large field data is sliced, and at least one slice data obtained after the slicing process is stored in the ENTOAST table, and the physical addresses of the physical areas of each slice data are used for concatenation. When all slice data of the large field data are stored in the ENTOAST table, the data starting address of the large field data in the off-row data table (that is, the physical address of the last stored slice data) is returned, and the data starting address is stored in the corresponding tuple of the data main table, ending the data storage process.
[0170] Optionally, when inserting slice data of the target data into the ENTOAST table in sequence, tuple data marked as deleted in the traversed pages can be deleted during the process of traversing the storage locations, and the physical space occupied by the tuple data can be released.
[0171] Referring to Figure 11, the large field data update process based on the ENTOAST mechanism provided by the embodiment of the present application is: monitor data update messages. If it is monitored that the new version tuple needs to be used to update the old version tuple in the database, the visibility of the old version tuple is first judged in the data master table. If the old version tuple is visible, the data in the old version tuple is updated based on the new version tuple; if the old version tuple is not visible, it means that the old version tuple is not stored in the data master table, and the data update process is directly ended.
[0172] If the old version tuple exists in the data master table, the new version tuple's tuple length is determined to be greater than the data length threshold. If the tuple length is less than or equal to the data length threshold, the old version tuple is directly updated with the new version tuple. If the tuple length is greater than the data length threshold, the new version tuple is sliced and stored out-of-row. Then, based on the out-of-row storage address, the TID stored in the old version tuple is updated in the data master table.
[0173] Furthermore, after storing the off-row storage address of the new version tuple in the data master table, it is necessary to mark the data corresponding to the TID in the off-row data table as invalid based on the TID of the old version tuple, thereby ending the data update process.
[0174] It should be understood that this is only an example of updating the data of the entire tuple. As explained above, when updating a tuple, the data of all fields can be updated, or only the data of at least one field can be updated. The update logic is similar.
[0175] Referring to Figure 12, the large field data query process based on the ENTOAST mechanism provided by the embodiment of the present application is: listen to the data query request. If it is listened to that the data of a certain tuple needs to be queried in the database, first determine whether the tuple is visible in the database main table. If the tuple is visible, then based on the data main table and the off-row data table, obtain the data of the tuple; if the tuple is not visible, it means that the tuple is not stored in the data main table, and then end the data query process.
[0176] If the tuple exists in the main data table and the target field in the tuple has an off-row storage address TID, at least one slice of data is retrieved from the off-row data table based on the off-row storage address TID. Then, the at least one slice of data is spliced. Based on the spliced data, the data query result for the tuple is output, ending the data query process.
[0177] Optionally, when obtaining slice data of the target field in the ENTOAST table, tuple data marked as deleted in the traversed pages may be deleted, and the physical space occupied by the tuple data may be released.
[0178] It should be understood that this is only an example of querying the data of the entire tuple. As explained above, when querying a tuple, you can query to obtain data of all fields, or you can query only data of at least one field. The query logic is similar.
[0179] To sum up, the data management method provided in the embodiment of the present application directly associates the data main table and the off-row data table based on the off-row storage address of the data. During the data management process, for a certain field in the tuple, the data stored in the off-row database can be directly managed based on the off-row storage address recorded in the data main table, thereby improving data management efficiency in high-concurrency data management scenarios.
[0180] Figure 13 is a schematic diagram of the structure of a data management device provided in an embodiment of the present application. The data management device can be implemented as part or all of a data node using software, hardware, or a combination of both. Referring to Figure 13 , the data management device 1300 includes a request receiving module 1301, a data storage module 1302, an address acquisition module 1303, and an address storage module 1304.
[0181] A request receiving module 1301 is configured to receive a data storage request for a first tuple, where the first tuple includes data of multiple fields;
[0182] A data storage module 1302 is configured to store the first data in an off-row data table if the first data meets an off-row storage condition, wherein the first data is data of any field in a plurality of fields;
[0183] An address acquisition module 1303 is configured to acquire a first storage address of the first data in the out-of-row data table;
[0184] The address storage module 1304 is configured to store the first storage address in the first tuple of the data main table.
[0185] Optionally, the data storage module 1302 includes:
[0186] a data slicing unit, configured to perform slicing processing on the first data to obtain at least one slice data;
[0187] The data storage unit is used to store at least one slice data in an off-row data table.
[0188] Optionally, the at least one slice data includes first slice data and second slice data connected sequentially; the data storage unit is specifically configured to:
[0189] storing the second slice data in a first physical area in the out-of-row data table;
[0190] storing the first slice data and the physical address of the first physical area in the second physical area of the out-of-row data table;
[0191] The address storage module 1303 is specifically used for:
[0192] The physical address of the second physical area is stored in the first tuple of the data master table.
[0193] Optionally, in the data management device 1300:
[0194] The request receiving module 1301 is further configured to receive a data update request for the first tuple, where the data update request is configured to request updating the first data using the second data;
[0195] The data storage module 1302 is further configured to store the second data in the out-of-row data table if the second data meets the out-of-row storage condition, and obtain a second storage address of the second data in the out-of-row data table;
[0196] The address storage module 1303 is further configured to update the first storage address in the first tuple of the data master table to a second storage address.
[0197] Optionally, after receiving the data update request of the first tuple, the data management apparatus 1300 further includes:
[0198] The data update module is configured to update the first tuple in the data main table based on the second data if the second data does not meet the out-of-row storage condition.
[0199] Optionally, after updating the first tuple in the data master table, the data management apparatus 1300 further includes:
[0200] The data updating module is further configured to delete the first data in the out-of-row data table based on the first storage address.
[0201] Optionally, the data management device 1300 further includes:
[0202] The request receiving module is further configured to receive a data query request of the first tuple, where the data query request is used to query data of a target field among the multiple fields;
[0203] A data reading module is used to read the data of the target field from the off-row data table based on the off-row storage address if the off-row storage address of the target field is stored in the data main table;
[0204] The result output module is further used to output the data query result of the first tuple based on the data of the target field.
[0205] Optionally, the data of the target field includes a plurality of slice data; the result output module includes:
[0206] A data splicing unit, configured to splice the plurality of slice data according to the order in which the plurality of slice data are read;
[0207] The result output unit is used to output the query result of the first tuple based on the spliced data.
[0208] In an embodiment of the present application, the data management device directly associates the data master table and the off-row data table based on the storage address of the data in the off-row data table. During the data management process, the data in the off-row database can be directly managed based on the storage address recorded in the data master table, thereby improving data management efficiency in high-concurrency data management scenarios.
[0209] It should be noted that the data management device provided in the above embodiment, when managing data stored in a database, is illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the data management device provided in the above embodiment and the data management method embodiment are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0210] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, or a magnetic tape), an optical medium (e.g., a digital versatile disc (DVD)), or a semiconductor medium (e.g., a solid state disk (SSD)). It is worth noting that the computer-readable storage medium mentioned in the embodiments of the present application may be a non-volatile storage medium, in other words, a non-transient storage medium.
[0211] It should be understood that the "plurality" mentioned herein refers to two or more. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in order to facilitate the clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit them to be different.
[0212] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in the embodiments of this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions.
[0213] The above description is an embodiment provided for this application and is not intended to limit this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this application should be included in the scope of protection of this application.
Claims
1. A data management method, characterized in that: The method comprises: receiving a data storage request for a first tuple, the first tuple including data of a plurality of fields; If the first data meets the out-of-row storage condition, the first data is stored in the out-of-row data table, and the first data is data of any field in the multiple fields; Obtain a first storage address of the first data in the out-of-row data table; The first storage address is stored in the first tuple of the data master table.
2. The method according to claim 1, characterized in that The storing the first data into an out-of-row data table includes: Performing slicing processing on the first data to obtain at least one slice data; The at least one slice data is stored in the out-of-row data table.
3. The method according to claim 2, characterized in that The at least one slice data includes first slice data and second slice data connected in sequence; The storing the at least one slice data into the out-of-row data table comprises: storing the second slice data in a first physical area in the out-of-row data table; storing the first slice data and a physical address of the first physical area in a second physical area in the out-of-row data table; The storing the first storage address in the first tuple of the data main table includes: The physical address of the second physical area is stored in the first tuple of the data master table.
4. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: receiving a data update request for the first tuple, wherein the data update request is used to request to update the first data by using second data; If the second data meets the out-of-row storage condition, the second data is stored in the out-of-row data table, and a second storage address of the second data in the out-of-row data table is obtained; The first storage address in the first tuple of the data master table is updated to the second storage address.
5. The method according to any one of claims 1 to 3, characterized in that: After receiving the data update request of the first tuple, the method further includes: If the second data does not satisfy the out-of-row storage condition, the first tuple in the data master table is updated based on the second data.
6. The method according to claim 4 or 5, characterized in that After updating the first tuple in the data master table, the method further includes: Based on the first storage address, the first data in the out-of-row data table is deleted.
7. The method according to any one of claims 1 to 6, characterized in that: The method further comprises: receiving a data query request for the first tuple, wherein the data query request is used to query data of a target field among the multiple fields; If the out-of-row storage address of the target field is stored in the data master table, then based on the out-of-row storage address, the data of the target field is read from the out-of-row data table; Based on the data of the target field, a data query result of the first tuple is output.
8. The method according to claim 7, characterized in that The data of the target field includes a plurality of slice data; and outputting the data query result of the first tuple based on the data of the target field includes: splicing the plurality of slice data according to the reading order of the plurality of slice data; Based on the concatenated data, the query result of the first tuple is output.
9. A data management device, characterized in that: The device comprises: A request receiving module, configured to receive a data storage request for a first tuple, wherein the first tuple includes data of a plurality of fields; A data storage module, configured to store the first data in an out-of-row data table if the first data meets an out-of-row storage condition, wherein the first data is data of any field in the multiple fields; An address acquisition module, used for acquiring a first storage address of the first data in the out-of-row data table; The address storage module is used to store the first storage address in the first tuple of the data main table.
10. The device according to claim 9, characterized in that The data storage module comprises: A data slicing unit, configured to perform slicing processing on the first data to obtain at least one slice data; A data storage unit is used to store the at least one slice data in the out-of-row data table.
11. The device according to claim 10, characterized in that The at least one slice data includes first slice data and second slice data connected in sequence; the data storage unit is specifically used for: storing the second slice data in a first physical area in the out-of-row data table; storing the first slice data and a physical address of the first physical area in a second physical area in the out-of-row data table; The address storage module is specifically used for: The physical address of the second physical area is stored in the first tuple of the data master table.
12. The device according to any one of claims 9 to 11, characterized in that: The device comprises: The request receiving module is further used to receive a data update request for the first tuple, wherein the data update request is used to request to update the first data through second data; The data storage module is further configured to store the second data in the out-of-row data table if the second data meets the out-of-row storage condition, and obtain a second storage address of the second data in the out-of-row data table; The address storage module is further used to update the first storage address in the first tuple of the data main table to the second storage address.
13. The device according to any one of claims 9 to 11, characterized in that: The device also includes: A data updating module is used to update the first tuple in the data main table based on the second data if the second data does not meet the out-of-row storage condition.
14. The device according to claim 12 or 13, characterized in that The device also includes: The data updating module is further used to delete the first data in the out-of-row data table based on the first storage address.
15. The device according to any one of claims 9 to 14, characterized in that: The device also includes: The request receiving module is further used to receive a data query request for the first tuple, wherein the data query request is used to query data of a target field among the multiple fields; A data reading module, configured to read the data of the target field from the off-row data table based on the off-row storage address if the off-row storage address of the target field is stored in the data master table; The result output module is further used to output the data query result of the first tuple based on the data of the target field.
16. The device according to claim 15, characterized in that The data of the target field includes a plurality of slice data; the result output module includes: A data splicing unit, configured to splice the plurality of slice data according to a reading order of the plurality of slice data; A result output unit is used to output the query result of the first tuple based on the concatenated data.
17. A computer device, characterized in that: The computer device includes a memory and a processor; The memory is used to store computer programs; The processor is used to execute the computer program to implement the data management method according to any one of claims 1 to 8.
18. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the data management method according to any one of claims 1 to 8 is implemented.
19. A computer program product, characterized in that The computer program product stores computer instructions, and when the computer instructions are executed by a processor, the data management method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Data management method and device, equipment, storage medium and computer program
CN119938664A
Postgresql-based big-field particular value indexing system and method
CN105488087A
Large-field data processing method and device, equipment and storage medium
CN113076325A
Large-field data processing method and device, equipment and storage medium
CN113076326A
Trimming blackhole clusters
US11704315B1