Data completion method and apparatus, electronic device, and storage medium

By dividing and positioning the area according to the primary key columns of the target table during the data completion process, the problem of excessive system resource consumption in the existing technology is solved, and a simplified data completion process and improved efficiency are achieved.

WO2025108008A1PCT designated stage expired Publication Date: 2025-05-30BEIJING OCEANBASE TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/127880
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-22
Filing Date
2024-10-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing technology requires creating column indexes for all columns during the data completion process, resulting in excessive system resource consumption.

Method used

By region division based on the primary key column of the target table, the starting row offset value of each target area is determined, and the offset value is used to locate and write data to the baseline sort string table, index creation and data sorting of all columns is avoided.

Benefits of technology

The data completion process is simplified, the system resource consumption and time cost are reduced, and data processing efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024127880_30052025_PF_FP_ABST
    Figure CN2024127880_30052025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of databases, and specifically provides a data completion method and apparatus, an electronic device, and a storage medium. The data completion method comprises: on the basis of a main key column in a target table to be processed, performing region division on the target table to obtain at least one target region; separately determining the number of rows of each target region; on the basis of the number of rows respectively corresponding to the at least one target region, respectively determining a starting row offset value of each target region; and, on the basis of the starting row offset value of each target region, respectively writing corresponding data in each target region in at least one target column group in the target table into a baseline sorted string table corresponding to a corresponding target column group. In this way, corresponding column indexes do not need to be separately created for all columns, and the data of columns does not need to be sorted by means of column indexes, which simplifies the tedious steps of data completion, and reduces the consumption of system resources and time costs.
Need to check novelty before this filing date? Find Prior Art

Description

Data completion method, device, electronic device and storage medium Technical Field

[0001] The present application relates to the field of database technology, and in particular to a method, device, electronic device, and storage medium for data completion. Background Art

[0002] In actual applications, in order to support different types of businesses at the same time, several columns in the table are usually combined to obtain column groups (CGs) according to business needs. One or more CGs can be set in the same table. When executing the Data Definition Language (DDL) for a table, it is usually necessary to update multiple CGs. Specifically, the updated data of each CG is written to the corresponding baseline sorted string table (SSTable). This process is called data completion.

[0003] In the prior art, a column index is usually created for each column in a table so that data can be completed for the CG based on the column index.

[0004] However, using this method requires creating column indexes for all columns, which consumes a lot of system resources.

[0005] Summary of the Invention

[0006] The purpose of the embodiments of the present application is to provide a method, device, electronic device, and storage medium for data completion, so as to reduce the consumption of system resources when performing data completion.

[0007] On the one hand, an embodiment of the present application provides a method for data completion, including: dividing the target table into regions according to the primary key column in the target table to be processed to obtain at least one target region; determining the number of rows in each target region respectively; determining the starting row offset value of each target region respectively according to the number of rows corresponding to at least one target region; the starting row offset value is used to locate the target region; according to the starting row offset value of each target region, writing the data corresponding to at least one target column combination in each target region in the target table into the baseline sorting string table corresponding to the corresponding target column combination.

[0008] In one embodiment, the target table is divided into regions based on the primary key column in the target table to be processed to obtain at least one target region, including: sampling the data in the primary key column, the data in the primary key column being sorted sequentially; comparing adjacent sampling values ​​in each sampling value to obtain corresponding comparison results; determining at least one data partition row based on the comparison results; and dividing the target table based on the at least one data partition row to obtain at least one target region.

[0009] In one embodiment, the starting row offset value of each target area is determined according to the number of rows corresponding to at least one target area, including: according to the order of at least one target area from top to bottom in the target table, for each target area, the following steps are performed: if the target area is the first area, the starting row offset value of the target area is determined to be a specified starting value; if the target area is not the first area, the starting row offset value of the target area is determined according to the number of rows corresponding to all target areas before the target area and the specified starting value.

[0010] In one embodiment, data corresponding to at least one target column combination in each target area is written into a baseline sorting string table corresponding to the corresponding target column combination according to the starting row offset value of each target area, including: for each target area, the following steps are performed by the thread corresponding to the target area: the number of combinations of at least one target column combination is determined as the initial batch size; according to the batch size, based on the starting row offset value of the target area, data corresponding to at least one target column combination in the target area is written into the corresponding baseline sorting string table; if it is determined that data writing fails, the latest successful number of the target column combination that has successfully written data is obtained; according to the batch size and the latest successful number, a new batch size is determined; according to the new batch size, a data write operation is performed on the target column combination for which data has not been successfully written.

[0011] In one embodiment, according to the new batch size, a data write operation is performed for the target column combination to which data has not been successfully written, including: looping the following steps until it is determined that the data is written successfully and there are no target column combinations to which data has not been successfully written: according to the new batch size, based on the starting row offset value of the target area, the data of the remaining target column combinations in the target area are written into the corresponding baseline sorting string table respectively; the remaining target column combinations are the target column combinations to which data has not been successfully written; if it is determined that the data write has failed, the latest successful number of the target column combinations to which data has been successfully written is obtained, and a new batch size is determined according to the batch size and the latest successful number; if it is determined that the data is written successfully and there are target column combinations to which data has not been successfully written, the new batch size is updated based on the set ratio; if it is determined that the data is written successfully and there are no target column combinations to which data has not been successfully written, it is determined that the data completion process is ended.

[0012] In one embodiment, determining a new batch size based on the batch size and the latest successful number includes: determining the product of the batch size and a set weight; and determining the maximum value of the product and the latest successful number as the new batch size.

[0013] In one embodiment, the method further includes: if it is determined that data writing fails, obtaining the sum of batch sizes corresponding to all threads; if the sum is not higher than the number of threads, determining that memory resources are insufficient and terminating the data completion process.

[0014] On the one hand, an embodiment of the present application provides a data completion device, including: a division unit, used to divide the target table into regions according to the primary key column in the target table to be processed, to obtain at least one target region; a determination unit, used to respectively determine the number of rows in each target region; a positioning unit, used to respectively determine the starting row offset value of each target region according to the number of rows corresponding to at least one target region; the starting row offset value is used to position the target region; a writing unit, used to write the data corresponding to each target region of at least one target column combination in the target table according to the starting row offset value of each target region into the baseline sorting string table corresponding to the corresponding target column combination.

[0015] In one embodiment, the partitioning unit is used to: sample the data in the primary key column, where the data in the primary key column is sorted sequentially; compare adjacent sampling values ​​in each sampling value to obtain corresponding comparison results; determine at least one data partitioning row based on the comparison results; and partition the target table based on the at least one data partitioning row to obtain at least one target area.

[0016] In one embodiment, the positioning unit is used to: perform the following steps for each target area in the order from top to bottom of at least one target area in the target table: if the target area is the first area, determine the starting row offset value of the target area as a specified starting value; if the target area is not the first area, determine the starting row offset value of the target area based on the number of rows corresponding to all target areas before the target area and the specified starting value.

[0017] In one embodiment, the write unit is used to: for each target area, respectively, perform the following steps through the thread corresponding to the target area: determine the number of combinations of at least one target column combination as the initial batch size; according to the batch size, based on the starting row offset value of the target area, write the corresponding data of the at least one target column combination in the target area into the corresponding baseline sorting string table; if it is determined that the data write fails, obtain the latest successful number of the target column combination that successfully writes data; determine a new batch size based on the batch size and the latest successful number; and perform a data write operation on the target column combination that has not successfully written data based on the new batch size.

[0018] In one embodiment, the writing unit is used to: loop through the following steps until it is determined that the data is written successfully and there are no target column combinations into which data has not been successfully written: according to the new batch size and based on the starting row offset value of the target area, write the data of the remaining target column combinations in the target area into the corresponding baseline sorting string table respectively; the remaining target column combinations are the target column combinations into which data has not been successfully written; if it is determined that the data writing has failed, obtain the latest successful number of the target column combinations into which data has been successfully written, and determine a new batch size based on the batch size and the latest successful number; if it is determined that the data is written successfully and there are target column combinations into which data has not been successfully written, update the new batch size based on the set ratio; if it is determined that the data is written successfully and there are no target column combinations into which data has not been successfully written, determine that the data completion process ends.

[0019] In one embodiment, the writing unit is used to: determine the product of the batch size and the set weight; and determine the maximum value of the product and the latest success number as the new batch size.

[0020] In one embodiment, the writing unit is further used to: if it is determined that the data writing fails, obtain the sum of the batch sizes corresponding to all threads; if the sum is not higher than the number of threads, determine that the memory resources are insufficient and end the data completion process.

[0021] On the one hand, an embodiment of the present application provides an electronic device, comprising: a processor; and a memory storing computer instructions, the computer instructions being used to enable the processor to execute the steps of the method provided in any of the various optional implementations of data completion described above.

[0022] On the one hand, an embodiment of the present application provides a storage medium storing computer instructions, which are used to enable a computer to execute the steps of the method provided in any of the various optional implementations of data completion described above.

[0023] In the data completion method, device, electronic device, and storage medium provided in the embodiments of the present application, the target table is divided into regions according to the primary key column in the target table to be processed to obtain at least one target region; the number of rows in each target region is determined respectively; the starting row offset value of each target region is determined respectively according to the number of rows corresponding to at least one target region; the starting row offset value is used to locate the target region; and according to the starting row offset value of each target region, the data corresponding to each target region of at least one target column combination in the target table is written into the baseline sorting string table corresponding to the corresponding target column combination. In this way, the table is divided into regions according to the primary key column, and the data in each region is located according to the offset of the starting row of each region. There is no need to create corresponding column indexes for all columns, nor is there any need to sort the data of each column according to the column index. This simplifies the tedious steps of data completion and reduces the consumption of system resources and time costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the specific implementation methods or the description of the prior art. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0025] FIG1 is a flow chart of a method for data completion in an embodiment of the present application.

[0026] FIG2 is an example diagram of determining a starting row offset value in an embodiment of the present application.

[0027] FIG3 is an example diagram of a method for adaptively adjusting batch times in an embodiment of the present application.

[0028] FIG4 is a structural block diagram of a data completion device in an embodiment of the present application.

[0029] FIG5 is a schematic structural diagram of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION

[0030] The technical solutions of the present application will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application. In addition, the technical features involved in the different embodiments of the present application described below can be combined with each other as long as they do not conflict with each other.

[0031] First, some of the terms involved in the embodiments of the present application are explained to facilitate understanding by those skilled in the art.

[0032] Terminal device: can be a mobile terminal, fixed terminal or portable terminal, such as a mobile phone, station, unit, device, multimedia computer, multimedia tablet, Internet node, communicator, desktop computer, laptop computer, notebook computer, netbook computer, tablet computer, personal communication system device, personal navigation device, personal digital assistant, audio / video player, digital camera / camcorder, positioning device, television receiver, radio broadcast receiver, e-book device, gaming device or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is also foreseeable that the terminal device can support any type of user interface (such as wearable device), etc.

[0033] Server: It can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, and big data and artificial intelligence platforms.

[0034] OceanBase (OB) is a distributed relational database that is continuously available, scalable, and high-performance.

[0035] Online Transaction Processing (OLTP): Generally refers to transaction operations with a small number of rows and short processing time, such as bank transactions.

[0036] Online Analytical Processing (OLAP): Generally refers to complex data analysis with many rows and a long processing time, such as report statistics.

[0037] Log Structured Merge Tree (LSM Tree): A data storage structure composed of multiple layers of independent data structures.

[0038] Baseline SStable: It is a data structure persisted on disk in the LSM Tree and consists of multiple macroblocks.

[0039] Macroblock: Contains multiple rows of sorted data and is persisted on disk. The macroblock size can be fixed at 2M and consists of multiple microblocks.

[0040] Microblock: It is the basic unit of macroblocks, contains multiple rows of sorted data, persists on disk, has a variable size, usually a few KB, and is the smallest unit for reading SSTable data from disk.

[0041] The technical concept of this application is described below.

[0042] In some database scenarios, in order to support different types of businesses (such as OLTP and OLAP) at the same time, some databases (such as OB databases) have developed column storage table functions, which allow users to specify several CGs in the same table.

[0043] For example, you can use the following statement to set CG: create table t1(c1 int primary key,c2 varchar(256),c3 int)with column group(all columns,each column). This statement creates Table 1 with CG. In terms of data organization at the storage layer, Table 1 has four CGs, namely (c1,c2,c3), (c1), (c2), and (c3). The data of each CG is organized into an SSTable. In other words, the data in each CG is actually stored in the corresponding SSTable.

[0044] When executing DDL on a table, data completion is usually required for multiple CGs. For example, when executing alter table t1 modify column c2 int on Table 1, data completion is required for both CGs (c1, c2, c3) and c2.

[0045] Under traditional technology, column indexes are usually created for all columns when a table is created. If a column index is not created when the table is created, it is created through the BUILD process when an index is needed.

[0046] However, creating column indexes for each column and sorting them according to column data consumes a large amount of system resources. Furthermore, when the BUILD process creates column indexes one by one, it needs to scan the main table data multiple times, which causes a large amount of input and output (IO) overhead and a longer build time.

[0047] Based on the defects of the above-mentioned related technologies, the embodiments of the present application provide a method, device, electronic device and storage medium for data completion, aiming to reduce the consumption of system resources when performing data completion.

[0048] The embodiment of the present application provides a method for data completion, which can be applied to electronic devices. The present application does not limit the type of electronic devices, which can be any type of device suitable for implementation, such as a terminal device or a server, etc. The present application will not elaborate on this. In one application scenario, the embodiment of the present application can be applied to an OB database. When executing DDL for a target table, the target column combination that needs data reorganization in the target table is determined, and the target table is divided into regions according to the primary key columns, and the data in the target region is located according to the offset of the starting row of each target region. Then, according to the offset corresponding to each target region, the corresponding data of each target column combination in the target region can be written into the SSTable corresponding to each target column combination to achieve data completion.

[0049] Refer to Figure 1, which is a flow chart of a data completion method in an embodiment of the present application, which is applied to an electronic device. The method is described below in conjunction with Figure 1. The specific implementation process of the method is as follows.

[0050] Step 100: Divide the target table into regions according to the primary key columns in the target table to be processed to obtain at least one target region.

[0051] In one embodiment, when executing step 100, the following steps may be used.

[0052] S1001: Sampling data in a primary key column to obtain multiple sampling values.

[0053] In one embodiment, the primary key of the target table is set, the primary key column in the target table is determined, the data in the primary key column is sorted, and when it is determined that data completion is required (DDL is executed for the target table), the data in the primary key column is sampled according to the set sampling interval to obtain multiple sampling values.

[0054] Among them, the primary key is used to uniquely identify a record (i.e., a row of data). The primary key value cannot be repeated and is not allowed to be empty. In the embodiment of the present application, the data in the primary key column is sorted sequentially. The sampling interval can be set to a fixed value or a non-fixed value (e.g., it can be continuously increasing). In actual applications, the sampling interval can be set according to the actual application scenario, such as setting the sampling interval to 5, which is not limited here.

[0055] DDL is used to define, delete, and modify relational schemas, as well as create various database objects. Database objects can include tables, clusters, indexes, views, functions, stored procedures, and triggers.

[0056] S1002: Compare adjacent sampling values ​​in each sampling value to obtain corresponding comparison results.

[0057] In one embodiment, the sample values ​​are all numerical values. According to the order of the sample values ​​in the primary key column, for each sample value, the following steps are performed: a difference between the sample value and the previous sample value is determined as a comparison result.

[0058] S1003: Determine at least one data partition row according to the comparison result.

[0059] As an example, if the difference between the sampling value and the previous sampling value is higher than a set difference (eg, 5), the row where the sampling value is located is determined as the data partition row.

[0060] In actual applications, the difference value can be set according to the actual application scenario and is not limited here.

[0061] S1004: Divide the target table according to at least one data partition row to obtain at least one target area.

[0062] This way, you can divide the table into multiple target areas.

[0063] Step 101: Determine the number of rows in each target area.

[0064] In one embodiment, a thread is assigned to each target region, and the data for each target region is shuffled to the corresponding thread. Each thread determines the number of rows to receive from the target region. It should be noted that the number of rows for different target regions can be the same or different. For example, one target region may have 10 rows, while another may have 20 rows.

[0065] Optionally, each thread can correspond to one or more target regions. Threads can be located on the same device or on different devices, without limitation.

[0066] In this way, multiple threads can be used to process each target area in parallel, thereby improving data processing efficiency.

[0067] Step 102: Determine a starting row offset value for each target area according to the number of rows corresponding to at least one target area.

[0068] The starting row offset value is used to locate the target area. In one embodiment, the following steps are performed for each target area according to the order of at least one target area from top to bottom in the target table.

[0069] If the target region is the first region, the starting row offset value of the target region is determined to be a specified starting value. As an example, the specified starting value may be 0.

[0070] If the target area is not the first area, the starting row offset of the target area is determined based on the number of rows corresponding to all target areas before the target area and the specified starting value.

[0071] As an example, if the start value is specified to be 0, the number of rows in the first area of ​​the target table is 10, and the number of rows in the second target area is 15, then the start row offset value of the third target area is 0+10+15=25.

[0072] As another example, each thread counts the number of rows in its corresponding target region. A designated thread obtains the statistical results of the other threads, determines the starting row offset value of each target region based on the statistical results of each thread, and sends the starting row offset value of each target region to the thread corresponding to the target region.

[0073] Optionally, the designated thread can be a pre-specified thread, a randomly selected thread, or the thread that last completed row count. In actual applications, the designated thread can be determined based on the actual application scenario and is not limited here.

[0074] Refer to Figure 2 for an example of determining a starting row offset value. In Figure 2, Thread 1, Thread 2, and Thread 3 each count the number of rows in their corresponding target regions (i.e., the first, second, and third regions). If the number of rows in each target region is 100, the starting row offset value for the first region is determined to be 0, the starting row offset value for the second region is 100, and the starting row offset value for the third region is determined to be 200.

[0075] Since the CG of the non-primary key column needs to identify each data row, in the embodiment of the present application, the starting row offset value of each target area is determined based on the accumulated number of rows in each target area. This process can be called row count, so that each data row in each target area can be identified based on the starting row offset value of each target area, thereby facilitating subsequent data completion without the need to establish a column index for each column separately, simplifying tedious operations, reducing system overhead, and improving data processing efficiency.

[0076] Step 103: According to the starting row offset value of each target area, the data corresponding to at least one target column combination in the target table in each target area are written into the baseline sorting string table corresponding to the corresponding target column combination.

[0077] In one embodiment, for each target area, the following steps are performed by a thread corresponding to the target area.

[0078] S1031: Determine the number of combinations of at least one target column combination as an initial batch size.

[0079] In one implementation, the number of target CGs to be processed in the target table, that is, the number of combinations, is obtained and determined as the initial batch size batch_size.

[0080] As an example, if the target table contains 4 CGs, the initial batch_size is determined to be 4.

[0081] Among them, CG is a combination of multiple columns in the table, and the target CG is the CG in each CG that needs to be reintegrated.

[0082] S1032: According to the batch size and based on the starting row offset value of the target area, combine the corresponding data of at least one target column in the target area and write them into the corresponding baseline sorting string table.

[0083] As an example, the target CG includes the first CG and the second CG. The first CG is column C2, and the second CG is columns C3 and C4. The number of target CGs is 2, and the batch size is 2. The thread writes the data of column C1 in the target area to the first baseline SSTable corresponding to the first CG, and writes the data of columns C3 and C4 in the target area to the second baseline SSTable corresponding to the second CG.

[0084] S1033: If it is determined that the data writing fails, the latest successful number of target column combinations that have successfully written data is obtained.

[0085] It should be noted that due to memory resource limitations, the data of each target CG may not be completely successfully written into the corresponding baseline SSTable. However, the data of some target CGs may be successfully written into the corresponding baseline SSTable. Therefore, when it is determined that data writing has failed, the number of this part of the target CGs (that is, the target CGs that have successfully written data at present) is obtained, that is, the latest successful number.

[0086] As an example, if the batch size is 4, 3 target CGs are written successfully, and one target CG fails to be written, the latest successful number is determined to be 3.

[0087] Furthermore, if it is determined that the data has been written successfully and the data in each target CG has been successfully written into the corresponding baseline SSTable, it is determined that the data has been completed and the data completion process ends.

[0088] S1034: Determine a new batch size based on the batch size and the latest successful quantity.

[0089] In one embodiment, the product of the batch size and the set weight is determined; the product of the batch size and the set weight is determined; and the maximum value of the product and the latest success number is determined as the new batch size.

[0090] The value range of the weight is set to (0,1].

[0091] As an example, the weight is set to 0.5. In actual applications, the weight can be set according to the actual application scenario and is not limited here.

[0092] In this way, if data writing fails, the batch size can be reduced by setting weights.

[0093] S1035: Perform a data write operation on the target column combination to which data has not been successfully written according to the new batch size.

[0094] In one embodiment, when executing S1035 , the following steps may be executed in a loop until it is determined that the data is written successfully and there is no target column combination to which the data is not written successfully.

[0095] S1035-1: According to the new batch size, based on the starting row offset value of the target area, the data of the remaining target columns in the target area are combined and written into the corresponding baseline sort string table respectively, and any one of steps S1035-2, S1035-3 and S1035-4 is executed.

[0096] The remaining target column combinations are target column combinations to which data has not been successfully written.

[0097] As an example, the target CG includes the first CG, the second CG, and the third CG. Among them, the second CG and the third CG are written successfully, and the first CG fails to be written. Then the first CG is the remaining target CG. Since the first CG only contains the C2 column, the data of the C2 column in the target area is written to the first baseline SSTable corresponding to the first CG.

[0098] S1035-2: If it is determined that the data writing fails, the latest successful number of target column combinations that have successfully written data is obtained, and a new batch size is determined based on the batch size and the latest successful number, and S1035-1 is executed.

[0099] Specifically, when executing S1035-3, the specific steps can be referred to S1033 and S1034, which will not be repeated here.

[0100] S1035-3: If it is determined that the data is written successfully and there is a target column combination where the data is not written successfully, then based on the set ratio, a new batch size is updated and S1035-1 is executed.

[0101] The set ratio may be no less than 1, for example, the set ratio is 2.

[0102] As an example, if data writing is determined to be successful and there are target column combinations where data writing has not been successful, the current batch size (i.e., the new batch size) multiplied by the set ratio is used as the new batch size. In actual applications, the set ratio can be set according to the actual application scenario and is not limited here.

[0103] Among them, the target column combination in which data is not successfully written refers to the target column combination in which the corresponding data in the target area is not successfully written into the corresponding baseline SStable.

[0104] S1035-4: If it is determined that the data is written successfully and there is no target column combination where the data is not written successfully, it is determined that the data completion process is ended.

[0105] Furthermore, during the data completion process, when it is determined that data writing fails, the sum of the batch sizes corresponding to all threads can be obtained. If the sum is not higher than the number of threads, it is determined that memory resources are insufficient and the data completion process is terminated.

[0106] The above embodiment is described below by taking a thread performing data completion on a target area as an example. Referring to FIG3 , an example diagram of a method for adaptively adjusting batch times is shown. The specific implementation process of the method may include the following steps.

[0107] Step 300: Determine the number of combinations of at least one target column combination as an initial batch size.

[0108] Step 301: According to the batch size, combine the corresponding data of at least one target column in the target area and write them into the corresponding baseline SStable respectively.

[0109] Step 302: Determine whether the data is written successfully. If so, execute step 303; otherwise, execute step 306.

[0110] Step 303: Multiply the batch size by the set ratio to obtain a new batch size.

[0111] Step 304: Determine whether there is a target CG to which data is to be written. If so, execute step 301; otherwise, execute step 305.

[0112] Among them, the target CG of the data to be written refers to the CG of the corresponding baseline SStable whose corresponding data in the target area is to be written.

[0113] Step 305: End the data completion process.

[0114] Step 306: Obtain the sum of the batch sizes corresponding to all threads.

[0115] Step 307 : Determine whether the total is not greater than the number of threads. If so, execute step 305 ; otherwise, execute step 308 .

[0116] Step 308: Determine the product of the batch size and the set weight.

[0117] Step 309: Determine the latest successful number of target column combinations that have successfully written data.

[0118] Step 310 : Determine the maximum value of the product and the latest successful number as the new batch size, and execute step 301 .

[0119] Specifically, for the specific steps of step 300 to step 310, please refer to the above steps 100 to step 103, which will not be repeated here.

[0120] It should be noted that the data of each row of the target CG in the corresponding target area is written into the corresponding baseline SStable through a thread. This process can be called rescan. Since it is necessary to maintain a buffer of at least the macroblock size (such as 2M), in scenarios with smaller memory specifications, in order to avoid memory explosion, it is necessary to scan the sorted data multiple times. In scenarios with larger memory specifications, it is hoped to reduce the number of scans as much as possible to save disk IO overhead. The maximum number of columns in the table can be 4096. Since the memory consumption of the data completion link is difficult to reserve accurately, in order to reduce overhead, some strategies (such as greedy algorithms) can be used for data completion. Specifically, the number N (N is a positive integer) of all target CGs can be used as batch_size, and rescan can be performed based on the batch_size. If it is determined that the data is written successfully, only rescan is required once. If data writing is determined to have failed, the process obtains the most recent successful number of CGs that successfully wrote data (last_succ_cg_count) and the most recent batch size (last_batch_size), and determines batch_size = max(last_batch_size / 2, last_succ_cg_count, 1). Furthermore, if the sum of the batch_sizes for all threads is less than or equal to the number of threads, memory resources are severely insufficient, and no further retries are made. The process returns a failure and exits.

[0121] In the embodiment of the present application, there is no need to establish a column index for each column separately, only a primary key needs to be set, which saves a lot of computing overhead. Furthermore, since the SSTable corresponding to the CG follows the row order of the primary key, row offset is used to locate the main table row without redundantly storing the primary key value, saving the storage overhead of the primary key value. Furthermore, it is also possible to complete the data of multiple CGs in a batch manner during a data scan, which greatly saves the construction overhead of columnar storage, and can also adaptively adjust the size of batch_size, further reducing the number of rescans, greatly reducing system overhead and time cost, and improving data processing efficiency and data completion performance.

[0122] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0123] Based on the same inventive concept, the embodiments of this application also provide a data completion device. Since the principles of the above-mentioned device and equipment for solving the problem are similar to those of a data completion method, the implementation of the above-mentioned device can refer to the implementation of the method, and the repeated parts will not be repeated. The device can be applied to electronic devices. This application does not limit the type of electronic device. It can be any device type suitable for implementation, such as a smartphone, tablet computer, etc. This application will not repeat this.

[0124] Refer to Figure 4, which is a block diagram of the structure of the data completion device in the embodiment of the present application. In some embodiments, the data completion device of the example of the present application includes: a division unit 401, which is used to divide the target table into regions according to the primary key column in the target table to be processed to obtain at least one target region; a determination unit 402, which is used to respectively determine the number of rows in each target region; a positioning unit 403, which is used to respectively determine the starting row offset value of each target region according to the number of rows corresponding to at least one target region; the starting row offset value is used to locate the target region; a writing unit 404, which is used to write the data corresponding to each target region of at least one target column combination in the target table according to the starting row offset value of each target region into the baseline sorting string table corresponding to the corresponding target column combination.

[0125] In one embodiment, the partitioning unit 401 is used to: sample the data in the primary key column, where the data in the primary key column is sorted sequentially; compare adjacent sampling values ​​in each sampling value to obtain corresponding comparison results; determine at least one data partition row based on the comparison results; and partition the target table based on the at least one data partition row to obtain at least one target area.

[0126] In one embodiment, the positioning unit 403 is used to: perform the following steps for each target area in the order from top to bottom of at least one target area in the target table: if the target area is the first area, determine the starting row offset value of the target area as the specified starting value; if the target area is not the first area, determine the starting row offset value of the target area based on the number of rows corresponding to all target areas before the target area and the specified starting value.

[0127] In one embodiment, the write unit 404 is used to: for each target area, respectively, perform the following steps through the thread corresponding to the target area: determine the number of combinations of at least one target column combination as the initial batch size; according to the batch size, based on the starting row offset value of the target area, write the corresponding data of the at least one target column combination in the target area into the corresponding baseline sort string table; if it is determined that the data write fails, obtain the latest successful number of the target column combination that successfully wrote the data; determine a new batch size based on the batch size and the latest successful number; and perform a data write operation on the target column combination that failed to write the data successfully based on the new batch size.

[0128] In one embodiment, the write unit 404 is used to: loop through the following steps until it is determined that the data is written successfully and there are no target column combinations to which data has not been successfully written: according to the new batch size and based on the starting row offset value of the target area, write the data of the remaining target column combinations in the target area into the corresponding baseline sorting string table respectively; the remaining target column combinations are the target column combinations to which data has not been successfully written; if it is determined that the data write has failed, obtain the latest successful number of target column combinations to which data has been successfully written, and determine a new batch size based on the batch size and the latest successful number; if it is determined that the data is written successfully and there are target column combinations to which data has not been successfully written, update the new batch size based on the set ratio; if it is determined that the data is written successfully and there are no target column combinations to which data has not been successfully written, determine that the data completion process ends.

[0129] In one embodiment, the writing unit 404 is configured to: determine the product of the batch size and the set weight; and determine the maximum value of the product and the latest success number as the new batch size.

[0130] In one embodiment, the writing unit 404 is further configured to: if it is determined that data writing fails, obtain the sum of the batch sizes corresponding to all threads; if the sum is not greater than the number of threads, determine that memory resources are insufficient and terminate the data completion process.

[0131] On the one hand, an embodiment of the present application provides an electronic device, comprising: a processor; and a memory storing computer instructions, the computer instructions being used to enable the processor to execute the steps of the method provided in any of the various optional implementations of data completion described above.

[0132] On the one hand, an embodiment of the present application provides a storage medium storing computer instructions, which are used to enable a computer to execute the steps of the method provided in any of the various optional implementations of data completion described above.

[0133] In the data completion method, device, electronic device, and storage medium provided in the embodiments of the present application, the target table is divided into regions according to the primary key column in the target table to be processed to obtain at least one target region; the number of rows in each target region is determined respectively; the starting row offset value of each target region is determined respectively according to the number of rows corresponding to at least one target region; the starting row offset value is used to locate the target region; and according to the starting row offset value of each target region, the data corresponding to each target region of at least one target column combination in the target table is written into the baseline sorting string table corresponding to the corresponding target column combination. In this way, the table is divided into regions according to the primary key column, and the data in each region is located according to the offset of the starting row of each region. There is no need to create corresponding column indexes for all columns, nor is there any need to sort the data of each column according to the column index. This simplifies the tedious steps of data completion and reduces the consumption of system resources and time costs.

[0134] An embodiment of the present application provides an electronic device, including: a processor; and a memory storing computer instructions, wherein the computer instructions are used to enable the processor to execute a method in any of the above embodiments.

[0135] An embodiment of the present application provides a storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the method of any of the above embodiments.

[0136] FIG5 shows a schematic structural diagram of an electronic device 5000. Referring to FIG5, the electronic device 5000 includes a processor 5010 and a memory 5020, and optionally, may further include a power supply 5030, a display unit 5040, and an input unit 5050.

[0137] The processor 5010 is the control center of the electronic device 5000 , which connects various components using various interfaces and lines, and performs various functions of the electronic device 5000 by running or executing software programs and / or data stored in the memory 5020 .

[0138] In the embodiment of the present application, the processor 5010 executes the various steps in the above embodiment when calling the computer program stored in the memory 5020.

[0139] Optionally, the processor 5010 may include one or more processing units. Preferably, the processor 5010 may integrate an application processor and a modem processor, wherein the application processor primarily processes the operating system, user interface, and applications, and the modem processor primarily processes wireless communications. It is understood that the modem processor may not be integrated into the processor 5010. In some embodiments, the processor and memory may be implemented on a single chip. In some embodiments, they may also be implemented on separate chips.

[0140] The memory 5020 may mainly include a program storage area and a data storage area. The program storage area may store an operating system, various applications, etc., and the data storage area may store data created based on the use of the electronic device 5000. In addition, the memory 5020 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device.

[0141] The electronic device 5000 also includes a power supply 5030 (such as a battery) for supplying power to various components. The power supply can be logically connected to the processor 5010 through a power management system, thereby managing functions such as charging, discharging, and power consumption through the power management system.

[0142] The display unit 5040 can be used to display information input by the user or information provided to the user, as well as various menus of the electronic device 5000. In the embodiment of the present application, it is mainly used to display the display interface of each application in the electronic device 5000 and objects such as text and pictures displayed on the display interface. The display unit 5040 may include a display panel 5041. The display panel 5041 can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc.

[0143] The input unit 5050 can be used to receive information such as numbers or characters input by the user. The input unit 5050 may include a touch panel 5051 and other input devices 5052. The touch panel 5051, also known as a touch screen, can receive user touch operations on or near it (for example, operations performed by the user using a finger, a stylus, or any other suitable object or accessory on or near the touch panel 5051).

[0144] Specifically, the touch panel 5051 can detect user touch operations and the signals generated by the touch operations, convert these signals into touch point coordinates, send them to the processor 5010, and receive and execute commands sent by the processor 5010. In addition, the touch panel 5051 can be implemented using various types, such as resistive, capacitive, infrared, and surface acoustic wave. Other input devices 5052 can include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control buttons, power buttons, etc.), a trackball, a mouse, a joystick, etc.

[0145] Of course, the touch panel 5051 may cover the display panel 5041. When the touch panel 5051 detects a touch operation on or near it, it transmits the information to the processor 5010 to determine the type of touch event. The processor 5010 then provides a corresponding visual output on the display panel 5041 based on the type of touch event. Although in FIG5 , the touch panel 5051 and the display panel 5041 are two independent components to implement the input and output functions of the electronic device 5000, in some embodiments, the touch panel 5051 and the display panel 5041 may be integrated to implement the input and output functions of the electronic device 5000.

[0146] The electronic device 5000 may further include one or more sensors, such as a pressure sensor, a gravity acceleration sensor, a proximity light sensor, etc. Of course, according to the needs of specific applications, the electronic device 5000 may also include other components such as a camera. Since these components are not the key components used in the embodiments of the present application, they are not shown in FIG5 and will not be described in detail.

[0147] Those skilled in the art will understand that FIG5 is merely an example of an electronic device and does not limit the electronic device. The electronic device may include more or fewer components than shown in the figure, or may combine certain components, or may include different components.

[0148] For the convenience of description, the above parts are divided into modules (or units) according to their functions and described separately. Of course, when implementing this application, the functions of each module (or unit) can be implemented in the same or multiple software or hardware.

[0149] Obviously, the above embodiments are merely examples for clarity of explanation and are not intended to limit the embodiments. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all embodiments here. Obvious variations or modifications arising therefrom remain within the scope of protection of this application.

Claims

1. A method for data completion, the method comprising: Divide the target table into regions according to the primary key column in the target table to be processed, and obtain at least one target region; Determine the number of rows for each target area respectively; Determine the starting row offset value of each target area according to the number of rows corresponding to the at least one target area; the starting row offset value is used to locate the target area; According to the starting row offset value of each target area, data corresponding to at least one target column combination in the target table in each target area are respectively written into the baseline sorting string table corresponding to the corresponding target column combination.

2. The method according to claim 1, wherein the target table is divided into regions according to the primary key column in the target table to be processed to obtain at least one target region, comprising: Sampling the data in the primary key column, where the data in the primary key column is sorted sequentially; Compare adjacent sampling values ​​in each sampling value respectively to obtain corresponding comparison results; Determine at least one data partition row according to the comparison result; The target table is divided according to the at least one data division row to obtain the at least one target area.

3. The method according to claim 1, wherein determining the starting row offset value of each target area according to the number of rows corresponding to the at least one target area comprises: According to the order of the at least one target area from top to bottom in the target table, for each target area, perform the following steps: If the target area is the first area, determining the starting row offset value of the target area to be a specified starting value; If the target area is not the first area, the starting row offset value of the target area is determined according to the number of rows corresponding to all target areas before the target area and the specified starting value.

4. The method according to any one of claims 1 to 3, wherein the step of writing the data corresponding to at least one target column combination in each target area in the target table into the baseline sorting string table corresponding to the corresponding target column combination according to the starting row offset value of each target area comprises: For each target area, the following steps are performed through the thread corresponding to the target area: determining the number of combinations of the at least one target column combination as an initial batch size; According to the batch size, based on the starting row offset value of the target area, combining the at least one target column with the corresponding data in the target area, and writing them into the corresponding baseline sorting string table respectively; If it is determined that the data writing fails, the latest successful number of the target column combination that has successfully written data is obtained; Determine a new batch size according to the batch size and the latest successful quantity; According to the new batch size, a data write operation is performed for a target column combination to which data has not been successfully written.

5. The method according to claim 4, wherein the step of performing a data write operation on a target column combination to which data has not been successfully written according to the new batch size comprises: The following steps are executed repeatedly until the data is written successfully and there is no target column combination to which the data is not written successfully: According to the new batch size, based on the starting row offset value of the target area, the remaining target column combinations are combined with the data in the target area, and are written into the corresponding baseline sorting string table respectively; the remaining target column combinations are target column combinations for which data is not currently written successfully; If it is determined that the data writing fails, the latest successful number of the target column combination that has successfully written data is obtained, and a new batch size is determined according to the batch size and the latest successful number; If it is determined that the data is written successfully, and there is a target column combination where the data is not written successfully, then based on the set ratio, the new batch size is updated; If it is determined that the data is written successfully and there is no target column combination to which the data is not written successfully, it is determined that the data completion process ends.

6. The method according to claim 5, wherein determining a new batch size according to the batch size and the latest successful number comprises: Determining the product of the batch size and a set weight; The maximum value of the product and the latest successful number is determined as the new batch size.

7. The method according to claim 4, further comprising: If it is determined that data writing fails, the sum of the batch sizes corresponding to all threads is obtained; If the sum is not greater than the number of threads, it is determined that memory resources are insufficient and the data completion process ends.

8. A data completion device, comprising: A partitioning unit, configured to partition the target table into regions according to the primary key column in the target table to be processed, to obtain at least one target region; A determination unit, used to determine the number of rows of each target area respectively; A positioning unit, used to determine the starting row offset value of each target area according to the number of rows corresponding to the at least one target area; The starting row offset value is used to locate the target area; The writing unit is used to write the data corresponding to at least one target column combination in the target table in each target area into the baseline sorting string table corresponding to the corresponding target column combination according to the starting row offset value of each target area.

9. An electronic device, comprising: processor; as well as A memory storing computer instructions, wherein the computer instructions are used to enable the processor to execute the method according to any one of claims 1-7.

10. A storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Table data extraction method and device, storage medium and electronic equipment

    CN113869014A

  • Partition table establishment method and device and data writing and reading method and device for partition table

    CN114490674A

  • Incremental data synchronization field completion method and device, electronic equipment and storage medium

    CN116894028A

  • Data completion method and device, electronic equipment and storage medium

    CN117807061A