Data insertion method and device, equipment, medium and product

By dividing the data set according to the maximum number of groups when inserting data in the database, the data set is divided according to the maximum number of groups and generating corresponding SQL statements, the problem of low data insertion performance in the prior art is solved, and the execution efficiency of data insertion is improved.

CN120011339APending Publication Date: 2025-05-16CETC JINCANG (BEIJING) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510072357.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When the prior art inserts a large amount of data into the database, performance is affected, and the SQL statements caused by batch insertion are too long, resulting in more time-consuming database semantic analysis and low execution efficiency.

Method used

By receiving the data insertion request, the structured query language statement is parsed, and multiple groups of data to be inserted are divided into data sets according to the maximum number of groups, and the SQL statement corresponding to each data set is merged and generated, and sent to the database for execution.

Benefits of technology

It significantly reduces the number of SQL statements and the number of transmissions, reduces the overhead of the database parsing and optimization, and thus improves the execution efficiency of data insertion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011339A_ABST
    Figure CN120011339A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data insertion method and device, equipment, a medium and a product. The method comprises the following steps: receiving a data insertion request, and analyzing to obtain a structured query language statement according to the data insertion request; the data insertion request is used for requesting to insert multiple groups of to-be-inserted data into a database; according to the maximum group number, dividing the multiple groups of to-be-inserted data to obtain a data set; wherein the group number of the data to be inserted in each data set is not greater than the maximum group number; merging the to-be-inserted data in each data set, and obtaining a statement corresponding to the data set based on the structured query language statement; and sending the statements corresponding to the data set to the database, so that the database executes data insertion based on the statements. The method can improve the execution efficiency of data insertion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of database technology, and in particular to a data insertion method, device, equipment, medium and product. Background Art

[0002] In modern information systems, databases are the core components for storing and managing large amounts of data. With the advent of the big data era, various applications need to frequently insert large amounts of data into databases, which places higher demands on the performance and efficiency of databases.

[0003] In actual applications, there are two ways to insert a large amount of data into the database through a persistence layer framework (such as Mybatis):

[0004] The first method is to use the For loop statement to insert data one by one. However, each insertion operation requires establishing a connection with the database, which will affect system performance. The second method is to use the foreach tag provided by the persistence layer framework to perform batch insertion in the mapper file. However, this method will cause the structured query language statement (SQL statement) to be too long, resulting in a long database semantic analysis, resulting in low execution efficiency of data insertion. Summary of the invention

[0005] Embodiments of the present application provide a data insertion method, apparatus, device, medium, and product to improve the execution efficiency of data insertion.

[0006] In a first aspect, an embodiment of the present application provides a data insertion method, comprising: receiving a data insertion request, and parsing a structured query language statement according to the data insertion request; the data insertion request is used to request to insert multiple groups of data to be inserted into a database; according to a maximum number of groups, the multiple groups of data to be inserted are divided to obtain data sets; wherein the number of groups of data to be inserted in each data set is not greater than the maximum number of groups; the data to be inserted in each data set are merged, and a statement corresponding to the data set is obtained based on the structured query language statement; the statement corresponding to the data set is sent to the database, so that the database executes data insertion based on the statement.

[0007] In a possible implementation, a structured query language statement includes an insert operation field, an insert position field, a keyword field, and a data field; the method further includes: generating a list field according to the structure of a single group of data to be inserted in the data field; extracting a main statement from the structured query language statement; the main statement includes the insert operation field, the insert position field, the keyword field, and the list field of the structured query language statement; according to the list field, obtaining the data volume of the single group of data to be inserted, and setting the maximum number of groups according to the data volume of the single group of data to be inserted.

[0008] In a possible implementation, extracting a main statement from a structured query language statement includes: locating the positions of an insert operation field and a keyword field in the structured query language statement; based on the positions of the insert operation field and the keyword field, extracting fields from the insert operation field to the keyword field in the structured query language statement, and concatenating the fields with a list field to obtain the main statement.

[0009] In a possible implementation, the method further includes: determining a maximum number of groups according to a user-defined number of groups.

[0010] In one possible implementation, the maximum number of groups is set according to the data volume of a single group of data to be inserted, including: determining a value range of the data volume of the single group of data to be inserted; wherein different value ranges correspond to different methods of setting the maximum number of groups; according to the value range of the data volume of the single group of data to be inserted, the maximum number of groups is set using a corresponding setting method; wherein the value range and the maximum number of groups are negatively correlated.

[0011] In a possible implementation, the maximum number of groups is set according to the value range of the data volume of a single group of data to be inserted, using a corresponding setting method, including: if the data volume of a single group of data to be inserted is within a first value range, then a preset first value is used as the maximum number of groups; if the data volume of a single group of data to be inserted is within a second value range, then the data volume of the single group of data to be inserted is normalized and the highest bit is calculated to obtain the maximum number of groups; if the data volume of a single group of data to be inserted is within a third value range, then the preset second value is used as the maximum number of groups; wherein the critical value of the first value range is smaller than the second value range, the critical value of the second value range is smaller than the third value range, and the first value is greater than the second value.

[0012] In a second aspect, an embodiment of the present application provides a data insertion device, comprising: a first processing module, used to receive a data insertion request, and parse the data insertion request to obtain a structured query language statement; the data insertion request is used to request to insert multiple groups of data to be inserted into a database; a division module, used to divide the multiple groups of data to be inserted into data sets according to a maximum number of groups; wherein the number of groups of data to be inserted in each data set is not greater than the maximum number of groups; a second processing module, used to merge the data to be inserted in each data set, and obtain a statement corresponding to the data set based on the structured query language statement; and send the statement corresponding to the data set to the database so that the database performs data insertion based on the statement.

[0013] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory, so that the processor executes the first aspect above and / or various possible implementations of the first aspect.

[0014] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the first aspect above and / or various possible implementations of the first aspect.

[0015] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above first aspect and / or various possible implementation methods of the first aspect.

[0016] In the data insertion method, device, equipment, medium and product provided by the embodiment of the present application, a data insertion request is first received, and according to the received data insertion request, a structured query language statement is parsed to obtain a data insertion request; the data insertion request is used to request to insert multiple groups of data to be inserted into the database; then, according to the maximum number of groups, the multiple groups of data to be inserted are divided to obtain data sets; wherein the number of groups of data to be inserted in each data set is not greater than the maximum number of groups; finally, the data to be inserted in each data set are merged, and the statement corresponding to the data set is obtained based on the structured query language statement; and the statement corresponding to the data set is sent to the database, so that the database executes data insertion based on the statement. The scheme of the present application, by setting the maximum number of groups, divides multiple groups of data to be inserted into multiple data sets, and generates SQL statements corresponding to each data set, which significantly reduces the number of SQL statements and the corresponding number of transmissions. The database parsing and optimization overhead is reduced, thereby improving the execution efficiency of data insertion. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0018] Figure 1 exemplarily shows a flow chart of the data insertion method provided in the first embodiment of the present application;

[0019] Figure 2 A flowchart of a data insertion method provided in Embodiment 2 of the present application;

[0020] Figure 3 A flowchart of another data insertion method provided in Embodiment 2 of the present application;

[0021] Figure 4 The structural diagram of the data insertion device provided in the third embodiment of the present application is exemplarily shown in FIG.

[0022] Figure 5 This is a schematic diagram of the structure of an electronic device provided in Example 4 of the present application.

[0023] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0024] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0025] The terms "including" and "having" in this application are used to express an open-ended inclusion, and mean that in addition to the listed elements / components / etc., there may be additional elements / components / etc.; the terms "first" and "second" etc. are only used as marks or distinctions, and are not intended to limit the order or quantity of their objects. In addition, the different elements and areas in the drawings are only shown schematically, and are therefore not limited to the sizes or distances shown in the drawings. The technical solution is described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0026] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0027] In many application developments, persistence layer frameworks are widely used. When a large amount of data needs to be inserted into a database, persistence layer frameworks (such as Mybatis) usually use a For loop to insert data, or add data in batches in a collection.

[0028] As an example, in actual applications, when using the For loop statement to insert data one by one, each loop will call the Mybatis insert operation once. The advantage of this method is that it is simple and direct, but since each insertion is an independent operation, each insertion requires an interaction with the database. Therefore, when inserting a large amount of data, system performance will be affected. When adding in batches in the form of a collection, use the foreach tag provided by Mybatis to perform batch insertion in the mapper file. The foreach tag can be used to insert multiple sets of data into the database at one time. Mybatis will generate a structured query language statement (SQL statement) containing multiple sets of data to achieve batch insertion.

[0029] However, the above two data insertion methods have obvious performance bottlenecks. On the one hand, the generated structured query language statements (SQL statements) are very long, which causes the database to take too long to perform semantic analysis; on the other hand, the optimization of each set of data by the database optimizer layer cannot be directly reused, making it difficult to further improve the execution efficiency and unable to meet the performance requirements in scenarios such as high concurrency and large data volumes, resulting in low execution efficiency of data insertion.

[0030] The technical content provided by this application is intended to solve the above-mentioned technical problems of related technologies. In the embodiment of this application, a data insertion request is first received, and according to the received data insertion request, a structured query language statement is parsed to obtain a plurality of groups of data to be inserted into the database; then, according to the maximum number of groups, the plurality of groups of data to be inserted are divided to obtain data sets; wherein, the number of groups of data to be inserted in each data set is not greater than the maximum number of groups; finally, the data to be inserted in each data set are merged, and the statement corresponding to the data set is obtained based on the structured query language statement; and the statement corresponding to the data set is sent to the database so that the database executes data insertion based on the statement. The scheme of this application, by setting the maximum number of groups, divides the plurality of groups of data to be inserted into a plurality of data sets, and generates a SQL statement corresponding to each data set, which significantly reduces the number of SQL statements and the corresponding number of transmissions. The parsing and optimization overhead of the database is reduced, thereby improving the execution efficiency of data insertion.

[0031] Some aspects of the examples of this application relate to the above considerations. The following is an example introduction of the solution in combination with some examples.

[0032] Embodiment 1

[0033] Figure 1 The flowchart of the data insertion method provided in the first embodiment of the present application is exemplarily shown in FIG. 1 . The execution subject of this embodiment may be a data insertion device, such as Figure 1 As shown, the method includes:

[0034] Step 101: receiving a data insertion request, and parsing the data insertion request to obtain a structured query language statement; the data insertion request is used to request to insert multiple groups of data to be inserted into a database;

[0035] Step 102: Divide the multiple groups of data to be inserted into data sets according to the maximum number of groups; wherein the number of groups of data to be inserted in each data set is not greater than the maximum number of groups;

[0036] Step 103: merge the data to be inserted in each data set, and obtain a statement corresponding to the data set based on the structured query language statement; send the statement corresponding to the data set to the database, so that the database executes data insertion based on the statement.

[0037] In practical applications, the executor of the method can be a data insertion device. There are many ways to implement the data insertion device. For example, it can be implemented through a computer program, such as application software, etc.; or it can be implemented as a medium storing relevant computer programs, such as a USB flash drive, a cloud disk, etc.; or it can be implemented through a physical device that integrates or installs relevant computer programs, such as a chip, etc.

[0038] In this example, when the system receives a data insertion request, the request carries instructions for inserting multiple sets of data to be inserted into the database. Specifically, a data insertion request refers to an instruction sent by a client (such as an application, user interface, or other system module) to a database, requesting that one or more sets of new data be inserted into a specified database. These data to be inserted include but are not limited to the user inputting and submitting data through a front-end interface (such as a web form, mobile application interface, etc.), triggering a data insertion request; when the application executes certain business logic, it needs to insert the generated data records into the database; in data migration or synchronization operations, the system will generate data insertion requests in batches to migrate data from one database to another or synchronize it to different data stores; certain scheduled tasks (such as logging, data backup, etc.) will periodically generate data insertion requests to record relevant data in the database, etc.

[0039] Accordingly, in order to process this data efficiently, the system will parse the received data insertion request. Through the internal parsing module, the information in the data insertion request is disassembled and converted into a structured query language statement (SQL statement). This process involves operations such as identifying the request format, verifying the data type, and matching the SQL statement template. For example, if the data insertion request is to insert multiple sets of user information, the SQL statement obtained after parsing is in the form of "insertinto users (name, age, email) values(?,?,?,?,?),…(?,?,?,?,?,?)", in which the question mark "?" is used as a placeholder, waiting for subsequent data to be filled.

[0040] After receiving the SQL statement, the system will reasonably divide the multiple groups of data to be inserted according to the preset maximum number of groups. The maximum number of groups is a key performance tuning parameter, which determines the amount of data sent to the database each time. The system will divide the multiple groups of data to be inserted into several data sets, where the number of groups of data to be inserted in each data set is not greater than the maximum number of groups. For example, if the amount of data in multiple groups of data to be inserted is less than or equal to the maximum number of groups, all the data to be inserted will be divided into one data set; if the amount of data in multiple groups of data to be inserted is greater than the maximum number of groups, all the data to be inserted will be divided into multiple data sets with the maximum number of groups as a boundary until all the data to be inserted falls into the corresponding data set.

[0041] Correspondingly, after the data sets are divided, the system will merge the data to be inserted in each data set. This process is to integrate multiple groups of data to be inserted in the same data set, and then based on the SQL statement obtained by the previous analysis, the corresponding main SQL statement can be obtained to generate the complete SQL statement corresponding to the data set. For example, if the preset maximum number of groups is 10, then there are 10 groups of user data in a divided data set, and the merged SQL statement is "insert into users(name,age,email)values('user1','25','user1@example.com'),('user2','30','user2@example.com'), … , ('user10', '28', 'user10@example.com')". Finally, the system sends the SQL statements corresponding to these data sets to the database. After receiving the statement, the database will perform the corresponding data insertion operation according to the content of the statement and accurately store the data to be inserted into the specified database.

[0042] In the above example, by setting the maximum number of groups, and dividing multiple groups of data to be inserted into data sets according to the maximum number of groups, and then merging the data to be inserted in the data sets, generating statements corresponding to the data sets based on structured query language statements and sending them to the database, executing data insertion operations, and adopting array batch transmission, the execution efficiency of data insertion is improved, and the problem that the optimizer cannot be reused is solved.

[0043] Based on the above example, the structured query language statement includes an insert operation field, an insert position field, a keyword field and a data field; the method further includes:

[0044] Generate a list field according to the structure of a single set of data to be inserted in the data field;

[0045] Extracting a main statement from a structured query language statement; the main statement includes an insert operation field, an insert position field, a keyword field, and a list field of the structured query language statement;

[0046] According to the list field, the data volume of a single group of data to be inserted is obtained, and the maximum number of groups is set according to the data volume of the single group of data to be inserted.

[0047] Specifically, in database operations, SQL statements are the core tool for data insertion. A complete SQL statement usually includes an insert operation field (such as insert into), an insert position field (such as table name table_name), a keyword field (such as values), and a data field (specific data). Among them, the insert operation field represents the operation type, that is, performing a data insert operation into the database; the insert position field represents the target location for data insertion, that is, the name of the table where the data will be inserted; the keyword field represents the introduction of specific data, that is, the data after the keyword field is specified as the data to be inserted into the database; the data field represents the specific data, that is, multiple groups of data to be inserted. For example, an SQL statement is "insert intotable_name ('column1','column2') values ​​('value1', 'value2')". In this SQL statement, "insert into" is the insert operation field; table_name('column1', 'column2') is the insert position field; values ​​is the keyword field; ('value1', 'value2') is the data field. This means inserting ('value1', 'value2') into the ('column1', 'column2') table in the database.

[0048] Accordingly, the list field is generated based on the structure of a single set of data to be inserted in the data field. The system generates a list field that defines the structure of each set of data. The list field specifies the column name of the data to be inserted, which is used to clarify the structure of each set of data and ensure that the data is inserted in the correct position. For example, if each set of data contains three fields: name, age, and email, then the list field is (name, age, email). This list field will be used to build the main statement, which is the core part of the SQL statement.

[0049] After generating the list fields, the system will extract the main statement from the SQL statement. The main statement is the basic framework of the SQL statement, which defines the basic structure of the insert operation. The purpose of extracting the main statement is to reuse this structure in subsequent batch insert operations, thereby reducing the repeated construction of SQL statements, reducing semantic analysis, and improving data insertion efficiency. Among them, the main statement includes the insert operation field, insert position field, keyword field, and list field of the SQL statement. In an example, the main statement can be "insert into table_name values ​​(name, age, email)".

[0050] Accordingly, the system calculates the data volume of a single group of data to be inserted based on the list fields. The data volume refers to the number of fields contained in each group of data. For example, if each group of data contains 3 fields, then the data volume is 3. According to this data volume, set the maximum number of groups.

[0051] Through the solution of this example, the SQL statement structure is simplified by generating list fields and extracting the main statement. Dynamically setting the maximum number of groups balances the database load, ensures efficient processing of batch insertion, reduces response time, and improves the execution efficiency of data insertion operations.

[0052] Based on the above example, the main statement is extracted from the structured query language statement, including:

[0053] Locate the position of inserting the operation field and the keyword field in the structured query language statement;

[0054] Based on the positions of the insert operation field and the keyword field, the field between the insert operation field and the keyword field in the structured query language statement is extracted, and the field is concatenated with the list field to obtain the main statement.

[0055] Specifically, when extracting the main statement from the SQL statement, you first need to locate the position of the insert operation field and the keyword field. Combined with the above example, the insert operation field is usually "insert into", which is a fixed part in the SQL statement used to identify the insert operation. The keyword field is "values", which is used to introduce specific data values. By parsing the SQL statement, you can find the position of these two fields. For example, in the SQL statement "insert into table_name values('user1','25','user1@example.com'),('user2','30','user2@example.com'),…,('user10','28','user10@example.com')", "insert into" is located at the beginning of the SQL statement, and "values" is located after the field list. Locating the position of these two fields is a key step in extracting the main statement because it determines the range of fields that need to be extracted.

[0056] Accordingly, after locating the positions of the insert operation field and the keyword field, all fields between the insert operation field and the keyword field are extracted. For example, in the above SQL statement, the field between "insert into" and "values" is "insert into table_name values". After extracting these fields, they are concatenated with the list fields to obtain the main statement. The concatenated main statement is in the form of "insert into table_name values ​​(name, age, email)". Through the solution of this example, it can be ensured that each insert operation uses the same structure, reducing semantic analysis, thereby improving insertion efficiency and database performance.

[0057] Based on the above example, the maximum number of groups is set according to the amount of data to be inserted in a single group, including:

[0058] Determine the value range of the data amount of a single group of data to be inserted; wherein different value ranges correspond to different ways of setting the maximum number of groups;

[0059] According to the value range of the data amount of a single group of data to be inserted, a corresponding setting method is adopted to set the maximum number of groups; wherein the value range and the maximum number of groups are negatively correlated.

[0060] In this example, when setting the maximum number of groups based on the amount of data in a single group to be inserted, you first need to determine the value range of the amount of data in a single group to be inserted. The amount of data refers to the number of fields contained in each group of data. For example, if each group of data contains 3 fields (such as name, age, email), the amount of data is 3. Different value ranges correspond to different maximum number of group settings. The division of the value range can be determined based on actual business needs and database performance test results. For example, the amount of data can be divided into the following value ranges:

[0061] (1) The data volume is 0;

[0062] (2) Small data volume: 1 <= data volume <= 255;

[0063] (3) Medium data volume: 255 < data volume <= 16383;

[0064] (4) Large data volume: data volume > 16383.

[0065] The division of these value ranges is based on a comprehensive consideration of database processing capabilities and network transmission efficiency, and is not limited here. Data in the small data range is usually inserted more efficiently, and the maximum number of groups can be appropriately increased; data in the large data range needs to reduce the maximum number of groups to avoid excessive pressure on the database.

[0066] According to the value range of the amount of data to be inserted in a single group, the corresponding setting method is used to determine the maximum number of groups. The value range and the maximum number of groups are negatively correlated, that is, the larger the amount of data, the smaller the maximum number of groups.

[0067] In one example, according to the value range of the amount of data in a single group of data to be inserted, a corresponding setting method is adopted to set the maximum number of groups, including:

[0068] If the amount of data in a single group of data to be inserted is within the first value range, the preset first value is used as the maximum number of groups;

[0069] If the data volume of a single group of data to be inserted is within the second value range, the data volume of the single group of data to be inserted is normalized and the highest bit is calculated to obtain the maximum number of groups;

[0070] If the amount of data in a single group to be inserted is within the third value range, the preset second value is used as the maximum number of groups; wherein the critical value of the first value range is smaller than the second value range, the critical value of the second value range is smaller than the third value range, and the first value is greater than the second value.

[0071] The specific settings are as follows:

[0072] When the data volume is 0, the maximum number of groups is set to 1024; when the data volume is small (1 <= data volume <= 255), a larger maximum number of groups can be set, such as 128; when the data volume is medium (255 < data volume <= 16383), the data volume of a single group of data to be inserted is normalized to calculate the highest bit to obtain the maximum number of groups, that is, Integer.highestOneBit((Short.MAX_VALUE-1) / data volume). The specific calculation method is the same as the existing technology and no additional explanation is given here. When the data volume is large (data volume > 16383), the maximum number of groups is set to 1. For the specific maximum number of groups, please refer to Table 1.

[0073] Table 1

[0074]

[0075] The above is only an example of setting the maximum number of groups. In actual applications, the maximum number of groups can be set according to actual conditions, as long as the critical value of the first value range is less than the second value range, the critical value of the second value range is less than the third value range, and the first value is greater than the second value. Through the solution of this example, the system can dynamically adjust the maximum number of groups according to the amount of data to be inserted in a single group, thereby maximizing the efficiency of batch data insertion while ensuring database performance.

[0076] In one example, the method further includes: determining a maximum number of groups based on a customized number of groups.

[0077] Specifically, the maximum number of groups can be implemented based on the user-defined number of groups. Users can dynamically adjust the maximum number of groups based on specific business needs and the actual performance of the database. To achieve this function, the system provides a user interface that allows users to enter or select the desired maximum number of groups. After the system receives the number of groups entered by the user, it will apply it as a parameter to subsequent data insertion operations. When performing batch data insertion, the system will divide the data to be inserted into multiple data sets based on this user-defined maximum number of groups, and the number of data groups in each set does not exceed the maximum number of groups set by the user. Through the solution of this example, users can flexibly control the efficiency of data insertion and resource usage according to different business scenarios and performance requirements, thereby achieving optimal system performance.

[0078] In the data insertion method provided in this embodiment, a data insertion request is first received, and according to the received data insertion request, a structured query language statement is parsed to obtain a data insertion request; the data insertion request is used to request to insert multiple groups of data to be inserted into the database; then, according to the maximum number of groups, the multiple groups of data to be inserted are divided to obtain data sets; wherein, the number of groups of data to be inserted in each data set is not greater than the maximum number of groups; finally, the data to be inserted in each data set are merged, and the statement corresponding to the data set is obtained based on the structured query language statement; and the statement corresponding to the data set is sent to the database, so that the database executes data insertion based on the statement. The scheme of the present application, by setting the maximum number of groups, divides multiple groups of data to be inserted into multiple data sets, and generates SQL statements corresponding to each data set, which significantly reduces the number of SQL statements and the corresponding number of transmissions. The database parsing and optimization overhead is reduced, thereby improving the execution efficiency of data insertion.

[0079] Embodiment 2

[0080] The data insertion method provided by the present application is described in detail below with a specific embodiment.

[0081] Figure 2 A flowchart of a data insertion method provided in Embodiment 2 of the present application is shown in FIG. Figure 2 As shown, the process is as follows: (Take the case of inserting 12,000 groups of data into the database this time, and the maximum number of groups is set according to the amount of data to be inserted in a single group as an example)

[0082] Step 201, receiving a data insertion request sent by a user; and parsing the SQL statement "insert into table_name values ​​(?,?,?,?,?),…,(?,?,?,?,?)" according to the data insertion request;

[0083] Step 202, according to the extraction method of the above embodiment, extract the main statement "insert into table_namevalues ​​(?,?,?,?,?),…,(?,?,?,?,?)" from the SQL statement "insert into table_namevalues ​​(x,x,x,x,x)", wherein the list field is (x,x,x,x,x);

[0084] Step 203, according to the list field, the data amount of a single group of data to be inserted is obtained as 5, and the maximum number of groups obtained from Table 1 is 128;

[0085] Step 204, according to the maximum number of groups 128, the 12000 groups of data are divided into 94 data sets, of which 93 data sets include 128 groups of data, and the remaining 1 data set includes 12000-93*128=96 groups of data;

[0086] Step 205, merging the data in the above 94 data sets to obtain 94 data insertion statements corresponding to the 94 data sets;

[0087] Step 206, sending the above 94 statements to the database, so that the database executes data insertion based on the above 94 statements.

[0088] Figure 3 A flowchart of another data insertion method provided in Embodiment 2 of the present application is shown in FIG. Figure 3 As shown, the process is as follows: (Take the example of inserting 12,000 groups of data into the database this time, and the maximum number of groups is set to 1,000 according to the user's custom setting)

[0089] Step 301, receiving a data insertion request sent by a user; and according to the data insertion request, parsing to obtain an SQL statement "insert into table_name values ​​(?,?,?,?,?),…,(?,?,?,?,?,?)";

[0090] Step 302, according to the extraction method of the above embodiment, extract the main statement "insert into table_namevalues ​​(?,?,?,?,?),…,(?,?,?,?,?)" from the SQL statement "insert into table_namevalues ​​(x,x,x,x,x)";

[0091] Step 303, according to the maximum number of groups 1000 set by the user, the 12000 groups of data are divided into 12 data sets, wherein each data set includes 1000 groups of data;

[0092] Step 304, merging the data in each of the above data sets to obtain 12 data insert statements corresponding to the 12 data sets;

[0093] Step 305: Send the above 12 statements to the database, so that the database executes data insertion based on the above 12 statements.

[0094] The specific method of data insertion can refer to the content of the aforementioned embodiment. In summary, the data insertion method provided in this example divides multiple groups of data to be inserted into multiple data sets by setting the maximum number of groups, and generates SQL statements corresponding to each data set, which significantly reduces the number of SQL statements and the corresponding number of transmissions. It reduces the database parsing and optimization overhead, thereby improving the execution efficiency of data insertion.

[0095] Embodiment 3

[0096] Figure 4 The structure diagram of the data insertion device provided in the third embodiment of the present application is exemplarily shown in FIG. Figure 4 As shown, the device comprises:

[0097] The first processing module 41 is used to receive a data insertion request and parse the data insertion request to obtain a structured query language statement; the data insertion request is used to request to insert multiple groups of data to be inserted into the database;

[0098] A division module 42 is used to divide the multiple groups of data to be inserted into data sets according to the maximum number of groups; wherein the number of groups of data to be inserted in each data set is not greater than the maximum number of groups;

[0099] The second processing module 43 is used to merge the data to be inserted in each data set, and obtain the statement corresponding to the data set based on the structured query language statement; send the statement corresponding to the data set to the database, so that the database executes data insertion based on the statement.

[0100] In practical applications, there are many ways to implement the data insertion device. For example, it can be implemented through a computer program, such as application software, etc.; or it can be implemented as a medium storing relevant computer programs, such as a USB flash drive, a cloud disk, etc.; or it can be implemented through a physical device that integrates or installs relevant computer programs, such as a chip, etc.

[0101] In this example, when the system receives a data insertion request, the request carries instructions for inserting multiple sets of data to be inserted into the database. Specifically, a data insertion request refers to an instruction sent by a client (such as an application, user interface, or other system module) to a database, requesting that one or more sets of new data be inserted into a specified database. These data to be inserted include but are not limited to the user inputting and submitting data through a front-end interface (such as a web form, mobile application interface, etc.), triggering a data insertion request; when the application executes certain business logic, it needs to insert the generated data records into the database; in data migration or synchronization operations, the system will generate data insertion requests in batches to migrate data from one database to another or synchronize it to different data stores; certain scheduled tasks (such as logging, data backup, etc.) will periodically generate data insertion requests to record relevant data in the database, etc.

[0102] Accordingly, in order to process this data efficiently, the system will parse the received data insertion request. Through the internal parsing module, the information in the data insertion request is disassembled and converted into structured query language statements (SQL statements). This process involves operations such as identifying the request format, verifying the data type, and matching the SQL statement template.

[0103] After receiving the SQL statement, the system will reasonably divide the multiple groups of data to be inserted according to the preset maximum number of groups. The maximum number of groups is a key performance tuning parameter that determines the amount of data sent to the database each time. The system will divide the multiple groups of data to be inserted into several data sets, where the number of groups of data to be inserted in each data set is not greater than the maximum number of groups.

[0104] Correspondingly, after dividing the data sets, the system will merge the data to be inserted in each data set. This process is to integrate multiple groups of data to be inserted in the same data set, and then based on the SQL statement obtained by the previous analysis, the corresponding main SQL statement can be obtained to generate the complete SQL statement corresponding to the data set. Finally, the system sends the SQL statements corresponding to these data sets to the database. After receiving the statement, the database will perform the corresponding data insertion operation according to the content of the statement, and accurately store the data to be inserted into the specified database.

[0105] In the above example, by setting the maximum number of groups, and dividing multiple groups of data to be inserted into data sets according to the maximum number of groups, and then merging the data to be inserted in the data sets, generating statements corresponding to the data sets based on structured query language statements and sending them to the database, executing data insertion operations, and adopting array batch transmission, the execution efficiency of data insertion is improved, and the problem that the optimizer cannot be reused is solved.

[0106] Based on the above example, the structured query language statement includes an insert operation field, an insert position field, a keyword field and a data field; the device also includes:

[0107] A generating module, used for generating a list field according to the structure of a single group of data to be inserted in a data field;

[0108] The extraction module is used to extract the main statement from the structured query language statement; the main statement includes the insert operation field, the insert position field, the keyword field and the list field of the structured query language statement;

[0109] The third processing module is used to obtain the data volume of a single group of data to be inserted according to the list field, and set the maximum number of groups according to the data volume of the single group of data to be inserted.

[0110] Specifically, in database operations, SQL statements are the core tool for inserting data. A complete SQL statement usually includes an insert operation field (such as insert into), an insert position field (such as table name table_name), a keyword field (such as values), and a data field (specific data). Among them, the insert operation field represents the type of operation, that is, performing a data insert operation into the database; the insert position field represents the target location for data insertion, that is, the name of the table where the data will be inserted; the keyword field represents the introduction of specific data, that is, specifying that the data after the keyword field is the data to be inserted into the database; the data field represents the specific data, that is, multiple groups of data to be inserted.

[0111] Correspondingly, the list field is generated based on the structure of a single group of data to be inserted in the data field. The system generates a list field that defines the structure of each group of data. The list field specifies the column name of the data to be inserted, which is used to clarify the structure of each group of data and ensure that the data is inserted in the correct position.

[0112] After generating the list fields, the system will extract the main statement from the SQL statement. The main statement is the basic framework of the SQL statement, which defines the basic structure of the insert operation. The purpose of extracting the main statement is to reuse this structure in subsequent batch insert operations, thereby reducing the repeated construction of SQL statements, reducing semantic analysis, and improving data insertion efficiency. Among them, the main statement includes the insert operation field, insert position field, keyword field, and list field of the SQL statement.

[0113] Accordingly, the system calculates the data volume of a single group of data to be inserted based on the list fields. The data volume refers to the number of fields contained in each group of data.

[0114] Through the solution of this example, the SQL statement structure is simplified by generating list fields and extracting the main statement. Dynamically setting the maximum number of groups balances the database load, ensures efficient processing of batch insertion, reduces response time, and improves the execution efficiency of data insertion operations.

[0115] Based on the above examples, the extraction module is specifically used for:

[0116] Locate the position of inserting the operation field and the keyword field in the structured query language statement;

[0117] Based on the positions of the insert operation field and the keyword field, the field between the insert operation field and the keyword field in the structured query language statement is extracted, and the field is concatenated with the list field to obtain the main statement.

[0118] Specifically, when extracting the main statement from the SQL statement, you first need to locate the position of the insert operation field and the keyword field. The insert operation field is usually "insert into", which is a fixed part in the SQL statement used to identify the insert operation. The keyword field is "values", which is used to introduce specific data values. By parsing the SQL statement, you can find the position of these two fields. Locating the position of these two fields is a key step in extracting the main statement because it determines the range of fields that need to be extracted.

[0119] Accordingly, after locating the positions of the insert operation field and the keyword field, all fields between the insert operation field and the keyword field are extracted. Through the solution of this example, it can be ensured that each insert operation uses the same structure, reducing semantic analysis, thereby improving insertion efficiency and database performance.

[0120] Based on the above example, the third processing module is specifically used for:

[0121] Determine the value range of the data amount of a single group of data to be inserted; wherein different value ranges correspond to different ways of setting the maximum number of groups;

[0122] According to the value range of the data amount of a single group of data to be inserted, a corresponding setting method is adopted to set the maximum number of groups; wherein the value range and the maximum number of groups are negatively correlated.

[0123] In this example, when setting the maximum number of groups based on the amount of data to be inserted in a single group, you first need to determine the value range of the amount of data to be inserted in a single group. The amount of data refers to the number of fields contained in each group of data. Different value ranges correspond to different maximum number of group settings. The division of the value range can be determined based on actual business needs and database performance test results.

[0124] The division of these value ranges is based on a comprehensive consideration of database processing capabilities and network transmission efficiency, and is not limited here. Data in the small data range is usually inserted more efficiently, and the maximum number of groups can be appropriately increased; data in the large data range needs to reduce the maximum number of groups to avoid excessive pressure on the database.

[0125] According to the value range of the amount of data to be inserted in a single group, the corresponding setting method is used to determine the maximum number of groups. The value range and the maximum number of groups are negatively correlated, that is, the larger the amount of data, the smaller the maximum number of groups.

[0126] In one example, the third processing module is specifically configured to:

[0127] If the amount of data in a single group of data to be inserted is within the first value range, the preset first value is used as the maximum number of groups;

[0128] If the data volume of a single group of data to be inserted is within the second value range, the data volume of the single group of data to be inserted is normalized and the highest bit is calculated to obtain the maximum number of groups;

[0129] If the amount of data in a single group to be inserted is within the third value range, the preset second value is used as the maximum number of groups; wherein the critical value of the first value range is smaller than the second value range, the critical value of the second value range is smaller than the third value range, and the first value is greater than the second value.

[0130] The specific settings are as follows:

[0131] When the data volume is 0, the maximum number of groups is set to 1024; when the data volume is small (1 <= data volume <= 255), a larger maximum number of groups can be set, such as 128; when the data volume is medium (255 < data volume <= 16383), the data volume of a single group of data to be inserted is normalized to calculate the highest bit to obtain the maximum number of groups, that is, Integer.highestOneBit((Short.MAX_VALUE-1) / data volume). The specific calculation method is the same as the existing technology and no additional explanation is given here. When the data volume is large (data volume > 16383), the maximum number of groups is set to 1. For the specific maximum number of groups, please refer to Table 1.

[0132] The above is only an example of setting the maximum number of groups. In actual applications, the maximum number of groups can be set according to actual conditions, as long as the critical value of the first value range is less than the second value range, the critical value of the second value range is less than the third value range, and the first value is greater than the second value. Through the solution of this example, the system can dynamically adjust the maximum number of groups according to the amount of data to be inserted in a single group, thereby maximizing the efficiency of batch data insertion while ensuring database performance.

[0133] In one example, the device further includes: a fourth processing module, configured to determine a maximum number of groups based on a user-defined number of groups.

[0134] Specifically, the maximum number of groups can be implemented based on the user-defined number of groups. Users can dynamically adjust the maximum number of groups based on specific business needs and the actual performance of the database. To achieve this function, the system provides a user interface that allows users to enter or select the desired maximum number of groups. After the system receives the number of groups entered by the user, it will apply it as a parameter to subsequent data insertion operations. When performing batch data insertion, the system will divide the data to be inserted into multiple data sets based on this user-defined maximum number of groups, and the number of data groups in each set does not exceed the maximum number of groups set by the user. Through the solution of this example, users can flexibly control the efficiency of data insertion and resource usage according to different business scenarios and performance requirements, thereby achieving optimal system performance.

[0135] In the data insertion device provided by this embodiment, the first processing module first receives a data insertion request, and parses the received data insertion request to obtain a structured query language statement; the data insertion request is used to request to insert multiple groups of data to be inserted into the database; then the division module divides the multiple groups of data to be inserted into data sets according to the maximum number of groups; wherein the number of groups of data to be inserted in each data set is not greater than the maximum number of groups; finally, the second processing module merges the data to be inserted in each data set, and obtains the statement corresponding to the data set based on the structured query language statement; and then sends the statement corresponding to the data set to the database, so that the database executes data insertion based on the statement. The scheme of the present application, by setting the maximum number of groups, divides multiple groups of data to be inserted into multiple data sets, and generates SQL statements corresponding to each data set, which significantly reduces the number of SQL statements and the corresponding number of transmissions. The database parsing and optimization overhead is reduced, thereby improving the execution efficiency of data insertion.

[0136] Embodiment 4

[0137] Figure 5 A schematic diagram of the structure of an electronic device provided in Embodiment 4 of the present application, such as Figure 5 As shown, the electronic device includes:

[0138] The electronic device includes a processor 291 and a memory 292; it may also include a communication interface 293 and a bus 294. The processor 291, the memory 292, and the communication interface 293 may communicate with each other through the bus 294. The communication interface 293 may be used for information transmission. The processor 291 may call the logic instructions in the memory 292 to execute the method of the above example.

[0139] In addition, the logic instructions in the above-mentioned memory 292 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.

[0140] The memory 292 is a computer-readable storage medium that can be used to store software programs and computer executable programs, such as program instructions / modules corresponding to the methods in the embodiments of the present application. The processor 291 executes functional applications and data processing by running the software programs, instructions, and modules stored in the memory 292, that is, implementing the methods in the above method examples.

[0141] The memory 292 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and an application required for at least one function; the data storage area may store data created according to the use of the terminal device, etc. In addition, the memory 292 may include a high-speed random access memory and may also include a non-volatile memory.

[0142] An embodiment of the present application further provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the method in any embodiment.

[0143] The embodiments of the present application also provide a computer program product, which implements the method in any embodiment when the computer program is executed by a processor.

[0144] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the following claims.

[0145] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A data insertion method, characterized in that: include: Receiving a data insertion request, and parsing the data insertion request to obtain a structured query language statement; the data insertion request is used to request to insert multiple groups of data to be inserted into the database; According to the maximum number of groups, the multiple groups of data to be inserted are divided into data sets; wherein the number of groups of data to be inserted in each data set is not greater than the maximum number of groups; The data to be inserted in each data set are merged, and a statement corresponding to the data set is obtained based on the structured query language statement; the statement corresponding to the data set is sent to the database, so that the database executes data insertion based on the statement.

2. The method according to claim 1, characterized in that The structured query language statement includes an insert operation field, an insert position field, a keyword field and a data field; the method further includes: Generate a list field according to the structure of a single group of data to be inserted in the data field; Extracting a main sentence from the structured query language statement; the main sentence includes an insert operation field, an insert position field, a keyword field and the list field of the structured query language statement; According to the list field, the data volume of a single group of data to be inserted is obtained, and the maximum number of groups is set according to the data volume of the single group of data to be inserted.

3. The method according to claim 2, characterized in that The extracting the main statement from the structured query language statement comprises: Locating the positions of the insert operation field and the keyword field in the structured query language statement; Based on the positions of the insert operation field and the keyword field, the field between the insert operation field and the keyword field in the structured query language statement is extracted, and the field is concatenated with the list field to obtain the main statement.

4. The method according to claim 1, characterized in that: The method further comprises: The maximum number of groups is determined according to the user-defined number of groups.

5. The method according to claim 2, characterized in that: The step of setting the maximum number of groups according to the amount of data of the single group to be inserted comprises: Determine a value range of the amount of data of the single group of data to be inserted; wherein different value ranges correspond to different ways of setting the maximum number of groups; According to the value range of the data amount of the single group of data to be inserted, the maximum number of groups is set by adopting a corresponding setting method; wherein the value range and the maximum number of groups are negatively correlated.

6. The method according to claim 5, characterized in that The setting of the maximum number of groups by adopting a corresponding setting method according to a value range of the amount of the single group of data to be inserted includes: If the amount of the single group of data to be inserted is within the first value range, the preset first value is used as the maximum number of groups; If the data amount of the single group of data to be inserted is within the second value range, normalizing the highest bit of the data amount of the single group of data to be inserted to obtain the maximum number of groups; If the amount of data in the single group to be inserted is within the third value range, the preset second value is used as the maximum number of groups; wherein the critical value of the first value range is smaller than the second value range, the critical value of the second value range is smaller than the third value range, and the first value is greater than the second value.

7. A data insertion device, characterized in that: include: A first processing module is used to receive a data insertion request and parse the data insertion request to obtain a structured query language statement; the data insertion request is used to request to insert multiple groups of data to be inserted into the database; A division module, used for dividing the plurality of groups of data to be inserted into data sets according to the maximum number of groups; wherein the number of groups of data to be inserted in each data set is not greater than the maximum number of groups; The second processing module is used to merge the data to be inserted in each data set, and obtain the statement corresponding to the data set based on the structured query language statement; send the statement corresponding to the data set to the database, so that the database executes data insertion based on the statement.

8. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 6 when executed by a processor.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.