Table data segmentation method and equipment based on database and medium

By obtaining the data types and encoded characters of the table data to be divided in the database, and combining with the preset wide table subtable strategy tree, the problem of cumbersome processing of wide table data in the existing technology is solved, the data segmentation efficiency and accuracy are improved, and the dynamic changes of business are adapted to.

CN119988381APending Publication Date: 2025-05-13HIGHGO SOFTWARE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510078767.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art, application layer adaptation is required when processing wide table data, resulting in cumbersome and inefficient in database table data segmentation, making it difficult to quickly respond to business dynamic changes.

Method used

By obtaining the data type of the table to be divided, determining the cutting position, and matching the encoded characters in the preset special encoding library, adding it to the cutting position. Determine the table sub-table strategy based on the encoded characters and the cutting position, introduce the data into different leaf nodes of the preset wide table sub-table strategy tree, process it and insert it into the target database.

Benefits of technology

The efficiency and accuracy of table data segmentation have been improved, the table partitioning process has been standardized and ordered, and the data is dynamically adjusted to adapt to business changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988381A_ABST
    Figure CN119988381A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a table data segmentation method and equipment based on a database and a medium, belongs to the technical field of data processing, and solves the problem that in the prior art, when database table data segmentation is carried out, an application layer needs to be adapted, so that dynamic changes of services are difficult to respond quickly. Obtaining to-be-segmented table data, determining a data type corresponding to the to-be-segmented table data, and determining a segmentation position of the to-be-segmented table data based on the data type; based on the data type, determining a corresponding coded character in a preset special coding library, and adding the coded character to the cutting position; determining a current table division strategy according to the coded characters and the cutting positions, and introducing the table data to be segmented into different leaf nodes of a preset wide table division strategy tree based on the table division strategy; and processing the sub-table data corresponding to the different coded characters through different leaf node processes so as to insert the sub-table data into a target table of a target database.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a table data segmentation method, device and medium based on a database. Background Art

[0002] In a database, a wide table refers to a table with hundreds or even thousands of table attributes (a table has hundreds or even thousands of columns), while a small table refers to a table with a few or dozens of table attributes. According to the current data development trend, in a specific industry, the table attributes (number of columns) of a single table in the database may sometimes reach hundreds or thousands. After a certain period of development, the amount of data in a single wide table will be too large to affect performance.

[0003] In the prior art, when processing wide table data, it usually goes through the process of data acquisition, data scanning and processing, data column position counting and segmentation, and data storage. However, after the wide table is split into small tables, the application layer usually needs to be adapted to achieve the purpose of directly inserting subsequent data into the small table, which makes the database table data segmentation process more cumbersome and inefficient, making it difficult to quickly respond to dynamic changes in business. Summary of the invention

[0004] The embodiments of the present application provide a database-based table data segmentation method, device and medium for solving the following technical problems: In the prior art, when processing wide table data, the application layer usually needs to be adapted, which makes the database table data segmentation process more cumbersome and inefficient, making it difficult to quickly respond to dynamic changes in the business.

[0005] The present application embodiment adopts the following technical solutions:

[0006] The embodiment of the present application provides a table data segmentation method based on a database. The method includes obtaining the table data to be segmented, and determining the data type corresponding to the table data to be segmented, and determining the cutting position of the table data to be segmented based on the data type; determining the corresponding coding characters in the preset special coding library based on the data type, and adding the coding characters to the cutting position; determining the current table splitting strategy based on the coding characters and the cutting position, and introducing the table data to be segmented into different leaf nodes of the preset wide table splitting strategy tree based on the table splitting strategy; processing the table splitting data corresponding to different coding characters through different leaf nodes, so as to insert the table splitting data into the target table of the target database.

[0007] When acquiring data, the embodiment of the present application adds coding characters to the data stream according to the table partitioning strategy to identify which sub-table the subsequent data belongs to, thereby improving the efficiency and accuracy of table data segmentation. In processing the data stream, the subsequent data is directly introduced into different small table data processing streams according to special coding characters, thereby improving the efficiency of data flow. A structured framework is provided by presetting a wide table partitioning strategy tree, and data is introduced into different leaf nodes according to established strategies, so that the table partitioning process is standardized and ordered. The embodiment of the present application solves the problem of wide table partitioning and adaptation, and provides different data cutting solutions to achieve the purpose of dynamically adjusting the subsequent data entry into the warehouse.

[0008] In one implementation of the present application, the data type corresponding to the data in the table to be split is determined, and the cutting position of the data in the table to be split is determined based on the data type, specifically including: determining the data to be identified corresponding to each column in the data in the table to be split, and performing feature extraction on the data to be identified; converting the data features corresponding to each extracted column into a feature vector, and inputting the feature vector into a preset data classification model to determine the data type corresponding to the feature vector through the preset data classification model; determining the column sequence number corresponding to the data to be identified, dividing the data to be identified corresponding to the same data type into the same data set based on the column sequence number, and determining multiple cutting positions corresponding to the current data type based on the serial number in the same data set.

[0009] In one implementation of the present application, based on the data type, the corresponding coding characters are determined in a preset special coding library, and the coding characters are added to the cutting position, specifically including: based on the data type, matching is performed in the preset special coding library to determine the coding character type corresponding to the data type; wherein the preset special coding library includes multiple data types, coding character types corresponding to the multiple data types, and coding character sets corresponding to the coding character types; in the coding character sets corresponding to the coding character types, reference coding characters that are not used in the current task are determined; based on the number of corresponding serial numbers in the data set, the same number of required coding characters are screened out from the reference coding characters, and the required coding characters are matched with the serial numbers; based on the matching results, the required coding characters are added to the cutting position corresponding to the serial number.

[0010] In one implementation of the present application, based on the matching result, the required coded characters are added to the cutting position corresponding to the serial number, specifically including: based on a preset callback function, determining the cutting position corresponding to the serial number; and, based on the preset callback function, adding the required coded characters to the cutting position corresponding to the serial number.

[0011] In one implementation of the present application, a current table splitting strategy is determined based on the coded characters and the cutting position, specifically including: receiving a table splitting instruction sent by a user; wherein the table splitting instruction includes a table splitting scheme number or a table splitting requirement; in the case where a table splitting scheme number exists in the table splitting instruction, a corresponding table splitting strategy is determined in a preset table splitting scheme library based on the table splitting scheme number; wherein the preset table splitting scheme library includes multiple table splitting scheme numbers, and also includes table splitting strategies corresponding to the multiple table splitting scheme numbers; in the case where a table splitting requirement exists in the table splitting instruction, multiple historical table splitting strategies are determined in a historical table splitting task library based on the table splitting requirement, so as to determine the current table splitting strategy based on the multiple historical table splitting strategies.

[0012] In one implementation of the present application, multiple historical sub-table strategies are determined in a historical sub-table task library based on the sub-table requirements, and the current sub-table strategy is determined based on the multiple historical sub-table strategies, specifically including: based on the sub-table requirements, multiple reference sub-table requirements are determined in the historical sub-table task library to construct a sub-table requirement set; based on the sub-table requirement set, a query is performed in the historical sub-table task library to determine a reference sub-table strategy set corresponding to the sub-table requirement set; based on the reference segmentation table data corresponding to each reference sub-table strategy in the reference sub-table strategy set, a reference segmentation data type, a reference sub-table segmentation quantity, and a reference sub-table are determined. corresponding data volume; based on the reference segmentation data type, determining in the reference segmentation table data the first candidate segmentation table data whose similarity value with the data type of the table data to be segmented is greater than a preset similarity value threshold; based on the reference sub-table segmentation number, determining in the first candidate segmentation table data the second candidate segmentation table data that belongs to the preset sub-table segmentation number range; performing weighted processing on the reference segmentation data type, the reference sub-table segmentation number and the data volume corresponding to the reference sub-table respectively corresponding to the reference segmentation table data in the second candidate segmentation table data, so as to determine the current segmentation strategy based on the weighted processing result.

[0013] In one implementation of the present application, the data of the table to be split is introduced into different leaf nodes of a preset wide table splitting strategy tree based on the table splitting strategy, specifically including: based on the table splitting strategy, splitting the data of the table to be split into multiple sub-table split data; traversing the preset wide table splitting strategy tree based on the multiple sub-table split data to determine the leaf nodes corresponding to each sub-table split data in the preset wide table splitting strategy tree; based on the position of the corresponding leaf node, moving each sub-table split data to a corresponding storage location.

[0014] In one implementation of the present application, sub-table data corresponding to different coded characters are processed through different leaf nodes to insert the sub-table data into a target table of a target database, specifically including: establishing a connection with the target database; obtaining database parameters corresponding to the target database; wherein the database parameters include at least the target database structure data, the data volume of the target database, and the performance parameters of the target database; determining the data insertion volume based on the database parameters through a preset machine learning algorithm; based on the data insertion volume, splitting the processed sub-table data into multiple batches of data, and inserting the multiple batches of data into the target table in sequence according to the structure and field order of the target table.

[0015] An embodiment of the present application provides a table data segmentation device based on a database, comprising: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can: obtain table data to be segmented, and determine the data type corresponding to the table data to be segmented, and determine the cutting position of the table data to be segmented based on the data type; determine the corresponding coding character in a preset special coding library based on the data type, and add the coding character to the cutting position; determine the current table segmentation strategy according to the coding character and the cutting position, and introduce the table data to be segmented into different leaf nodes of a preset wide table segmentation strategy tree based on the table segmentation strategy; process the segmented table data corresponding to different coding characters through different leaf nodes, so as to insert the segmented table data into a target table of a target database.

[0016] A non-volatile computer storage medium provided by an embodiment of the present application stores computer executable instructions, wherein the computer executable instructions are configured to: obtain table data to be segmented, and determine a data type corresponding to the table data to be segmented, so as to determine a cutting position of the table data to be segmented based on the data type; determine a corresponding coding character in a preset special coding library based on the data type, and add the coding character to the cutting position; determine a current table segmentation strategy based on the coding character and the cutting position, and introduce the table data to be segmented into different leaf nodes of a preset wide table segmentation strategy tree based on the table segmentation strategy; process the table segmentation data corresponding to different coding characters through different leaf node processes, so as to insert the table segmentation data into a target table of a target database.

[0017] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects: When acquiring data, the embodiments of the present application add coding characters to the data stream according to the table partitioning strategy to identify which sub-table the subsequent data belongs to, thereby improving the efficiency and accuracy of table data segmentation. In processing the data stream, the subsequent data is directly introduced into different small table data processing streams according to special coding characters, thereby improving the efficiency of data flow. A structured framework is provided by presetting a wide table partitioning strategy tree, and data is introduced into different leaf nodes according to established strategies, so that the table partitioning process is standardized and ordered. The embodiments of the present application solve the problems of wide table partitioning and adaptation, and provide different data cutting solutions to achieve the purpose of dynamically adjusting the subsequent data entry into the warehouse. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative labor. In the drawings:

[0019] Figure 1 A flowchart of a table data segmentation method based on a database provided in an embodiment of the present application;

[0020] Figure 2 A framework diagram of a table data segmentation process based on a database provided in an embodiment of the present application;

[0021] Figure 3 A schematic diagram of the structure of a database-based table data segmentation device provided in an embodiment of the present application.

[0022] Reference numerals:

[0023] 200 is a table data segmentation device based on a database, 201 is a processor, and 202 is a memory. DETAILED DESCRIPTION

[0024] The embodiments of the present application provide a table data segmentation method, device and medium based on a database.

[0025] In order to enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of them. Based on the embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this application.

[0026] Figure 1 A flowchart of a table data segmentation method based on a database is provided in an embodiment of the present application, such as Figure 1 As shown, the table data segmentation method based on the database includes the following steps:

[0027] S101 , obtaining table data to be segmented, and determining a data type corresponding to the table data to be segmented, so as to determine a cutting position of the table data to be segmented based on the data type.

[0028] In one implementation of the present application, the data to be identified corresponding to each column in the table data to be segmented is determined, and features are extracted for the data to be identified. The data features corresponding to each extracted column are converted into feature vectors, and the feature vectors are input into a preset data classification model to determine the data type corresponding to the feature vectors through the preset data classification model. The column sequence number corresponding to the data to be identified is determined, and the data to be identified corresponding to the same data type is divided into the same data set based on the column sequence number, and multiple cutting positions corresponding to the current data type are determined based on the sequence number in the same data set.

[0029] Specifically, after obtaining the data in the table to be segmented, each column of data needs to be analyzed separately, where the data to be identified is the specific data item in each column. For each column of data to be identified, feature extraction is performed. For example, for numerical data, statistical features such as mean, variance, maximum value, minimum value, etc. can be extracted; for string data, text features such as word frequency, TF-IDF (word frequency-inverse document frequency), text length, etc. can be extracted; for date / time data, time features such as year, month, day, week, hour, etc. can be extracted.

[0030] Further, the extracted data features of each column are converted into a feature vector, and the feature vector is input into a preset data classification model to determine the data type corresponding to the feature vector through the preset data classification model. Among them, the preset data classification model in the embodiment of the present application is a support vector machine, and other data classification models can also be selected according to needs, and the embodiment of the present application does not limit this.

[0031] Furthermore, the data type corresponding to the feature vector is determined by a preset data classification model, and the column number corresponding to the data to be identified, that is, the position of each column of data in the table, is determined, and the data to be identified corresponding to the same data type is divided into the same data set based on the column number. In the same data set, multiple cutting positions corresponding to the current data type are determined based on the sequence number, and the data in the table to be divided is cut into multiple sub-tables based on the multiple cutting positions.

[0032] In one implementation of the present application, based on the data type, a match is performed in a preset special coding library to determine the coded character type corresponding to the data type; wherein the preset special coding library includes multiple data types, coded character types corresponding to the multiple data types, and coded character sets corresponding to the coded character types. In the coded character set corresponding to the coded character type, a reference coded character that is not used in the current task is determined. Based on the number of corresponding serial numbers in the data set, the same number of required coded characters is screened out from the reference coded characters, and the required coded characters are matched with the serial numbers. Based on the matching result, the required coded characters are added to the cutting position corresponding to the serial number.

[0033] Specifically, the embodiment of the present application is provided with a preset special coding library, wherein the preset special coding library is a database containing multiple data types, each data type corresponds to a coded character type, and each coded character type corresponds to a coded character set. Based on the data type of the determined data set, a query is performed in the preset special coding library, and for each data set, it is determined which coded characters have not been used in the current task, which usually requires tracking the coded characters that have been allocated or used, and excluding these characters from the coded character set.

[0034] Further, based on the number of corresponding serial numbers in the data set, the same number of required coded characters are selected from the reference coded characters, and then the selected coded characters are matched one-to-one with the serial numbers in the data set, that is, each serial number has a corresponding coded character. The matched coded characters are added before or after the determined cutting position.

[0035] S102: Based on the data type, determine the corresponding coding characters in the preset special coding library, and add the coding characters to the cutting position.

[0036] In one implementation of the present application, based on a preset callback function, a cutting position corresponding to the serial number is determined, and based on the preset callback function, a required coded character is added to the cutting position corresponding to the serial number.

[0037] Specifically, the preset callback function is a set of predefined functions that can usually be called under specific events or conditions. In the embodiment of the present application, the preset callback function is used to determine the cutting position corresponding to the serial number. That is, when a given condition or event occurs, for example, a specific serial number in the data set is processed, the callback function will be triggered and calculate the position where the serial number should be cut. After the cutting position is determined, the required coded characters are added to the cutting position through the preset callback function.

[0038] S103: Determine the current table splitting strategy according to the coded characters and the cutting position, and introduce the table data to be split into different leaf nodes of the preset wide table table splitting strategy tree based on the table splitting strategy.

[0039] In one implementation of the present application, a table sharding instruction sent by a user is received; wherein the table sharding instruction includes a table sharding scheme number or a table sharding requirement. In the case where a table sharding scheme number exists in the table sharding instruction, a corresponding table sharding strategy is determined in a preset table sharding scheme library based on the table sharding scheme number; wherein the preset table sharding scheme library includes a plurality of table sharding scheme numbers, and also includes table sharding strategies corresponding to a plurality of table sharding scheme numbers. In the case where a table sharding requirement exists in the table sharding instruction, a plurality of historical table sharding strategies are determined in a historical table sharding task library based on the table sharding requirement, so as to determine a current table sharding strategy based on the plurality of historical table sharding strategies.

[0040] Specifically, the user sends a table sharding instruction through a web page or a command line tool, and the table sharding instruction contains the table sharding scheme number or the specific table sharding requirement that the user wants to execute. When the table sharding instruction contains the table sharding scheme number, the system will search for the table sharding strategy corresponding to the number in the preset table sharding scheme library, where the preset table sharding scheme library is a database or data structure that stores multiple table sharding scheme numbers and their corresponding table sharding strategies. Each table sharding strategy contains table sharding rules, such as by date, by ID, etc., target table name, number of partitions and other detailed information. The system matches the table sharding scheme number and finds the corresponding table sharding strategy in the preset table sharding scheme library. If a matching table sharding strategy is found, the strategy is used for subsequent table sharding operations.

[0041] Furthermore, when the table partitioning instruction contains specific table partitioning requirements, the system will search for historical table partitioning strategies similar to these requirements in the historical table partitioning task library, wherein the historical table partitioning task library is a database or data structure that stores previously executed table partitioning tasks and their corresponding table partitioning strategies. According to the historical table partitioning strategies, the table partitioning strategy that should be used currently is determined in combination with the current table partitioning requirements. Specifically, the system first parses the table partitioning requirements and extracts key information, such as data volume, query mode, etc., and then searches for historical table partitioning strategies similar to these key information in the historical table partitioning task library. By comparing and analyzing these historical strategies, the system determines a table partitioning strategy that best suits the current requirements.

[0042] In one implementation of the present application, based on the sub-table requirements, multiple reference sub-table requirements are determined in the historical sub-table task library to construct a sub-table requirement set. Based on the sub-table requirement set, a query is performed in the historical sub-table task library to determine the reference sub-table strategy set corresponding to the sub-table requirement set. Based on the reference segmentation table data corresponding to each reference sub-table strategy in the reference sub-table strategy set, the reference segmentation data type, the reference sub-table segmentation number, and the amount of data corresponding to the reference sub-table are determined. Based on the reference segmentation data type, in the reference segmentation table data, the first candidate segmentation table data whose similarity value with the data type of the table data to be segmented is greater than a preset similarity value threshold is determined. Based on the reference sub-table segmentation number, in the first candidate segmentation table data, the second candidate segmentation table data belonging to the preset sub-table segmentation number range is determined. The reference segmentation data type, the reference sub-table segmentation number, and the amount of data corresponding to the reference sub-table corresponding to the reference segmentation table data in the second candidate segmentation table data are weighted, so as to determine the current segmentation strategy based on the weighted processing result.

[0043] Specifically, a table sharding requirement submitted by a user is received, which includes information such as the data type of the table to be split, the expected number of sub-tables, the expected data volume range of each sub-table, etc. According to the table sharding requirement submitted by the user, multiple similar reference table sharding requirements are searched in the historical table sharding task library to construct a table sharding requirement set.

[0044] Furthermore, for each reference sub-table requirement in the sub-table requirement set, the corresponding reference sub-table strategy is queried in the historical sub-table task library, and these reference sub-table strategies include strategies adopted by previously executed sub-table tasks similar to the current requirement. For each reference sub-table strategy, key information of the corresponding reference segmentation table data is extracted, including the reference segmentation data type, the reference sub-table segmentation number, and the amount of data corresponding to the reference sub-table. The similarity values ​​between the data types of the table data to be segmented and each reference segmentation table data are compared, wherein the similarity values ​​can be calculated by comparing factors such as the structure of the data type, the number of fields, and the field type. When the similarity value is greater than a preset similarity value threshold, the reference segmentation table data is selected as the first candidate segmentation table data. Among the first candidate segmentation table data, those data whose reference sub-table segmentation number falls within the preset sub-table segmentation number range are further screened, and these data are selected as the second candidate segmentation table data.

[0045] Furthermore, for each reference partition table data in the second candidate partition table data, the system performs weighted processing according to its reference partition data type, reference sub-table partition quantity, and reference sub-table corresponding data volume. The weighted processing can assign different weights based on the importance of these factors, and according to the weighted processing results, a comprehensive evaluation is performed to obtain the most suitable partition table strategy at present.

[0046] In one implementation of the present application, based on the table partitioning strategy, the table data to be partitioned is partitioned into a plurality of sub-table partitioned data. Based on the plurality of sub-table partitioned data, a preset wide table partitioning strategy tree is traversed to determine the leaf nodes corresponding to each sub-table partitioned data in the preset wide table partitioning strategy tree. Based on the position of the corresponding leaf node, each sub-table partitioned data is moved to a corresponding storage location.

[0047] Specifically, according to the selected table partitioning strategy, the original large table data is divided into multiple smaller sub-tables. The pre-set wide table partitioning strategy tree in the embodiment of the present application is a logical structure used to define how to distribute the data to different target tables according to its different characteristics. Each node of the tree represents a decision point, which determines whether the data should be moved to the left subtree or the right subtree according to the value of a certain attribute of the data.

[0048] Further, the preset wide table partitioning strategy tree is traversed, that is, the decision rules of the tree are applied to each sub-table partitioning data until the leaf node of the tree is reached, and each leaf node represents the final data storage location. By traversing the strategy tree, each sub-table partitioning data will eventually be assigned to a leaf node. This leaf node identifies the specific location or table where the data should be stored. Once the leaf node corresponding to each sub-table partitioning data is determined, it is necessary to move the data from the original table or temporary storage location to the final storage location.

[0049] S104: Process the sub-table data corresponding to different coded characters through different leaf nodes to insert the sub-table data into the target table of the target database.

[0050] In one implementation of the present application, a connection is established with a target database to obtain database parameters corresponding to the target database; wherein the database parameters include at least target database structure data, target database data volume, and target database performance parameters. The amount of data to be inserted is determined based on the database parameters by presetting a machine learning algorithm. Based on the amount of data to be inserted, the processed sub-table data is split into multiple batches of data, and the multiple batches of data are sequentially inserted into the target table according to the structure and field order of the target table.

[0051] Specifically, a connection is established with the target database to obtain the target database structure data of the target table, such as table structure, field type, index, etc.; the data volume of the target database is obtained, and the performance parameters of the target database, such as CPU usage, memory usage, disk I / O, etc. are obtained.

[0052] Furthermore, the embodiment of the present application first uses historical data to train a machine learning model, which can predict the optimal amount of data to be inserted based on the current database parameters. In the actual data insertion process, the trained model is used to determine the amount of data to be inserted each time. According to the determined amount of data to be inserted, the sub-table data is divided into multiple smaller data sets. Large amounts of data are divided into smaller batches so that each insertion operation can be performed efficiently, while reducing performance issues caused by inserting too much data at a time. Ensure that the data is inserted into the target table in the correct format and order, thereby maintaining the integrity and consistency of the data.

[0053] Figure 2 A table data segmentation process framework diagram based on a database is provided in an embodiment of the present application, such as Figure 2 As shown:

[0054] Wide table access layer: It is the data input interface of the wide table processing module. Its external connection can be upper-layer applications, source databases, data files, etc. But they have one thing in common, that is, the data to be processed is wide table data;

[0055] Table partitioning strategy management module: It uses a decision tree to construct how to split a wide table into multiple small tables.

[0056] This includes how many small tables are included, what special encoding characters are used to identify the data in the table, how many small table processes are started at the bottom layer, etc.

[0057] Data stream processing layer: adapts to different data cutting schemes and does not require data conversion for specific data content.

[0058] The details are as follows:

[0059] Upper-layer application: callback policy module configuration, directly adding special coded characters in the data stream;

[0060] Source database: When exporting data, add special coding characters directly to the data stream, or in the acquired data stream, process each line of data according to the processing strategy, add special coding characters at the beginning of the data position that needs to be cut, and then pass it to the lower-level processing module.

[0061] Data file: Read the data file, process each row of data according to the data partitioning strategy, add special encoding characters, and then introduce the data into the data stream.

[0062] Data stream processing module: According to the table partitioning strategy, different special encoding characters and subsequent data are introduced into the strategy processing node, that is, the specific small table data processing process, for data insertion processing.

[0063] Figure 3 The following is a schematic diagram of a structure of a table data segmentation device based on a database provided in an embodiment of the present application. Figure 3 As shown, a table data segmentation device 200 based on a database includes: at least one processor 201; and a memory 202 that is communicatively connected to the at least one processor 201; wherein the memory 202 stores instructions that can be executed by the at least one processor 201, and the instructions are executed by the at least one processor 201 so that the at least one processor 201 can: obtain table data to be segmented, and determine the data type corresponding to the table data to be segmented, and determine the cutting position of the table data to be segmented based on the data type; determine the corresponding coding character in a preset special coding library based on the data type, and add the coding character to the cutting position; determine the current table segmentation strategy based on the coding character and the cutting position, and introduce the table data to be segmented into different leaf nodes of a preset wide table segmentation strategy tree based on the table segmentation strategy; process the segmented table data corresponding to different coding characters through different leaf nodes, so as to insert the segmented table data into the target table of the target database.

[0064] A non-volatile computer storage medium provided by an embodiment of the present application stores computer executable instructions, wherein the computer executable instructions are configured to: obtain table data to be segmented, and determine a data type corresponding to the table data to be segmented, so as to determine a cutting position of the table data to be segmented based on the data type; determine a corresponding coding character in a preset special coding library based on the data type, and add the coding character to the cutting position; determine a current table segmentation strategy based on the coding character and the cutting position, and introduce the table data to be segmented into different leaf nodes of a preset wide table segmentation strategy tree based on the table segmentation strategy; process the table segmentation data corresponding to different coding characters through different leaf node processes, so as to insert the table segmentation data into a target table of a target database.

[0065] Each embodiment in this application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device, equipment, and non-volatile computer storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0066] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the embodiments of the present application may have various changes and variations. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application.

Claims

1. A table data segmentation method based on a database, characterized in that: The method comprises: Acquire the table data to be segmented, and determine the data type corresponding to the table data to be segmented, so as to determine the cutting position of the table data to be segmented based on the data type; Based on the data type, determining a corresponding coding character in a preset special coding library, and adding the coding character to the cutting position; Determine a current table partitioning strategy according to the coded character and the cutting position, and introduce the table data to be partitioned into different leaf nodes of a preset wide table table partitioning strategy tree based on the table partitioning strategy; The sub-table data corresponding to the different coded characters are processed through the different leaf nodes to insert the sub-table data into the target table of the target database.

2. A method for segmenting table data based on a database according to claim 1, characterized in that: The step of determining the data type corresponding to the table data to be segmented, and determining the cutting position of the table data to be segmented based on the data type, specifically includes: Determine the data to be identified corresponding to each column in the table data to be segmented, and perform feature extraction on the data to be identified; Convert the data features corresponding to each extracted column into a feature vector, and input the feature vector into a preset data classification model to determine the data type corresponding to the feature vector through the preset data classification model; Determine the column serial number corresponding to the data to be identified, divide the data to be identified corresponding to the same data type into the same data set based on the column serial number, and determine multiple cutting positions corresponding to the current data type based on the serial number in the same data set.

3. A method for segmenting table data based on a database according to claim 2, characterized in that: The step of determining a corresponding coding character in a preset special coding library based on the data type and adding the coding character to the cutting position specifically includes: Based on the data type, matching is performed in the preset special coding library to determine the coded character type corresponding to the data type; wherein the preset special coding library includes multiple data types, coded character types corresponding to multiple data types respectively, and also includes a coded character set corresponding to the coded character type; Determining, in the coded character set corresponding to the coded character type, a reference coded character that is not used in the current task; Based on the number of corresponding serial numbers in the data set, selecting the same number of required coded characters from the reference coded characters, and matching the required coded characters with the serial numbers; Based on the matching result, the required encoding character is added to the cutting position corresponding to the serial number.

4. The method for segmenting table data based on a database according to claim 3, characterized in that: The adding the required coded characters to the cutting position corresponding to the serial number based on the matching result specifically includes: Based on a preset callback function, determining a cutting position corresponding to the serial number; And, based on the preset callback function, the required encoding character is added to the cutting position corresponding to the serial number.

5. The method for segmenting table data based on a database according to claim 1, characterized in that: The determining of the current table partitioning strategy according to the coded character and the cutting position specifically includes: Receiving a table partitioning instruction sent by a user; wherein the table partitioning instruction includes a table partitioning scheme number or a table partitioning requirement; In the case where the table subdivision scheme number exists in the table subdivision instruction, a corresponding table subdivision strategy is determined in a preset table subdivision scheme library based on the table subdivision scheme number; wherein the preset table subdivision scheme library includes a plurality of table subdivision scheme numbers, and also includes a plurality of table subdivision strategies corresponding to the table subdivision scheme numbers respectively; In the case that the table sharding requirement exists in the table sharding instruction, a plurality of historical table sharding strategies are determined in a historical table sharding task library based on the table sharding requirement, so as to determine the current table sharding strategy based on the plurality of historical table sharding strategies.

6. A method for segmenting table data based on a database according to claim 5, characterized in that: The determining of a plurality of historical table partitioning strategies in a historical table partitioning task library based on the table partitioning requirements, and determining the current table partitioning strategy based on the plurality of historical table partitioning strategies, specifically includes: Based on the sub-table requirement, a plurality of reference sub-table requirements are determined in the historical sub-table task library to construct a sub-table requirement set; Based on the table sharding requirement set, query in the historical table sharding task library to determine the reference table sharding strategy set corresponding to the table sharding requirement set; Determine the reference segmentation data type, the reference sub-table segmentation number, and the data volume corresponding to the reference sub-table based on the reference segmentation table data corresponding to each reference sub-table strategy in the reference sub-table strategy set; Based on the reference segmentation data type, determining first candidate segmentation table data in the reference segmentation table data, the first candidate segmentation table data having a similarity value greater than a preset similarity value threshold value with the data type of the table data to be segmented; Based on the reference sub-table segmentation quantity, determining second candidate segmentation table data belonging to a preset sub-table segmentation quantity range in the first candidate segmentation table data; The reference segmentation data type, the reference sub-table segmentation quantity and the data volume corresponding to the reference segmentation table data in the second candidate segmentation table data are weighted to determine the current segmentation strategy based on the weighted processing result.

7. The method for segmenting table data based on a database according to claim 1, characterized in that: The step of introducing the table data to be split into different leaf nodes of the preset wide table splitting strategy tree based on the table splitting strategy specifically includes: Based on the table partitioning strategy, the table data to be partitioned is partitioned into a plurality of sub-table partitioned data; Traversing the preset wide table partitioning strategy tree based on the plurality of sub-table partitioning data to determine leaf nodes corresponding to the sub-table partitioning data in the preset wide table partitioning strategy tree; Based on the position of the corresponding leaf node, each of the sub-table segmented data is moved to a corresponding storage position.

8. The method for segmenting table data based on a database according to claim 1, characterized in that: The processing of the sub-table data corresponding to the different coded characters through the different leaf nodes to insert the sub-table data into the target table of the target database specifically includes: Establishing a connection with the target database; Obtaining database parameters corresponding to the target database; wherein the database parameters at least include target database structure data, target database data volume, and target database performance parameters; Determine the amount of data to be inserted based on the database parameters by presetting a machine learning algorithm; Based on the data insertion amount, the processed sub-table data is split into a plurality of batch data, and the plurality of batch data are sequentially inserted into the target table according to the structure and field sequence of the target table.

9. A table data segmentation device based on a database, characterized in that: The device comprises a memory for storing computer program instructions and a processor for executing the program instructions, wherein when the computer program instructions are executed by the processor, the device is triggered to execute the method according to any one of claims 1 to 8.

10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions can execute the method according to any one of claims 1 to 8.