A data processing method, device, equipment and storage medium

Through the fully connected neural network model, the Chinese definition of database fields is converted into a unified English definition, which solves the problem of inconsistent data attribute definitions caused by different developers' understanding of database development specifications, and realizes the unified and standardized data table design.

CN113568914BActive Publication Date: 2025-06-24SHANGHAI PUDONG DEVELOPMENT BANK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110863400.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-29
Publication Date
2025-06-24
Estimated Expiration
2041-07-29

AI Technical Summary

Technical Problem

In database application systems, different developers have different understandings of database development specifications, resulting in inconsistent definitions of the same data attribute in different data tables, affecting users' understanding of data and long-term maintenance of the database.

Method used

By obtaining the data field requirements table, the fully connected neural network model is used to convert the Chinese definition and attribute description information of the field into a unified English definition, and a field definition with unified standards is generated, thereby realizing the unified and standardized data table design.

Benefits of technology

It realizes the unified and standardized data table design, and improves the long-term maintenance of the database and the user's understanding of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113568914B_ABST
    Figure CN113568914B_ABST
Patent Text Reader

Abstract

The present invention discloses a data processing method, apparatus, device and storage medium. The method includes: obtaining a data field requirement table, wherein the data field requirement table includes: Chinese definitions of fields and attribute description information of fields; inputting the data field requirement table into a fully-connected neural network model to obtain English definitions of the fields, and the fully-connected neural network model is obtained by iteratively training a neural network model through a target sample set, and the target sample set includes: Chinese definitions of field samples, English definitions of field samples and attribute description information of field samples. Through the technical solution of the present invention, it is possible to combine database development specifications, make a field mapping table into a training sample set, and use a fully-connected neural network to generate field definitions with unified specifications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of computer technologies, and in particular, to a data processing method, apparatus, device, and storage medium. Background Art

[0002] In the initial stage of the database application system put into use, a large amount of initial data usually needs to be input at one time. In this process, relevant data often needs to be obtained from local or other systems. The traditional method is to use the data input function of the database application system itself to complete this work, or write specific code for specific services to achieve batch data input. The development of the input function includes the definition of the database, the definition of the data table, and the definition of the index. The database definition is generally completed by using "sql" scripts, and the specific service code is generally completed by using big data processing frameworks such as SpringBatch and Spark. The development efficiency of both depends mainly on the technical proficiency of developers and their understanding of business knowledge.

[0003] The main work of the batch project is to import different data files into the database according to certain processing rules. The data files are described in matrix form, where each row of data represents an independent piece of data, and each column of data represents the same attribute. To ensure that after the data is imported into the corresponding data table (the database contains multiple data tables), each column of data has a clear and accurate definition (abbreviation, precision size, whether it can be a null value, etc.), developers need to conduct a detailed design of the data table in combination with the database development specifications and the actual meaning of the data attributes. However, due to the different understandings of the specifications by different developers and the differences in personal experience, there will be a phenomenon that the definitions of the same data attribute are not unified in different data tables. This problem is not conducive to users' understanding of the data and the long-term maintenance of the database. Summary of the Invention

[0004] The embodiments of the present invention provide a data processing method, apparatus, device, and storage medium, so as to be able to combine database development specifications, make a field mapping table into a training sample set, and use a fully connected neural network for model training, so that the network takes data attributes (fields) as inputs and generates field definitions with unified specifications to achieve the standardization and normalization of data table design.

[0005] In a first aspect, the embodiments of the present invention provide a data processing method, including:

[0006] Obtain a data field requirement table, where the data field requirement table includes: the Chinese definition of the field and the attribute description information of the field;

[0007] Input the data field requirement table into a fully connected neural network model to obtain the English definition of the field. The fully connected neural network model is obtained by iteratively training the neural network model with a target sample set, and the target sample set includes: the Chinese definition of the field sample, the English definition of the field sample, and the attribute description information of the field sample.

[0008] Further, iteratively training the neural network model with the target sample set includes:

[0009] Input the Chinese definition of the field sample and the attribute description information of the field sample in the target sample set into the neural network model to obtain a predicted English definition;

[0010] Train the parameters of the neural network model according to the objective function formed by the predicted English definition and the English definition of the field sample;

[0011] Return to execute the operation of inputting the Chinese definition of the field sample and the attribute description information of the field sample in the target sample set into the neural network model to obtain a predicted English definition until a fully connected neural network model is obtained.

[0012] Further, after inputting the data field requirement table into the fully connected neural network model to obtain the English definition of the field, it further includes:

[0013] Determine the precision of the field and the constraint limits of the field according to the English definition of the field;

[0014] Create a database according to the Chinese definition of the field, the English definition of the field, the precision of the field, and the constraint limits of the field.

[0015] Further, after creating a database according to the Chinese definition of the field, the English definition of the field, the precision of the field, and the constraint limits of the field, it further includes:

[0016] Generate an index design table according to the English definition of the field;

[0017] Determine the target item set according to the support degree of each field combination in the index design table;

[0018] Determine the field combination with the highest confidence in the target item set as the target combined index.

[0019] Further, after determining the field combination with the highest confidence in the target item set as the target combined index, it further includes:

[0020] Generate a data table according to the English definition of the field and the target combined index;

[0021] Import the data corresponding to the English definition of the field into the database through the Spark framework.

[0022] Further, iteratively training the neural network model through the target sample set includes:

[0023] Obtain a target sample set, where the target sample set includes: the Chinese definition of the field sample, the English definition of the field sample, and the attribute description information of the field sample;

[0024] Perform semantic segmentation on the Chinese definition of the field sample and the attribute description information of the field sample to obtain an English word combination corresponding to the Chinese definition of the field sample and an English word combination corresponding to the attribute description information of the field sample;

[0025] Determine a field description matrix according to the English word combination corresponding to the Chinese definition of the field sample and the English word combination corresponding to the attribute description information of the field sample;

[0026] Obtain the first eigenvalue of the field description matrix;

[0027] Input the first eigenvalue into the neural network model to obtain a first vector;

[0028] Obtain the predicted English definition corresponding to the first vector;

[0029] If the predicted English definition is the same as the English definition of the field sample, it is determined to be correct;

[0030] When the correct rate is greater than the set threshold, obtain a fully connected neural network model.

[0031] In a second aspect, an embodiment of the present invention further provides a data processing device, and the device includes:

[0032] An acquisition module, configured to acquire a data field requirement table, where the data field requirement table includes: the Chinese definition of the field and the attribute description information of the field;

[0033] A determination module, configured to input the data field requirement table into a fully connected neural network model to obtain the English definition of the field, and the fully connected neural network model is obtained by iteratively training the neural network model through a target sample set, and the target sample set includes: the Chinese definition of the field sample, the English definition of the field sample, and the attribute description information of the field sample.

[0034] Further, the determination module is specifically configured to:

[0035] Input the Chinese definition of the field sample and the attribute description information of the field sample in the target sample set into the neural network model to obtain a predicted English definition;

[0036] Training the parameters of the neural network model with an objective function formed according to the predicted English definition and the English definition of the field sample;

[0037] Return the operation of inputting the Chinese definition and the attribute description information of the field sample in the target sample set into the neural network model to obtain the predicted English definition until a fully connected neural network model is obtained.

[0038] In a third aspect, an embodiment of the present invention further provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements any one of the data processing methods in the embodiments of the present invention.

[0039] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements any one of the data processing methods in the embodiments of the present invention.

[0040] In an embodiment of the present invention, by obtaining a data field requirement table, where the data field requirement table includes: the Chinese definition of the field and the attribute description information of the field; inputting the data field requirement table into a fully connected neural network model to obtain the English definition of the field, and the fully connected neural network model is obtained by iteratively training the neural network model with a target sample set, and the target sample set includes: the Chinese definition of the field sample, the English definition of the field sample, and the attribute description information of the field sample, so as to be able to combine the database development specification, make the field mapping table into a training sample set, and use a fully connected neural network for model training, so that the network takes data attributes (fields) as input and generates field definitions with unified specifications, so as to realize the standardization and normalization of data table design. Description of the Drawings

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0042] Figure 1 is a flowchart of a data processing method in an embodiment of the present invention;

[0043] Figure 1a is a flowchart of training a fully connected neural network model in an embodiment of the present invention;

[0044] Figure 1b is a flowchart of calculating frequent K-item sets in an embodiment of the present invention;

[0045] Figure 1c It is the structural diagram of the batch task program in the embodiment of the present invention;

[0046] Figure 1d It is the structural block diagram of the tool in the embodiment of the present invention;

[0047] Figure 2 It is the schematic structural diagram of a data processing device in the embodiment of the present invention;

[0048] Figure 3 It is the schematic structural diagram of an electronic device in the embodiment of the present invention;

[0049] Figure 4 It is the schematic structural diagram of a computer-readable storage medium containing a computer program in the embodiment of the present invention. Detailed implementation manners

[0050] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of description, only parts related to the present invention rather than all structures are shown in the drawings. Moreover, the embodiments in the present invention and the features in the embodiments can be combined with each other without conflict.

[0051] Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of the operations can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operations are completed, but it can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc. Moreover, the embodiments in the present invention and the features in the embodiments can be combined with each other without conflict.

[0052] The term "including" and its variants used in the present invention are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment".

[0053] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0054] Figure 1The flowchart of a data processing method provided by an embodiment of the present invention. This embodiment is applicable to the situation of data processing. This method can be executed by a data processing device in the embodiment of the present invention, and the device can be implemented in software and / or hardware, such as Figure 1 shown, the method specifically includes the following steps:

[0055] S110, obtain a data field requirement table, where the data field requirement table includes: the Chinese definition of the field and the attribute description information of the field.

[0056] S120, input the data field requirement table into a fully connected neural network model to obtain the English definition of the field. The fully connected neural network model is obtained by iteratively training the neural network model with a target sample set, and the target sample set includes: the Chinese definition of the field sample, the English definition of the field sample, and the attribute description information of the field sample.

[0057] Among them, the target sample set can be a field attribute mapping table drawn by organizing a database design file, as shown in Table 1:

[0058] Table 1

[0059]

[0060]

[0061] It can be seen from Table 1 that fields with similar meanings have the same definition in different data tables, and fields with the same Chinese definition have different attribute descriptions and English definitions in different data tables.

[0062] Such as Figure 1a shown, the way to input the data field requirement table into the fully connected neural network model to obtain the English definition of the field can be: obtain a field attribute mapping table, perform word sense segmentation on the Chinese definition of the field sample and the attribute description information of the field sample in the field attribute mapping table, and perform field vectorization on the segmented result. Determine a field description matrix according to the vectorized result, obtain the eigenvalues of the field description matrix, and train the parameters of the neural network model according to the eigenvalues of the field description matrix.

[0063] Optionally, iteratively training the neural network model with the target sample set includes:

[0064] Input the Chinese definition of the field sample and the attribute description information of the field sample in the target sample set into the neural network model to obtain a predicted English definition;

[0065] Train the parameters of the neural network model according to the objective function formed by the predicted English definition and the English definition of the field sample;

[0066] Return the operation of inputting the Chinese definition of the field sample and the attribute description information of the field sample in the target sample set into the neural network model to obtain the predicted English definition until the fully connected neural network model is obtained.

[0067] Optionally, after inputting the data field requirement table into the fully connected neural network model to obtain the English definition of the field, it further includes:

[0068] Determine the precision of the field and the constraint limit of the field according to the English definition of the field;

[0069] Create a database according to the Chinese definition of the field, the English definition of the field, the precision of the field, and the constraint limit of the field.

[0070] Optionally, after creating a database according to the Chinese definition of the field, the English definition of the field, the precision of the field, and the constraint limit of the field, it further includes:

[0071] Generate an index design table according to the English definition of the field;

[0072] Determine the target item set according to the support of each field combination in the index design table;

[0073] Determine the field combination with the maximum confidence in the target item set as the target combined index.

[0074] Among them, the method of generating an index design table according to the English definition of the field can be: query the field combination related to the English definition of the field from the database according to the English definition of the field, and generate an index design table according to the field combination related to the English definition of the field.

[0075] Among them, the target item set can be a frequent K-item set.

[0076] It should be noted that the field combination with a confidence greater than the confidence threshold in the target item set can also be determined as the target combined index. The embodiments of the present invention do not limit this.

[0077] Specifically, use the generated field definition as the input of the index definition module. According to the index association relationship already extracted by the Apriori calculation unit, select appropriate fields from the input fields as the combined index of the corresponding data table, and use this as the output of the module.

[0078] In a specific example, use Apriori (association analysis algorithm) to mine the index relationship in the index design table. The specific calculation process is as follows. As shown in Table 2, Table 2 is the index design table, and assume that A to E are five index fields.

[0079] Table 2

[0080] Table serial number Index field 1 A, C, D 2 B, C, E 3 A, B, C, E 4 B, E

[0081] Calculate the "support" of each field combination in Table 2, that is, the probability that one or more fields are used as indexes in the same data table. The calculation formula is as follows:

[0082] support(x, y) = num(xy) / num(All);

[0083] Among them, num(xy) is the number of occurrences of the xy combined index, and num(All) is the number of occurrences of all combined indexes in Table 2.

[0084] That is, the joint probability in mathematics. According to the above table, we can get:

[0085] support(A) = 2 / 4 = 0.5;

[0086] support(A, B) = 1 / 4 = 0.25.

[0087] Calculate the confidence of each field combination in this way.

[0088] Calculate the "frequent k-item set", that is, the item set that frequently appears in the dataset. It can be one or more. Since it contains k fields, it is called a k-item set. And the event that meets the minimum support threshold is called a frequent k-item set. As Figure 1b shown, when calculating the support, because the denominator part in the calculation formula is the same, the numerator part (that is, the number of occurrences) is used as the support. Among them, the minimum support is set to 2. When the support of a certain field combination is less than the minimum support, this field combination will be discarded. In this way, the most suitable field combination for use as an index can be selected from multiple field combinations.

[0089] (1) Obtain its confidence by calculating the support of the frequent k-item set. The calculation formula of the confidence is as follows:

[0090]

[0091] Among them, p(xy) is the probability of the xy combined index appearing, and p(y) is the probability of the y index appearing.

[0092] That is, the conditional probability in mathematics. According to the results of the frequent k-item set, 6 groups of confidences can be obtained:

[0093]

[0094]

[0095]

[0096]

[0097]

[0098]

[0099] (2) Set a confidence threshold (set to 0.8 in the example), and use the combinations with confidence greater than the confidence threshold among the 6 groups of confidences as the index setting rules. That is, when a data table has three fields B, C, and E at the same time, the system will use the three fields B, C, and E as a composite index, and the order of this composite index is (B, C, E) or (C, E, B). In the actual application process, since the number of data tables and fields is relatively large, the accuracy of the confidence will be relatively high, and there will be no situation where the same field combination has the same confidence.

[0100] Train a mature fully connected neural network with the data field requirements table (including the Chinese definitions and attribute descriptions of the fields) as the input, output a 2×5 matrix, perform inverse quantization processing on the matrix to obtain the corresponding English definitions. And perform matching in the target sample set based on this to obtain the accuracy and constraint limits of the field.

[0101] Specifically, index design is an important part of database development, and its design results will directly affect the query efficiency of data. Since the performance of a composite index is higher than that of the same number of single indexes, the design of the composite index is more critical. To improve the usage efficiency of the index, combined with the database development specifications and the index design principles, perform data mining on the index fields, analyze the potential correlation relationships therein, and output index results that meet the design specifications based on this.

[0102] Optionally, after determining the field combination with the maximum confidence in the target item set as the target composite index, it further includes:

[0103] Generate a data table according to the English definition of the field and the target composite index;

[0104] Import the data corresponding to the English definition of the field into the database through the Spark framework.

[0105] Specifically, the development of batch tasks involves the development of data import programs and database generation programs. Both have the characteristics of cumbersome development processes, strong code structuring, and high code content repetition rates during the implementation process. In the actual development process, developers will perform a large amount of repetitive development due to individual differences in development requirements. To address this issue, the embodiments of the present invention uniformly encapsulate the structural content in the development program, perform interface processing on the customized content, develop an automated design module, and thus replace the manual development process to improve the development efficiency.

[0106] Specifically, taking the English definition of the field as input, a corresponding batch task program is generated, such that each data file corresponds to a specified batch task. After secondary encapsulation, the structure of the generated batch task program is as follows Figure 1c shown in the figure: (1) Prepare, this part mainly prepares resources for the batch task. It mainly includes: database connection, FTP connection (used to pull specified data from the remote data warehouse to the local), and S3 connection (used to obtain object storage resources). To ensure the integrity of the batch task, when resource acquisition fails, the subsequent code will not be executed; (2) Execute, this part takes the field definition result as input, uses the Spark framework as the underlying data processing unit, and according to the field definition, imports the data in the data file into the database with specified precision; (3) Finish, this part mainly backs up the local files that have been imported into the database to object storage, and after the backup is completed, closes all connections and returns the currently occupied resources.

[0107] Specifically, taking the English definition and index definition of the field as input, using the field definition as the parameter for the database table creation statement and the index definition as the parameter for the index creation statement, a corresponding ".sql" script is generated, and thus the corresponding data table is generated.

[0108] To improve the definition precision of the field and index, the system provides a feedback module to manually verify the generated field definition and index definition. If the generated definition does not conform to the development specifications of the Shanghai Pudong Development Bank or conflicts with the data field requirements table, it is redefined in a manual design manner, and the definition result is recorded in the field attribute mapping table and the index design table, enabling the field definition module and the index definition module to perform secondary training (extraction), thereby improving the system precision.

[0109] Optionally, iteratively training the neural network model through the target sample set includes:

[0110] Obtaining the target sample set, where the target sample set includes: the Chinese definition of the field sample, the English definition of the field sample, and the attribute description information of the field sample;

[0111] Performing semantic segmentation on the Chinese definition of the field sample and the attribute description information of the field sample to obtain the English word combination corresponding to the Chinese definition of the field sample and the English word combination corresponding to the attribute description information of the field sample;

[0112] Determining the field description matrix according to the English word combination corresponding to the Chinese definition of the field sample and the English word combination corresponding to the attribute description information of the field sample;

[0113] Obtaining the first eigenvalue of the field description matrix;

[0114] Input the first eigenvalue into a neural network model to obtain a first vector;

[0115] Obtain the predicted English definition corresponding to the first vector;

[0116] If the predicted English definition is the same as the English definition of the field sample, it is determined to be correct;

[0117] When the accuracy rate is greater than a set threshold, obtain a fully connected neural network model.

[0118] In a specific example, since Chinese logic is complex and not easy to be processed by a computer, the Chinese definition of the field and its attribute description information are first subjected to semantic segmentation to make them into combinations of several English words. As shown in Table 3, Table 3 is the result of semantic segmentation for Table 1.

[0119] Table 3

[0120]

[0121]

[0122] Since English is relatively simpler than Chinese, after different field names are segmented, the same word combinations may be formed.

[0123] Field vectorization: Each word can be transformed into a vector through the tool "Word2Vec". For example, after the word "creation" is vectorized, it can be represented by [0.1, 0.2, 0.2, 0.8, 0.6], and for words with similar meanings, the distance between the vectors will be relatively small. The vector of the word "create" can be [0.1, 0.3, 0.2, 0.8, 0.6], and its Euclidean distance from "creation" can be expressed as:

[0124]

[0125] When the Euclidean distance is less than the distance threshold, the two are recognized as the same word. This kind of vectorization can calculate the similarity of different words and can also effectively avoid the interference of "multiple forms of the same word" on network training.

[0126] Field Matrixization: Generally, the Chinese definition of a field usually consists of 1 to 2 words, and the attribute description information consists of 4 to 5 words. Therefore, a 2×2 matrix is used to represent the Chinese definition of the field, and a 5×5 matrix is used to represent the attribute description information of the field. When the number of words is insufficient, a 0 vector is used for supplementation. When the number of words exceeds the upper limit, only the first part of the words is intercepted. Therefore, the matrix expressions of the field definition "creation time" and its attribute description information "Company effectiveregistration time" are as follows:

[0127]

[0128]

[0129] According to the uniqueness of word vectors, during the process of field matrixization, conjunctions and prepositions are excluded, etc., to increase the effective information volume of the matrix.

[0130] Obtaining Matrix Eigenvalues: First, the Chinese definition (2×5) matrix of the field and the attribute description information (5×5) matrix are spliced and fused to form a (7×5) field description matrix.

[0131]

[0132] After that, convolution processing (Convolution) is performed on the field description matrix. Since each word is described by a 1×5 vector, the convolution kernel specifications are 2×5, 3×5, and 4×5 respectively, and there are two of each type of convolution kernel. The description matrix is convolved using different convolution kernels and different sliding distances, so that each description matrix can generate multiple 1D vectors. These 1D vectors are subjected to pooling processing (Max-Pooling, that is, the maximum value in each vector is selected as the feature of the vector), and the maximum values of these vectors are fused into a 6D vector as the eigenvalue of the corresponding field description matrix.

[0133] Taking the above eigenvalue as the input of the neural network model, a 2×5-dimensional vector can be obtained. The Euclidean distance is calculated between the output vector and the field vector in the target sample set, and the minimum value is taken as the final mapping result. If the field of the mapping result is the same as the input field, it is determined to be correct, otherwise it is incorrect. When the correct rate reaches the set threshold, it is considered that the network has been trained maturely, that is, a fully connected neural network model is obtained.

[0134] In another specific example, such as Figure 1dAs shown in the figure, the design structure framework consists of 3 modules and 6 components, and specifically includes: a definition module, a feedback module, and a design module; the definition module includes: a field definition component and an index definition component, where the field definition component includes: a field attribute mapping table and a fully connected neural network; the index definition component includes: an Apriori calculation unit and an index design table, the feedback module includes: at least two feedback components, and the design module includes: a batch task design component and a database design component.

[0135] In terms of functions in the embodiments of the present invention, the system uses static files as input, operates in a multi-module low-coupling manner, abandons the disadvantages of manual intervention in the development process, and realizes the function of automatically generating batch tasks. In terms of design, the system introduces a negative feedback design pattern, enabling the system to continuously and autonomously iterate and update during use. In terms of performance, the definition module uses a variety of machine learning algorithms, overcoming the defects of high manual learning costs and different learning effects. In terms of efficiency, the system structurally encapsulates the database scripts and batch execution programs involved in batch tasks, reducing the code repetition in batch projects and increasing the difference between tasks.

[0136] The technical solution of this embodiment is to obtain a data field requirements table, where the data field requirements table includes: the Chinese definition of the field and the attribute description information of the field; input the data field requirements table into the fully connected neural network model to obtain the English definition of the field. The fully connected neural network model is obtained by iteratively training the neural network model with a target sample set, and the target sample set includes: the Chinese definition of the field sample, the English definition of the field sample, and the attribute description information of the field sample, so as to be able to combine the database development specifications, make the field mapping table into a training sample set, and use the fully connected neural network for model training, so that the network takes the data attribute (field) as input and generates a field definition with a unified specification, so as to achieve the standardization and normalization of data table design.

[0137] Figure 2 It is a schematic structural diagram of a data processing device provided by an embodiment of the present invention. This embodiment is applicable to data processing scenarios. The device can be implemented in software and / or hardware, and can be integrated into any device that provides data processing functions, such as Figure 2 As shown in the figure, the data processing device specifically includes: an acquisition module 210 and a determination module 220.

[0138] Among them, the acquisition module 210 is used to obtain a data field requirements table, where the data field requirements table includes: the Chinese definition of the field and the attribute description information of the field;

[0139] A determination module 220 is configured to input the data field requirement table into a fully-connected neural network model to obtain the English definition of the field. The fully-connected neural network model is obtained by iteratively training a neural network model with a target sample set, where the target sample set includes: the Chinese definition of the field sample, the English definition of the field sample, and the attribute description information of the field sample.

[0140] Optionally, the determination module is specifically configured to:

[0141] Input the Chinese definition of the field sample and the attribute description information of the field sample in the target sample set into the neural network model to obtain a predicted English definition;

[0142] Train the parameters of the neural network model according to an objective function formed by the predicted English definition and the English definition of the field sample;

[0143] Return to execute the operation of inputting the Chinese definition of the field sample and the attribute description information of the field sample in the target sample set into the neural network model to obtain a predicted English definition until a fully-connected neural network model is obtained.

[0144] The above product can execute the method provided in any embodiment of the present invention and has corresponding functional modules and beneficial effects for executing the method.

[0145] The technical solution of this embodiment realizes the ability to combine database development specifications, make a field mapping table into a training sample set, and use a fully-connected neural network for model training, so that the network takes data attributes (fields) as inputs and generates field definitions with unified specifications, in order to achieve the standardization and normalization of data table design by obtaining a data field requirement table, where the data field requirement table includes: the Chinese definition of the field and the attribute description information of the field; and inputting the data field requirement table into a fully-connected neural network model to obtain the English definition of the field. The fully-connected neural network model is obtained by iteratively training a neural network model with a target sample set, where the target sample set includes: the Chinese definition of the field sample, the English definition of the field sample, and the attribute description information of the field sample.

[0146] Figure 3 It is a schematic structural diagram of an electronic device in an embodiment of the present invention. Figure 3 A block diagram of an exemplary electronic device 12 suitable for implementing the embodiments of the present invention is shown. Figure 3 The shown electronic device 12 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.

[0147] Such as Figure 3As shown, the electronic device 12 is presented in the form of a general-purpose computing device. The components of the electronic device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 that connects different system components (including the system memory 28 and the processing unit 16).

[0148] The bus 18 represents one or more of several types of bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. By way of example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0149] The electronic device 12 typically includes a variety of computer system-readable media. These media can be any available media that can be accessed by the electronic device 12, including volatile and non-volatile media, removable and non-removable media.

[0150] The system memory 28 may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The electronic device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 34 may be used for reading and writing on a non-removable, non-volatile magnetic medium ( Figure 3 not shown, commonly referred to as a "hard disk drive"). Although Figure 3Not shown in the figure, a disk drive for reading and writing a removable non-volatile disk (such as a "floppy disk") and an optical disk drive for reading and writing a removable non-volatile optical disk (Compact Disc-Read Only Memory (CD-ROM), Digital Video Disc-Read Only Memory (DVD-ROM), or other optical media) can be provided. In these cases, each drive can be connected to the bus 18 through one or more data medium interfaces. The system memory 28 may include at least one program product having a set (such as at least one) of program modules configured to perform the functions of the embodiments of the present invention.

[0151] A program / utility 40 having a set (at least one) of program modules 42 can be stored, for example, in the system memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. The implementation of a network environment may be included in each or some combination of these examples. The program modules 42 generally perform the functions and / or methods in the embodiments described in the present invention.

[0152] The electronic device 12 can also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a display 24, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 12, and / or communicate with any device that enables the electronic device 12 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 22. Additionally, in this embodiment of the electronic device 12, the display 24 does not exist as an independent entity but is embedded in the mirror. When the display surface of the display 24 is not displaying, the display surface of the display 24 visually merges with the mirror surface. Moreover, the electronic device 12 can also communicate with one or more networks (such as a Local Area Network (LAN), a Wide Area Network (WAN), and / or a public network, such as the Internet) through the network adapter 20. As shown in the figure, the network adapter 20 communicates with other modules of the electronic device 12 through the bus 18. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, Redundant Arrays of Independent Disks (RAID) systems, tape drives, and data backup storage systems, etc.

[0153] The processing unit 16 executes various functional applications and data processing by running the programs stored in the system memory 28, for example, implementing the data processing method provided by the embodiments of the present invention:

[0154] Obtain a data field requirement table, where the data field requirement table includes: the Chinese definition of the field and the attribute description information of the field;

[0155] Input the data field requirement table into a fully connected neural network model to obtain the English definition of the field. The fully connected neural network model is obtained by iteratively training the neural network model with a target sample set, and the target sample set includes: the Chinese definition of the field sample, the English definition of the field sample, and the attribute description information of the field sample.

[0156] Figure 4 It is a schematic structural diagram of a computer-readable storage medium including a computer program in the embodiments of the present invention. The embodiments of the present invention provide a computer-readable storage medium 61, on which a computer program 610 is stored. When the program is executed by one or more processors, it implements the data processing method provided by all the embodiments of the present application:

[0157] Obtain a data field requirement table, where the data field requirement table includes: the Chinese definition of the field and the attribute description information of the field;

[0158] Input the data field requirement table into a fully connected neural network model to obtain the English definition of the field. The fully connected neural network model is obtained by iteratively training the neural network model with a target sample set, and the target sample set includes: the Chinese definition of the field sample, the English definition of the field sample, and the attribute description information of the field sample.

[0159] One or more computer-readable media in any combination can be adopted. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.

[0160] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, carrying computer-readable program code. Such a propagated data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0161] The program code contained on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, and the like, or any suitable combination of the foregoing.

[0162] In some embodiments, the client and the server may communicate using any currently known or future-developed network protocol, such as HTTP (Hyper Text Transfer Protocol), and may be interconnected with digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0163] The foregoing computer-readable medium may be included in the above-described electronic device; or may exist separately without being assembled into the electronic device.

[0164] The computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages, such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0165] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0166] The units involved in the embodiments described in the present disclosure can be implemented in software or in hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself.

[0167] The functions described above herein can be performed, at least in part, by one or more hardware logic components. By way of example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0168] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0169] Note that the above is only a preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, re-adjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments only. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A data processing method, characterized in that, Including: Obtain a data field requirements table, where the data field requirements table includes: the Chinese definition of the field and the attribute description information of the field; Input the data field requirements table into a fully connected neural network model to obtain the English definition of the field. The fully connected neural network model is obtained by iteratively training the neural network model with a target sample set, and the target sample set includes: the Chinese definition of the field sample, the English definition of the field sample, and the attribute description information of the field sample; Among them, the step of inputting the data field requirements table into the fully connected neural network model to obtain the English definition of the field includes: Perform semantic segmentation on the Chinese definition of the field and the attribute description information of the field in the data field requirements table to obtain the English word combination corresponding to the Chinese definition of the field and the English word combination corresponding to the attribute description information of the field; Determine a target field description matrix according to the English word combination corresponding to the Chinese definition and the English word combination corresponding to the attribute description information of the field; Input the target field description matrix into the fully connected neural network model to obtain the English definition of the field.

2. The method according to claim 1, characterized in that, Iteratively training the neural network model with a target sample set includes: Input the Chinese definition of the field sample and the attribute description information of the field sample in the target sample set into the neural network model to obtain a predicted English definition; Train the parameters of the neural network model according to the objective function formed by the predicted English definition and the English definition of the field sample; Return to the operation of inputting the Chinese definition of the field sample and the attribute description information of the field sample in the target sample set into the neural network model to obtain a predicted English definition until a fully connected neural network model is obtained.

3. The method according to claim 1, wherein After inputting the data field requirements table into the fully connected neural network model to obtain the English definition of the field, it further includes: Determine the precision of the field and the constraint limit of the field according to the English definition of the field; Create a database according to the Chinese definition of the field, the English definition of the field, the precision of the field, and the constraint limit of the field.

4. The method according to claim 3, characterized in that, After creating a database according to the Chinese definition of the field, the English definition of the field, the precision of the field, and the constraint limit of the field, it further includes: Generate an index design table according to the English definition of the field; Determine a target item set according to the support degree of each field combination in the index design table; Determine the field combination with the highest confidence in the target item set as the target combined index.

5. The method according to claim 4, wherein After determining the field combination with the highest confidence in the target item set as the target combined index, it further includes: Generate a data table according to the English definition of the field and the target combined index; Import the data corresponding to the English definition of the field into the database through the Spark framework.

6. The method according to claim 1, characterized in that, Iteratively training the neural network model with a target sample set includes: Obtain a target sample set, where the target sample set includes: the Chinese definition of the field sample, the English definition of the field sample, and the attribute description information of the field sample; Perform semantic segmentation on the Chinese definition of the field sample and the attribute description information of the field sample to obtain the combination of English words corresponding to the Chinese definition of the field sample and the combination of English words corresponding to the attribute description information of the field sample; Determine a field description matrix based on the combination of English words corresponding to the Chinese definition of the field sample and the combination of English words corresponding to the attribute description information of the field sample; Obtain the first eigenvalue of the field description matrix; Input the first eigenvalue into a neural network model to obtain a first vector; Obtain the predicted English definition corresponding to the first vector; If the predicted English definition is the same as the English definition of the field sample, it is determined to be correct; When the correct rate is greater than a set threshold, obtain a fully connected neural network model.

7. A data processing device, characterized in that, Including: An acquisition module for acquiring a data field requirement table, where the data field requirement table includes: the Chinese definition of the field and the attribute description information of the field; A determination module for inputting the data field requirement table into a fully connected neural network model to obtain the English definition of the field. The fully connected neural network model is obtained by iteratively training the neural network model with a target sample set, and the target sample set includes: the Chinese definition of the field sample, the English definition of the field sample, and the attribute description information of the field sample; Among them, the determination module is used to perform semantic segmentation on the Chinese definition of the field and the attribute description information of the field in the data field requirement table to obtain the combination of English words corresponding to the Chinese definition of the field and the combination of English words corresponding to the attribute description information of the field; determine a target field description matrix according to the combination of English words corresponding to the Chinese definition and the combination of English words corresponding to the attribute description information of the field; input the target field description matrix into the fully connected neural network model to obtain the English definition of the field.

8. The device according to claim 7, characterized in that, The determination module is specifically used for: Input the Chinese definition of the field sample and the attribute description information of the field sample in the target sample set into the neural network model to obtain a predicted English definition; Train the parameters of the neural network model according to the objective function formed by the predicted English definition and the English definition of the field sample; Return to execute the operation of inputting the Chinese definition of the field sample and the attribute description information of the field sample in the target sample set into the neural network model to obtain a predicted English definition until a fully connected neural network model is obtained.

9. An electronic device, characterized in that, Including: One or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the processors implement the method according to any one of claims 1-6.

10. A computer-readable storage medium containing a computer program, on which a computer program is stored, characterized in that, When the program is executed by one or more processors, it implements the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Field name processing method and device, computer storage medium and terminal

    CN109800332A