Data processing method and device, electronic equipment, readable storage medium and program product

By matching the metadata of the data table with the standard phrase library and root library to generate standard phrases, the problem of low management and utilization efficiency in data standardization is solved, efficient data management and sharing are achieved, and regulatory requirements are met.

CN120687585APending Publication Date: 2025-09-23MASHANG CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510231788.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

In existing technologies, there is a lack of effective means to manage and utilize data during data standardization, resulting in low data quality, difficulty in sharing, high management costs, and difficulty in meeting regulatory and compliance requirements.

Method used

By matching the metadata of the data table with the standard phrase library, a first word set is generated, and then matched with the standard root library to establish a first standard phrase, the standard phrase library is improved, and data management and utilization efficiency is improved.

Benefits of technology

By matching and processing metadata and establishing standard phrases, we can effectively improve data management and utilization efficiency, meet data sharing and regulatory requirements, and reduce management costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687585A_ABST
    Figure CN120687585A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device, electronic equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: matching metadata of a data table with a standard phrase library to obtain a first matching degree; under the condition that the first matching degree is smaller than a first threshold value, processing the metadata to obtain a first word set; matching the first word set with a standard root library to obtain a second matching degree; and under the condition that the second matching degree is greater than or equal to a second threshold value, creating a first standard word group based on the first word set. Through the method, the data standard library can be assisted to be perfected, so that the working efficiency is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a data processing method, device, electronic device, computer-readable storage medium, and computer program product. Background Art

[0002] The necessity of data standards is reflected in many aspects, such as improving data quality, promoting data sharing, reducing data management costs, improving data maintainability and scalability, and complying with regulations and compliance requirements. Creating and improving data standard information is a very important link in the data standardization process. Summary of the Invention

[0003] The embodiments of the present application provide a data processing method, device, electronic device, computer-readable storage medium, and computer program product, which can assist in improving the data standard library and thereby effectively improve work efficiency.

[0004] The technical solution of the embodiment of the present application is implemented as follows:

[0005] The present invention provides a data processing method, including:

[0006] Matching the metadata of the data table with the standard phrase library to obtain a first matching degree;

[0007] When the first matching degree is less than a first threshold, processing the metadata to obtain a first word set;

[0008] Matching the first word set with a standard root library to obtain a second matching degree;

[0009] In a case where the second matching degree is greater than or equal to a second threshold, a first standard phrase is created based on the first word set.

[0010] An embodiment of the present application provides a data processing device, including:

[0011] A matching module, configured to match metadata of the data table with a standard phrase library to obtain a first matching degree;

[0012] a processing module, configured to process the metadata to obtain a first word set when the first matching degree is less than a first threshold;

[0013] The matching module is further configured to match the first word set with a standard root library to obtain a second matching degree;

[0014] A creating module is configured to create a first standard phrase based on the first word set when the second matching degree is greater than or equal to a second threshold.

[0015] An embodiment of the present application provides an electronic device, including:

[0016] a memory for storing computer-executable instructions or computer programs;

[0017] The processor is used to implement the data processing method provided in the embodiment of the present application when executing the computer-executable instructions or computer programs stored in the memory.

[0018] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions, which is used to implement the data processing method provided in the embodiment of the present application when executed by a processor.

[0019] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, the data processing method provided in the embodiment of the present application is implemented.

[0020] The embodiments of the present application have the following beneficial effects:

[0021] First, the metadata of the data table is matched with the standard phrase library to confirm whether the metadata of the data table matches the standard phrases in the standard phrase library. When the first matching degree is less than the first threshold, it means that there is no standard phrase in the standard phrase library that matches the metadata of the data table. Then, the metadata is processed to obtain a first word set, and the first word set is matched with the standard root library to confirm the matching degree between the first word set and the standard root library. When the second matching degree is greater than or equal to the second threshold, it means that the similarity between the metadata and the standard phrase is high. Then, a first standard phrase (i.e., a new standard phrase to be reviewed) can be established based on the first word set, thereby establishing and improving the standard phrase library, so as to more effectively manage and utilize data based on the standard phrase library, thereby effectively improving work efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 1 is a schematic diagram of the architecture of a data processing system 100 provided in an embodiment of the present application;

[0023] Figure 2 is a structural diagram of an electronic device 500 provided in an embodiment of the present application;

[0024] Figure 3 Schematic diagram of the data processing method provided in the embodiment of the present application;

[0025] Figure 4 Schematic diagram of the data processing method provided in the embodiment of the present application;

[0026] Figure 5This is a schematic diagram of the principle of the data processing method provided in the embodiment of the present application;

[0027] Figure 6 It is a process diagram of the data processing method provided in an embodiment of the present application. DETAILED DESCRIPTION

[0028] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0029] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0030] It is understandable that in the embodiments of the present application, when user information and other related data are involved, when the embodiments of the present application are applied to specific products or technologies, user permission or consent must be obtained, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards.

[0031] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0032] In the following description, the terms "first\second\..." are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understandable that "first\second\..." can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0034] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0035] 1) Metadata: Metadata refers to data about data, that is, a description of the data's structure, attributes, and content. In a database, metadata typically includes table names, field names, data types, constraints, and default values.

[0036] 2) Standard phrase library: A standard phrase library is a predefined and compiled vocabulary collection that contains a series of standardized keywords or phrases that are commonly used in specific fields or industries to represent common concepts, attributes, operations or other business terms.

[0037] 3) Standard Root Library: The Standard Root Library is a vocabulary collection designed specifically for text processing, natural language processing, and data standardization. It contains a series of predefined roots or stems. A root is the core part of a word, typically containing its basic meaning, while a stem is the remaining portion of a word after removing prefixes and suffixes.

[0038] 4) Natural Language Processing (NLP): NLP is a branch of Artificial Intelligence (AI) that involves the interaction between computers and human language. The purpose of NLP is to enable computers to understand, interpret, generate, and process human language in order to perform tasks such as text analysis, sentiment analysis, machine translation, and speech recognition. NLP escaping refers to the specific processing of text data to eliminate or neutralize characters or symbols that may interfere with the execution of NLP tasks. These characters or symbols may include punctuation marks, special symbols, control characters, etc., which may have special meanings or functions in the language, but may not be needed or even harmful in some NLP applications.

[0039] 5) Word Vector Model: A word vector model is a technique in natural language processing that maps words or phrases in a vocabulary into a real-valued vector space, with each word represented as a fixed-length vector. This representation not only captures the semantic information of words but also analyzes the relevance and similarity between them by calculating the distance between vectors. Common word vector models include Word2Vec and GloVe. Word vector models are an important tool in natural language processing, providing a foundation for understanding text data and building complex language models.

[0040] The necessity of data standards is reflected in improving data quality, promoting data sharing, reducing data management costs, improving data maintainability and scalability, and complying with regulations and compliance requirements. Through data standardization, enterprises can manage and utilize data more effectively, improve business efficiency and competitiveness, and creating and improving data standard information is a very important link in the data standardization process.

[0041] Based on this, embodiments of the present application provide a data processing method, device, electronic device, computer-readable storage medium, and computer program product that can assist in improving the data standard library, thereby effectively improving work efficiency. The electronic device provided in the embodiments of the present application can be implemented as a server, or can be implemented collaboratively by a server and a terminal. The following description uses the data processing method provided in the embodiments of the present application collaboratively implemented by a server and a terminal as an example.

[0042] For example, see Figure 1 , Figure 1 This is a schematic diagram of the architecture of the data processing system 100 provided in an embodiment of the present application, which is used to support a data processing application, such as Figure 1 As shown, the data processing system 100 includes: a server 200, a network 300, a terminal 400 and a database 600. The terminal 400 is connected to the server 200 via the network 300, and the server 200 is connected to the database 600. The network 300 can be a local area network or a wide area network, or a combination of the two; the database 600 is used to store structured data (such as text information, mapping relationships and other types of data).

[0043] In some embodiments, a user initiates a processing request through terminal 400 and transmits the processing request to server 200 via network 300. Server 200 obtains a corresponding data table from database 600 based on the received processing request. Next, server 200 matches the metadata of the data table with the standard phrase library in database 600 to obtain a first matching degree. Subsequently, if the first matching degree is less than a first threshold, server 200 processes the metadata to obtain a first word set. Thereafter, server 200 matches the first word set with the standard root library in database 600 to obtain a second matching degree. Finally, if the second matching degree is greater than or equal to a second threshold, server 200 creates a first standard phrase (i.e., a new standard phrase to be reviewed) based on the first word set and returns the first standard phrase to terminal 400 via network 300 for display on terminal 400 for manual review by the user. When the first standard phrase passes the review, database 600 marks the first standard phrase as a pre-created standard phrase and adds the pre-created standard phrase to the standard library.

[0044] In some embodiments, a user initiates a processing request through terminal 400 and transmits the processing request to server 200 through network 300. Server 200 obtains a corresponding data table from database 600 based on the received processing request. Next, server 200 matches the metadata of the data table with the standard phrase library in database 600 to obtain a first matching degree. When the first matching degree is greater than or equal to a first threshold, database 600 marks the metadata as a standard phrase.

[0045] In other embodiments, the embodiments of the present application can also be implemented with the help of cloud technology. Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and network within a wide area network or local area network to realize data calculation, storage, processing, and sharing.

[0046] Cloud technology is a general term for network, information, integration, management platform, and application technologies used in the cloud computing business model. It can form a resource pool that can be used flexibly and conveniently on demand. Cloud computing technology will become a key support. The backend services of technical network systems require a large amount of computing and storage resources.

[0047] For example, Figure 1 The server 200 in the example can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal 400 can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, car terminal, etc., but is not limited to these. The terminal 400 and the server 200 can be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present application.

[0048] It should be noted that the data processing method provided in the embodiments of the present application can be applied to different data processing application scenarios such as data cleaning, data analysis, data integration, data visualization, and database management, and can more effectively manage and utilize data.

[0049] The following continues to describe the structure of the electronic device provided in the embodiment of the present application. Take the electronic device as an example, see Figure 2 , Figure 2 is a structural diagram of an electronic device 500 provided in an embodiment of the present application, Figure 2The electronic device 500 shown includes: at least one processor 510, a memory 540, and at least one network interface 520. The various components in the electronic device 500 are coupled together via a bus system 530. It is understood that the bus system 530 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 530 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, the bus system 530 is not described in detail. Figure 2 Various buses are labeled as bus system 530 .

[0050] The processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0051] The memory 540 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 540 may optionally include one or more storage devices that are physically remote from the processor 510.

[0052] The memory 540 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 540 described in the embodiments of the present application is intended to include any suitable type of memory.

[0053] In some embodiments, the memory 540 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.

[0054] Operating system 541, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;

[0055] A network communication module 542 for reaching other computing devices via one or more (wired or wireless) network interfaces 520 , exemplary network interfaces 520 including Bluetooth, Wireless LAN (WiFi), and Universal Serial Bus (USB);

[0056] In some embodiments, the apparatus provided in the embodiments of the present application may be implemented in software. Figure 2 The data processing device 543 stored in the memory 540 is shown. It can be software in the form of a program or plug-in, and includes the following software modules: a matching module 5431, a processing module 5432, a creation module 5433, a determination module 5434, a marking module 5435, an update module 5436, and an acquisition module 5437. These modules are logical and can be arbitrarily combined or further divided according to the functions implemented. It should be noted that in Figure 2 For the sake of convenience, all the above modules are shown at once, but it should not be considered that the data processing device 543 excludes the implementation that only includes the matching module 5431, the processing module 5432 and the creation module 5433. The functions of each module will be explained below.

[0057] The data processing method provided in the embodiment of the present application will be described in detail below in combination with the exemplary application and implementation of the server provided in the embodiment of the present application.

[0058] See also Figure 3 , Figure 3 This is a flow chart of the data processing method provided in the embodiment of the present application, which will be combined with Figure 3 The steps shown are explained.

[0059] In step 101, metadata of a data table is matched with a standard phrase library to obtain a first matching degree.

[0060] In some embodiments, first, a data table is obtained, and metadata of the data table (for example, including table name, column name, column data type, etc.) is extracted, and a standard phrase library is obtained; then, the metadata of the data table is matched with each phrase in the standard phrase library using an appropriate matching strategy (for example, exact matching, fuzzy matching, partial matching, pattern matching, etc.); then, the matching degree between the metadata and the standard phrase library is calculated according to the matching strategy, where the matching degree can be a score indicating the closeness of the match, such as a value between 0 and 1, where 1 represents a complete match, or a classification description can be used, such as complete mismatch, low matching degree, medium matching degree, high matching degree, and complete match; finally, the calculated matching degree is used as the first matching degree.

[0061] For example, suppose a data table is obtained, and the corresponding metadata is: table name: ˋOrdersˋ, column names: ˋOrderIDˋ, ˋCustomerNameˋ, ˋOrderDateˋ, ˋTotalAmountˋ, and the standard phrase library is obtained at the same time, which contains the following predefined words: ˋOrderNumberˋ, ˋCustomerNameˋ, ˋDateOfOrderˋ, ˋAmountˋ; then, the metadata of the data table is matched with the standard phrase library, and the table name in the metadata is matched with the standard phrase library. There is no word directly corresponding to ˋOrdersˋ in the standard phrase library, so the first matching degree of the table name is 0, that is, no match; matching the column name in the metadata with the standard phrase library, ˋOrderIDˋ and ˋOrderNumberˋ have very similar meanings, so it can be considered that ˋOrderIDˋ is ˋOrderNumberˋ. For example, the first match degree for OrderID can be a high match; CustomerName is exactly the same as CustomerName, meaning that CustomerName's first match degree is a perfect match; OrderDate and DateOfOrder both represent the order date. Although expressed differently, they describe the same information, meaning that the first match degree for OrderDate can be a high match; TotalAmount and Amount represent the total amount of the order. TotalAmount can be considered a more detailed representation of Amount, meaning that the first match degree for TotalAmount can be a medium match. In practical applications, algorithms can also be used to determine the match degree between metadata and a standard phrase library. Here, the match degree is a hypothetical similarity score that indicates the degree of match between metadata and terms in the standard phrase library.

[0062] For example, assuming exact matching, the calculation of the matching degree between metadata and the standard phrase library can be divided into complete matching and mismatching. For example, the metadata includes a column named "TotalValue". During the matching process with the standard phrase library, if "TotalValue" exists in the standard phrase library, the metadata is considered to be completely matched with the standard phrase library, that is, the first matching degree is a complete match (that is, 100%); assuming that there is no completely consistent "TotalValue" in the standard phrase, for example, there are other phrases such as "Total" and "TotalAmount", the metadata is considered to be mismatched with the standard phrase library, that is, the first matching degree is a complete mismatch.

[0063] In step 102, when the first matching degree is less than a first threshold, the metadata is processed to obtain a first word set.

[0064] In some embodiments, see Figure 4 , Figure 4 is a flow chart of the data processing method provided in the embodiment of the present application, such as Figure 4 As shown, Figure 3 Step 102 shown can be performed by Figure 4 Steps 1021 to 1023 shown are implemented by combining Figure 4 The steps shown are explained.

[0065] In step 1021, the metadata is segmented to obtain a plurality of first words.

[0066] In some embodiments, before the metadata is segmented, the metadata can be preprocessed (for example, unifying the format of characters, unifying the case of characters, removing or replacing non-alphabetic characters, etc.); then, different algorithms can be used to segment the metadata (for example, regular expressions, dictionary segmentation, etc.).

[0067] For example, suppose there is metadata called "CustomerOrderDate". First, preprocess the metadata and convert it to lowercase "customerorderdate". Then, use the dictionary segmentation method to match the preprocessed metadata with words in the dictionary. Suppose that "customer" is found in the dictionary and the boundary is determined. Suppose that "order" is found in the dictionary and the boundary is determined. Suppose that "date" is found in the dictionary and the boundary is determined. Finally, the segmented result is obtained, that is, multiple words in the first language, "customer", "order", and "date".

[0068] In step 1022 , a plurality of second words are generated based on the plurality of first words.

[0069] Here, the language of the first word and the language of the second word can be different. For example, the language of the first word can be English and the language of the second word can be Chinese. Then, multiple second words can be generated based on multiple first words. Multiple English words can be translated in sequence to obtain corresponding Chinese words. No specific limitation is made here.

[0070] In some embodiments, first, a suitable translation tool or service is selected. Here, it can be an online translation application programming interface (API), such as the Google Translate API, or an offline translation software. Then, the first words (e.g., English words) are sent to the translation tool for translation, and each word is translated to obtain the second words (e.g., Chinese words). Subsequently, all the translated words are collected and stored in a list or set. After that, the translation results may contain duplicate words, and these words need to be de-duplicated. Finally, the de-duplicated words are combined into a new list, and this list is the final set of second-language words.

[0071] For example, assume that the first words are English words, and the first-word list includes 'Persons', 'Customer', 'Person', 'Date'. The corresponding translation results obtained by translating the word list are '人', '顾客', '人', and '日期'. After de-duplicating the translation results, the de-duplicated results are '人', '顾客', and '日期'. Finally, the de-duplicated words are recombined into a new list, and the set of second words (i.e., Chinese words) is {'人', '顾客', '日期'}.

[0072] In step 1023, a first-word set is generated based on multiple first words and multiple second words.

[0073] In some embodiments, an empty set is created to store the translated words; the words in the obtained set of second-language words and the words in the set of first-language words are added to the created empty set to obtain the first-word set.

[0074] For example, assume that the first-word list of the first language includes 'Order', 'Customer', 'Date', and the translated second-word list of the second language includes '秩序', '顾客', and '日期'; combining the first-word list of the first language and the second-word list of the second language, the obtained first-word set is {'Order', 'Customer', 'Date', '秩序', '顾客', '日期'}.

[0075] In some embodiments, when the first matching degree is greater than or equal to the first threshold, the metadata is marked as a standard phrase.

[0076] For example, assuming that the metadata is "This product is an efficient and durable solar panel", the standard phrase is "efficient and durable solar panel", assuming that cosine similarity is used for calculation and the first matching degree is 85%, and assuming that the first threshold is 80%, the metadata is marked as the standard phrase.

[0077] It should be noted that the first threshold may be a preset value (eg, 100%, ie, a complete match is required), or may be dynamically set according to an actual scenario, and is not specifically limited here.

[0078] In some embodiments, the first threshold may be 100%, that is, the metadata is marked as a standard phrase only when the metadata completely matches the standard phrase.

[0079] In step 103, the first word set is matched with the standard root library to obtain a second matching degree.

[0080] In some embodiments, first, a standard root library is created or obtained, which contains a series of standard roots. These roots are usually the basic form of a word. For example, for a verb, the root may be its original form. Then, ensure that the format of the root library is suitable for matching operations, for example, it can be a simple list or database. Then, obtain a first word set and pre-process the words in the first word set (for example, unify the format of characters, unify the uppercase and lowercase characters, remove or replace non-alphabetic characters, etc.). Then, select an appropriate algorithm to compare the words in the first word set with the roots in the standard root library, which may include simple string matching or more complex algorithms such as regular expression matching, word form restoration, etc. Then, use the selected matching algorithm to match each word in the first word set with the roots in the standard root library, and for each matching operation, calculate the second matching degree according to the defined calculation method. Finally, store the matching result and the corresponding second matching degree of each word for subsequent analysis.

[0081] For example, assuming that the first word set includes "Running", "Jumping", and "Swimming", and the standard root library includes "Run", "Jump", and "Swim", through matching, "Running" is matched with "Run", and the second matching degree is calculated to be 100%; "Jumping" is matched with "Jump", and the second matching degree is calculated to be 100%; "Swimming" is matched with "Swim", and the second matching degree is calculated to be 100%.

[0082] It should be noted that if a word matches multiple roots, the final second match may need to be determined based on the best match or the most relevant match; in addition, if the word does not completely match the root, it may be necessary to use fuzzy matching or lemmatization technology to find the closest root and calculate the corresponding match.

[0083] In step 104 , when the second matching degree is greater than or equal to the second threshold, a first standard phrase is created based on the first word set.

[0084] Here, the first standard phrase refers to a new standard phrase to be reviewed, which will be further reviewed to determine whether it belongs to the standard phrase.

[0085] In some embodiments, a second threshold is determined to determine whether a word is close enough to a root word to be considered a candidate for a new standard phrase; a second matching degree between the acquired first word set and the standard root word library is compared with the second threshold value, and if the second matching degree is greater than or equal to the second threshold value, the first word set is marked as a first standard phrase.

[0086] For example, assuming that the second threshold is 80%, and the second matching degree between the first word set and the standard root word library is obtained to be 90% through calculation, the first word set is marked as the first standard phrase.

[0087] In some embodiments, after creating a first standard phrase based on the first word set, the following operations may be performed: when the first standard phrase passes the review, the first standard phrase is marked as a pre-created standard phrase; the pre-created standard phrase is added to the standard library to obtain an updated standard library; and the threshold is updated based on the updated standard library.

[0088] For example, after the first word combination is marked as the first standard phrase, the first word set will be reviewed to confirm whether it meets the conditions of the new standard phrase. After the review is passed, the first word set will be added to the standard library and called the new standard phrase, thereby obtaining an updated standard library. Subsequently, the threshold value of the updated standard library will be appropriately adjusted.

[0089] It should be noted that the review process can be submitted to the reviewer for manual review, or it can be analyzed in one step through automated tools to achieve automatic review. No specific restrictions are made here.

[0090] In some embodiments, after executing step 104, the following operations may also be performed: when the second matching degree is less than a second threshold, the words included in the first word set are processed to obtain a second word set; the second word set is matched with a standard root word library to obtain a third matching degree; when the third matching degree is greater than or equal to a third threshold, a second standard phrase is created based on the second word set.

[0091] Here, the second standard phrase is the same as the first standard phrase mentioned above, that is, a new standard phrase to be reviewed, which will be further reviewed to determine whether it belongs to the standard phrase.

[0092] In some embodiments, a second threshold is first determined. When the second matching degree is less than the second threshold, the words in the first word set are processed to obtain a second word set, wherein the processing includes translating the words, segmenting them, and rearranging and combining the segmented words; then, the words in the second word set are matched against the standard root library, and a third matching degree is calculated; thereafter, a third threshold is determined. When the third matching degree is greater than or equal to the third threshold, the second word set can be marked as a second standard phrase.

[0093] For example, assuming that the first word set is ["Running","Jumping","Swimming","Climbing"], the standard root library is ["Run","Jump","Swim","Climb"], the corresponding second matching degree is less than the second threshold, NLP escape processing is performed on the words in the first word combination, and stem extraction is performed on ˋClimbingˋ to obtain ˋClimbˋ (hypothetical simplified processing), and stem extraction is performed on ˋRunningˋ to obtain ˋRunˋ, and stem extraction is performed on ˋJumpingˋ to obtain ˋJumpˋ, and stem extraction is performed on ˋSwimmingˋ to obtain ˋSwimˋ. The second word set is obtained: ["Climb","Run","Jump","Swim"], and the matching degree is calculated to obtain a third matching degree of 100%, which is greater than the third threshold of 80%. The second word set is marked as the second standard phrase.

[0094] In some embodiments, after creating a second standard phrase based on the second word set, the following operations may be performed: when the second standard phrase passes the review, the second standard phrase is marked as a pre-created standard phrase; the pre-created standard phrase is added to the standard library to obtain an updated standard library; and the threshold is updated based on the updated standard library.

[0095] In some embodiments, when the third matching degree is less than the third threshold and the third matching degree is greater than or equal to the fourth threshold, and when the first number of data in the data table that meets the verification rule is greater than or equal to the fourth threshold, the metadata is marked as a standard phrase corresponding to the verification rule.

[0096] In some embodiments, each word in the second word set is matched with the standard root library, and a third matching degree is calculated. When the third matching degree is less than a third threshold and the third threshold is greater than a fourth threshold, data in the data table is obtained, and the verification rules corresponding to the data are determined. The data volume of the data that meets different verification rules is counted respectively, and the ratio between the data volume of the data that meets each verification rule and the data volume of the total data is calculated as a fourth matching degree. When the fourth matching degree is greater than a fifth threshold (that is, the first number of data in the data table that meets the verification rule is greater than or equal to the fourth threshold), the metadata is marked as a standard phrase corresponding to the verification rule.

[0097] For example, assuming that the verification rules matching the data in the data table include verification rule 1, verification rule 2 and verification rule 3, there are a total of 20 data in the data standard, 18 data satisfying verification rule 1, 3 data satisfying verification rule 2, and 10 data satisfying verification rule 3. Assuming that the fifth threshold is 75%, the proportion of data satisfying verification rule 1 is 90%, the proportion of data satisfying verification rule 2 is 10%, and the proportion of data satisfying verification rule 3 is 50%. Therefore, the metadata is marked as the standard phrase corresponding to verification rule 1.

[0098] In some embodiments, the above-mentioned determination of the verification rules corresponding to the data table can also be implemented in the following way: determine the verification rules corresponding to the type of data in the data table to obtain multiple verification rules; determine the second quantity of data corresponding to each verification rule; based on the second quantity, determine the verification rules corresponding to the data table from multiple verification rules.

[0099] In some embodiments, the data type of the data is first identified, and the type of the data retrieved from the data table is identified to determine the type of each data (for example, text, number, date, etc.); then, based on the identified data type, the corresponding verification rules are filtered out from the rule library, and the rule library should contain verification rules applicable to various data types; then, for each filtered verification rule, the amount of data that meets the rule in the data table is counted; then, the amount of data corresponding to each verification rule is compared, and the verification rule with the largest data amount is selected as the verification rule corresponding to the data table.

[0100] For example, suppose there is a data table product_info, which contains the following fields: product_id (numeric type), product_name (text type), release_date (date type). The rule base contains the following validation rules: Rule 1: product_id must be a positive integer, Rule 2: product_name must be between 2 and 100 characters long, Rule 3: release_date must be a valid date format. By checking the data in the product_info table, the data type of each column is identified. Based on the identified data type, the rule base filters out the validation rules. The corresponding validation rules are: product_id corresponds to rule 1, product_name corresponds to rule 2, and release_date corresponds to rule 3. For each validation rule, the amount of data that meets the rule is counted. Assume that the amount of data that satisfies rule 1 (product_id is a positive integer) is 95, the amount of data that satisfies rule 2 (product_name is between 2 and 100 characters) is 90, and the amount of data that satisfies rule 3 (release_date is a valid date format) is 100. Finally, by comparing the statistical results, rule 3 corresponds to the largest amount of data (100 records), so rule 3 is selected as the validation rule for the product_info data table.

[0101] In some embodiments, when the third matching degree is less than a fourth threshold, the first word set, the second word set, and the data are converted into word vectors to obtain word vectors; the similarity between the word vectors and the word vectors corresponding to the standard library is determined, and the similarity is used as the fifth matching degree, wherein the standard library includes a standard phrase library and a standard root library; when the fifth matching degree is greater than a sixth threshold, the metadata is marked as a third standard phrase.

[0102] Here, the third standard phrase is the same as the first standard phrase mentioned above, that is, a new standard phrase to be reviewed, which will be further reviewed to determine whether it belongs to the standard phrase.

[0103] In some embodiments, first, a word vector model, such as Word2Vec, Glo Ve or BERT, is selected or trained to convert words into word vectors; then, the word vector model is used to convert each word in the first word set and the second word set and the data in the data table into corresponding word vectors; then, the similarity between each output word vector and the word vector in the standard library (including the standard phrase library and the standard root library) is determined, which can be achieved by calculating cosine similarity, Euclidean distance or other similarity metrics; then, a sixth similarity threshold is determined, and the calculated similarity is compared with the sixth threshold. If the similarity is greater than the sixth threshold, the corresponding metadata is marked as a third standard phrase.

[0104] For example, assuming that the first word set is ["Running", "Jumping", "Swimming"], the second word set is wield["Run", "Jump", "Swim"], and a piece of data in the data table is "The person is running and jumping." Then, the words in the first word set are converted into word vectors 1, 2, and 3, and the words in the second word set are converted into word vectors 4, 5, and 6; then, the data "The person is running and jumping." is segmented to obtain a word list ["The", "person", "is", "running", "and", "jumping"], and each word is converted into a word vector; then, the similarity between each output word vector (from the data and the first and second word sets) and the word vectors in the standard library is calculated. Subsequently, assuming that the sixth threshold is 0.8, if the ratio of the number of sufficiently similar word vectors to the word vectors in the standard library to the total number is greater than 0.8, the metadata is marked as the third standard phrase.

[0105] In some embodiments, before executing the above-mentioned use of the word vector model to convert each word in the first word set and the second word set and the data in the data table into corresponding word vectors, the word vector model needs to be trained. The specific training process can be implemented in the following way: obtain a large text data set, where the data set can be a book, an article, a web page, or any other text source; clean and format the text, for example, remove punctuation, convert to lowercase, delete stop words, etc., to obtain sample data; select a suitable word vector model algorithm, such as Word2Vec, GloVe, FastText or BERT; create a model instance and load the preprocessed sample data, traverse all sentences in the sample data, and for each sentence, use the model algorithm to update the word vector, repeat the above vector update process until the model converges or reaches a predetermined number of iterations. During the training process, use the change of the loss function to evaluate the model performance, and finally save the trained model for subsequent use.

[0106] In some embodiments, after marking the metadata as a new standard phrase to be reviewed, the following operations may be performed: when the third standard phrase passes the review, the third standard phrase is marked as a pre-created standard phrase; the pre-created standard phrase is added to the standard library to obtain an updated standard library; and the threshold is updated based on the updated standard library.

[0107] In some embodiments, a dictionary library is obtained, wherein the dictionary library contains multiple words and multiple phrases; a mapping relationship between each word or phrase in the dictionary library and a standard library is determined, wherein the standard library includes a standard phrase library and a standard root library; and based on the mapping relationship, metadata corresponding to the standard library is marked as a standard phrase.

[0108] In some embodiments, first, a dictionary library containing multiple words and phrases is obtained from a data source; then, an appropriate algorithm (such as exhaustive, machine learning, deep learning, etc.) is used to determine the mapping relationship between each word or phrase in the dictionary library and a standard library (including a standard phrase library and a standard root library); finally, based on the mapping relationship, it is checked whether the words or phrases in the metadata correspond to the entries in the standard library. If so, these words or phrases in the metadata are marked as standard phrases.

[0109] For example, suppose the dictionary includes the words: "run", "jump", "swim", and the phrases: "running", "jumping", the standard phrase library includes "efficient runner", "agile jumper", "strong swimmer", and the standard root library includes "run", "jump", "swim". First, obtain the dictionary by querying the database or reading the file. Then, for each word or phrase in the dictionary, use an algorithm (such as string matching, word vector similarity calculation, etc.) to determine the mapping relationship between them and the entries in the standard library. For example, "run" is mapped to ["run", "efficientrunner"], and "running" is mapped to ["run", "efficientrunner"]. Next, assuming there is metadata "He is an efficient runner." and "She likes to jump and swim.", check whether the words or phrases in the metadata correspond to the standard phrases or roots in the mapping relationship. "efficient "runner" corresponds to "running" in the dictionary library, so it is marked as a standard phrase. "jump" and "swim" correspond to "jump" and "swim" in the dictionary library, so they are marked as standard phrases. Finally, "efficient runner" in the metadata is marked as a standard phrase.

[0110] The following describes an exemplary application of the embodiment of the present application in a practical application scenario, which describes the specific implementation process of the data processing method in the scenario of improving the data standard library.

[0111] The necessity of data standards lies in improving data quality, facilitating data sharing, reducing data management costs, enhancing data maintainability and scalability, and meeting regulatory and compliance requirements. Through data standardization, enterprises can more effectively manage and utilize data, thereby enhancing business efficiency and competitiveness. Creating and improving data standard information is a crucial step in the data standardization process.

[0112] Based on this, this application analyzes the relational database metadata and standard library by providing table metadata parsing, field Chinese-English word segmentation, data standard verification, and data standard library construction, establishes and improves the data standard library, and establishes a corresponding relationship between data standards and metadata, which is conducive to the quality inspection and rectification of metadata standards, thereby effectively improving work efficiency.

[0113] In some embodiments, see Figure 5 , Figure 5This is a schematic diagram of the principle of the data processing method provided in the embodiment of the present application. Figure 5 Describe the specific implementation process.

[0114] like Figure 5 As shown, the data processing method provided by the present application consists of five modules: a calling interface, word segmentation and analysis, data standards, Chinese-English escape, and a server. Among them, the calling interface module is used to provide data entry and exit calls, the word segmentation and analysis module is used to use a word segmentation tool to segment Chinese and English words, the data standard module contains word roots (English words and abbreviations, word types include data types, data length, size, etc.), standard phrases (standard phrases composed of one or more words separated by underscores (_) contain data type definitions), the Chinese-English escape module is used to perform Chinese and English analysis based on the English encoding and Chinese description of the metadata field, and the server module is the transit hub of the data processing device of the present application, which is responsible for the combination and collaboration of various modules of the entire device, and escapes the disassembled Chinese and English to match them with the data standards or performs word segmentation and analysis to match them with the data standards.

[0115] In some embodiments, first, the interface module is called to receive the incoming data table metadata and the data information in the data table; then, the server module directly matches the English in the received metadata with the standard phrases in the data standard library. If the match is successful, the metadata is marked as meeting the data standard. Otherwise, the Chinese-English translation module is called to parse the English and Chinese words in the metadata into words and word groups, or into Chinese words and Chinese phrases. Here, the English and Chinese descriptions in the metadata are required, and the Chinese descriptions are not necessarily required; then, the server module matches the English words after Chinese translation with the words in the data standard. If all match, the metadata is translated into English. The Chinese fields obtained by translation are matched with the Chinese fields obtained by Chinese-English translation. If the matching rate reaches 70% or above, the phrase corresponding to the metadata can be marked as a standard phrase. If the matching rate is lower than 70% or the words are partially matched, the standard phrases of manual business data can be analyzed and confirmed to improve the data standard words and standard phrases. If Chinese-English translation and word segmentation are not possible, the table data can be matched according to some specific rules or natural language translation can be performed, and then data standard matching can be performed after translation. Finally, if none of the above methods can be used to process the problem, a prompt message will be returned indicating that the problem cannot be processed and manual intervention is required.

[0116] For example, the standard word library can be expressed as follows:

[0117] English word: AMOUNT

[0118] English abbreviation: AMT

[0119] Chinese translation: amount

[0120] Chinese description (note): AMOUNT

[0121] Data type: BigDecimal

[0122] Length (precision): 2

[0123] The standard phrase library (data standard) is composed of standard words. The data standard field is composed of standard words. It can be expressed in the following ways:

[0124] English field name: REPAY_AMT

[0125] Chinese dictionary name: repayment amount

[0126] Reference standard words: REPAY, AMOUNT

[0127] Data type: BigDecimal

[0128] Length (precision): 2

[0129] In some embodiments, see Figure 6 , Figure 6 This is a process diagram of the data processing method provided in the embodiment of the present application. Figure 6 The specific process of data processing is explained.

[0130] Phase 1: Pass metadata and table data to the calling interface.

[0131] Here, the metadata English name includes English words or word abbreviations.

[0132] Phase 2: Match the English or Chinese fields included in the metadata with the standard phrase library in the data standard library. If all matches are found, the metadata is marked as a standard phrase.

[0133] Here, full match means that the metadata is completely consistent with one of the Chinese and English field names in the standard phrase library.

[0134] Phase 3: If the requirements of Phase 2 are not met, the English field is segmented using underscores (_) or camelCase. The segmented English array is translated into Chinese, and the Chinese in the field is segmented, parsed, and duplicates are removed. The phrases are then combined into Chinese phrases (including segmentation within the specified word roots (English words and abbreviations) and reference to third-party word segmentation tools). The Chinese and English phrases are then matched against the standard word root library. If the ratio of the number of matches to the number of phrases after segmentation (matching degree) is greater than or equal to 75% (the ratio can be adjusted according to the accuracy), a new standard phrase is created for review. After review, the field metadata is marked as a pre-created standard phrase.

[0135] Here, the definition of the standard phrase library can be achieved through manual intervention or through program-automated review.

[0136] Stage 4: If the conditions in Stage 3 are not met, and the matching ratio in Stage 3 is less than 75%, the Chinese and English phrases are NLP-translated, and then the translated Chinese and English phrases are matched against the standard word roots respectively. If the ratio of the number of matches to the number of translated phrases (matching ratio) is greater than or equal to 50%, it is marked as a pre-created standard phrase.

[0137] Here, NLP translation refers to adjusting the semantic order of Chinese and English words to make their expressions more standardized and the word segmentation results different.

[0138] Stage 5: If the conditions in stage 4 are not met, customized rules are applied to the table data for metadata with a matching degree between 25% and 50% (the range can be adjusted according to the accuracy) in stage 4. For example, specific rules such as ID number, mobile phone number, email address, address, name (all surnames), and URL are used for verification. If the table data meets any rule and the ratio of the amount of data meeting the verification rule to the table data is greater than 80%, the field is marked as a standard phrase corresponding to the rule.

[0139] Here, table data refers to the data stored in the data table. Assuming that there are 20 data in the data table, when 16 data satisfy a rule at the same time, it can be determined that the data table meets the data phrase corresponding to the rule.

[0140] Stage 6: If the requirements of stage 5 are not met, the field metadata English name, Chinese name (including translated ones), and table data are input into the word2vec word vector model. If the output data matches the standard library at a rate greater than 50%, the field metadata is marked as a pre-created standard phrase.

[0141] Here, before using the word vector model to convert each word in the first word set and the second word set and the data in the data table into corresponding word vectors, the word vector model needs to be trained to obtain a large text dataset, where the dataset can be a book, article, web page, or any other text source; clean and format the text, for example, remove punctuation, convert to lowercase, delete stop words, etc., to obtain sample data; select a suitable word vector model algorithm, such as Word2Ve c, GloVe, FastText or BERT; create a model instance and load the preprocessed sample data, traverse all sentences in the sample data, and for each sentence, use the model algorithm to update the word vector, repeat the above vector update process until the model converges or reaches a predetermined number of iterations. During the training process, use the change of the loss function to evaluate the model performance, and finally save the trained model for subsequent use.

[0142] Stage 7: The calculated data that do not meet the requirements of stages 3, 4, 5, and 6 are used as manual references (the analysis results of the fields that can be processed are used as references for manual determination of standard phrases).

[0143] Stage 8: Repeat stages 1 to 7 to continuously adjust the percentage values ​​to make the data processing method of this application more accurate.

[0144] The data processing method provided in this application can expand the standard word library, standard phrase library and bind them with metadata, and continuously improve the association between the data standard library and metadata.

[0145] The following continues to describe the exemplary structure of the data processing device 543 provided in the embodiment of the present application implemented as a software module. In some embodiments, such as Figure 2 As shown, the software modules stored in the data processing device 543 of the memory 540 may include: a matching module 5431 , a processing module 5432 and a creation module 5433 .

[0146] The matching module 5431 is used to match the metadata of the data table with the standard phrase library to obtain a first matching degree; the processing module 5432 is used to process the metadata to obtain a first word set when the first matching degree is less than a first threshold; the matching module 5431 is also used to match the first word set with the standard root library to obtain a second matching degree; the creation module 5433 is used to create a first standard phrase based on the first word set when the second matching degree is greater than or equal to the second threshold.

[0147] In some embodiments, the processing module 5432 is further used to segment the metadata to obtain multiple first words; generate multiple second words based on the multiple first words, where the language of the first words is different from the language of the second words; and generate a first word set based on the multiple first words and the multiple second words.

[0148] In some embodiments, the processing module 5432 is further used to process the words included in the first word set to obtain a second word set when the second matching degree is less than a second threshold; the matching module 5431 is further used to match the second word set with the standard root library to obtain a third matching degree; the creation module 5433 is further used to create a second standard phrase based on the second word set when the third matching degree is greater than or equal to a third threshold.

[0149] In some embodiments, the data processing device 543 further includes a determination module 5434 and a marking module 5435. The determination module 5434 is configured to determine a verification rule corresponding to the data table when the third matching degree is less than a third threshold and the third matching degree is greater than or equal to a fourth threshold; and the marking module 5436 is configured to mark the metadata as a standard phrase corresponding to the verification rule when the first number of data in the data table that meets the verification rule is greater than or equal to the fourth threshold.

[0150] In some embodiments, the determination module 5434 is also used to determine the verification rules corresponding to the type of data in the data table to obtain multiple verification rules; determine the second quantity of data corresponding to each verification rule; and based on the second quantity, determine the verification rules corresponding to the data table from multiple verification rules.

[0151] In some embodiments, the processing module 5432 is further used to perform word vector conversion on the first word set, the second word set and the data to obtain a word vector when the third matching degree is less than a fourth threshold; the matching module 5431 is further used to determine the similarity between the word vector and the word vector corresponding to the standard library, and use the similarity as the fifth matching degree, wherein the standard library includes a standard phrase library and a standard root library; the marking module 5435 is further used to mark the metadata as a third standard phrase when the fifth matching degree is greater than a sixth threshold.

[0152] In some embodiments, the data processing device 543 also includes an update module 5436, which is used to mark the new standard phrase to be reviewed as a pre-created standard phrase after the new standard phrase to be reviewed is reviewed and passed; add the pre-created standard phrase to the standard library to obtain an updated standard library; and update the threshold based on the updated standard library.

[0153] In some embodiments, the data processing device 543 also includes an acquisition module 5437, which is also used to acquire a dictionary library, wherein the dictionary library contains multiple words and multiple phrases; the determination module 5434 is also used to determine the mapping relationship between each word or phrase in the dictionary library and the standard library, wherein the standard library includes a standard phrase library and a standard root library; the marking module 5435 is also used to mark the metadata corresponding to the standard library as a standard phrase based on the mapping relationship.

[0154] In some embodiments, the marking module 5435 is further configured to mark the metadata as a standard phrase when the first matching degree is greater than or equal to a first threshold.

[0155] It should be noted that the description of the device in the embodiment of the present application is similar to the description of the method embodiment above, and has similar beneficial effects as the method embodiment, so it will not be repeated here. Figure 3 ,or Figure 4 The present invention should be understood by referring to the description of any one of the accompanying drawings.

[0156] An embodiment of the present application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the data processing method described in the embodiment of the present application.

[0157] The embodiment of the present application provides a computer-readable storage medium in which computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the data processing method provided in the embodiment of the present application, for example, Figure 3 ,or Figure 4 The data processing method is shown.

[0158] In some embodiments, the computer-readable storage medium may be a ferroelectric random access memory (FRAM), ROM, programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic surface memory, optical disk, or compact disc read-only memory (CD-ROM); or various devices including one or any combination of the above memories.

[0159] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0160] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).

[0161] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.

[0162] In summary, through the embodiments of the present application, the metadata and standard library in the data table are analyzed, the data standard library is established and improved, and corresponding tags are established between the data standards and the metadata, which is conducive to the quality inspection and rectification of the metadata standards, and can more conveniently manage and utilize data, thereby improving work efficiency.

[0163] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.

Claims

1. A data processing method, characterized in that: The method comprises: Matching the metadata of the data table with the standard phrase library to obtain a first matching degree; When the first matching degree is less than a first threshold, processing the metadata to obtain a first word set; Matching the first word set with a standard root library to obtain a second matching degree; In a case where the second matching degree is greater than or equal to a second threshold, a first standard phrase is created based on the first word set.

2. The method according to claim 1, characterized in that The processing of the metadata to obtain a first word set includes: Segmenting the metadata to obtain a plurality of first words; Based on the plurality of first words, generating a plurality of second words, wherein the language of the first words is different from the language of the second words; The first word set is generated based on the plurality of first words and the plurality of second words.

3. The method according to claim 1 or 2, characterized in that The method further comprises: When the second matching degree is less than the second threshold, processing the words included in the first word set to obtain a second word set; Matching the second word set with the standard root library to obtain a third matching degree; In a case where the third matching degree is greater than or equal to a third threshold, a second standard phrase is created based on the second word set.

4. The method according to claim 3, characterized in that The method further comprises: When the third matching degree is less than the third threshold value and the third matching degree is greater than or equal to a fourth threshold value, determining a verification rule corresponding to the data table; When a first number of data in the data table that meets the verification rule is greater than or equal to a fourth threshold, the metadata is marked as a standard phrase corresponding to the verification rule.

5. The method according to claim 4, characterized in that Determining the verification rule corresponding to the data table includes: Determine verification rules corresponding to the types of data in the data table to obtain multiple verification rules; Determining a second amount of data corresponding to each of the verification rules; Based on the second number, a verification rule corresponding to the data table is determined from the multiple verification rules.

6. The method according to claim 4 or 5, characterized in that The method further comprises: When the third matching degree is less than the fourth threshold, performing word vector conversion on the first word set, the second word set, and the data to obtain a word vector; Determining a similarity between the word vector and a word vector corresponding to a standard library, and using the similarity as a fifth matching degree, wherein the standard library includes the standard phrase library and the standard root library; When the fifth matching degree is greater than a sixth threshold, the metadata is marked as a third standard phrase.

7. A data processing device, characterized in that: The device comprises: A matching module, configured to match metadata of the data table with a standard phrase library to obtain a first matching degree; a processing module, configured to process the metadata to obtain a first word set when the first matching degree is less than a first threshold; The matching module is further configured to match the first word set with a standard root library to obtain a second matching degree; A creating module is configured to create a first standard phrase based on the first word set when the second matching degree is greater than or equal to a second threshold.

8. An electronic device, characterized in that: include: a memory for storing computer-executable instructions or computer programs; The processor is configured to implement the data processing method according to any one of claims 1 to 6 when executing the computer-executable instructions or computer programs stored in the memory.

9. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer-executable instructions or computer programs are executed by a processor, the data processing method according to any one of claims 1 to 6 is implemented.

10. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the data processing method according to any one of claims 1 to 6 is implemented.