Data standard acquisition method, computer equipment and storage medium

By obtaining metadata information, and filtering out standard fields that meet the conditions, the problem of inconsistent metadata in different systems is solved, and the accuracy and acquisition efficiency of data standards are improved.

CN120492425AActive Publication Date: 2025-08-15ZHEJIANG DAHUA TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510407384.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-08-15
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

The metadata definitions of different systems vary, resulting in complex and inefficient data processing across systems, and the lack of unified data standards, affecting the accuracy of data understanding and usage.

Method used

By obtaining relevant information of metadata, a transition data standard is generated, and a standard field that meets the preset filtering conditions is determined using the standard generation model to generate target data standards.

Benefits of technology

It improves the accuracy and acquisition efficiency of data standards, reduces user-defined methods, and ensures the accuracy and consistency of data standards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492425A_ABST
    Figure CN120492425A_ABST
Patent Text Reader

Abstract

The invention discloses a data standard acquisition method, computer equipment and a storage medium. The method comprises the following steps: acquiring related information of metadata; generating a transition data standard by using the related information of the metadata; wherein the transition data standard comprises a plurality of standard fields; determining a standard field meeting a preset screening condition from the transition data standard to obtain a screening standard field; and determining a target data standard based on the screening standard field. According to the scheme, the accuracy of obtaining the data standard can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a method for obtaining data standards, a computer device, and a storage medium. Background Art

[0002] In the process of information construction, there are usually a variety of data that need to be processed. Since data in different systems are usually developed and maintained independently, the metadata definitions of different systems are different. There is a lack of unified data standard definitions, which makes cross-system data processing complex and inefficient, and easily leads to errors in data understanding and use.

[0003] The current way of defining data standards has problems such as low accuracy or low efficiency. Summary of the Invention

[0004] The main technical problem solved by this application is to provide a method for obtaining data standards, a computer device and a storage medium, which can improve the accuracy of obtaining data standards.

[0005] In a first aspect, the present application provides a method for obtaining data standards, the method comprising: obtaining relevant information of metadata; using the relevant information of metadata to generate a transitional data standard; wherein the transitional data standard includes a plurality of standard fields; determining, from the transitional data standard, a standard field that meets preset filtering conditions to obtain a filtering standard field; and determining a target data standard based on the filtering standard field.

[0006] A second aspect of the present application provides a computer device, which includes a memory and a processor coupled to each other, wherein the memory stores program data, and the processor is used to execute the program data to implement any step of the above-mentioned data standard acquisition method.

[0007] A third aspect of the present application provides a computer-readable storage medium, which stores program data that can be executed by a processor, and the program data is used to implement any step of the above-mentioned data standard acquisition method.

[0008] The above scheme obtains relevant information of metadata; uses the relevant information of metadata to generate transition data standards, wherein the transition data standards include several standard fields, which can obtain standard fields with relatively accurate metadata, and then, from the transition data standards, determine the standard fields that meet the preset filtering conditions to obtain the filtering standard fields, and determine the target data standards based on the filtering standard fields. The filtering standard fields can further determine more accurate target data standards, which can improve the accuracy of obtaining data standards, reduce the number of user-defined methods for determining data standards, and improve the efficiency of obtaining data standards.

[0009] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions of this application, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be derived from these drawings without inventive effort. Among them:

[0011] Figure 1 This is a flow chart of an embodiment of a method for obtaining data standards of this application;

[0012] Figure 2 This application Figure 1 A flow chart of an embodiment of step S11;

[0013] Figure 3 This application Figure 1 A flow chart of an embodiment of step S12;

[0014] Figure 4 This application Figure 1 A flow chart of an embodiment of step S14;

[0015] Figure 5 It is a structural diagram of an embodiment of the data processing system of the present application;

[0016] Figure 6 This is a structural diagram of an embodiment of a device for obtaining data standards of the present application;

[0017] Figure 7 This is a schematic structural diagram of an embodiment of a computer device of the present application;

[0018] Figure 8 It is a structural diagram of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION

[0019] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0020] The terms "first" and "second" in this application are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units that are inherent to these processes, methods, products or devices.

[0021] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0022] The term "and / or" in this article is simply a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects are in an "or" relationship. In addition, "many" in this article means two or more than two. In addition, the term "at least one" in this article means any combination of at least two of any one or more of a plurality of. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0023] This application provides the following embodiments, and each embodiment is described in detail below.

[0024] See also Figure 1 , Figure 1 This is a flow chart of an embodiment of a method for obtaining data standards of this application. The method may include the following steps:

[0025] S11: Obtain metadata related information.

[0026] The original metadata may be collected to obtain metadata-related information based on the original metadata.

[0027] In some embodiments, the original metadata includes several original fields. For example, the original metadata includes metadata information such as table name, table note / description information, field name, field type and length, field note / description, etc. This application does not impose any restrictions on this.

[0028] In some embodiments, metadata-related information obtained based on original metadata includes at least one of the following: original metadata, metadata samples, and metadata features, wherein metadata samples are obtained by sampling original metadata, and metadata features are obtained by feature extraction of original metadata.

[0029] In some embodiments, see Figure 2 , step S11 of the above embodiment can be further expanded. To obtain metadata related information, this embodiment may include at least one of the following steps:

[0030] S111: Collect data from the original database to obtain original metadata.

[0031] Data can be collected from a raw database. The raw database can refer to multiple raw databases and can contain multiple raw data. For example, a database for managing various task types, a database for various table types, etc. can be an online database or an offline database. This application does not limit the raw database.

[0032] In some embodiments, a network interface can be used to connect to the original database, thereby accessing the original database and collecting data from the original database to obtain original metadata with access rights, such as a table list, table description, table structure (e.g., including detailed field names, field types, field descriptions, indexes, and other metadata information), and table data content. The network interface can be determined based on the specific original database. For example, the network interface can include an interface such as JDBC (Java Database Connectivity), which is not limited in this application.

[0033] S112: Extract features from the original metadata to obtain metadata features.

[0034] Collecting features of the raw metadata, that is, extracting feature values from the raw metadata to obtain metadata features. Optionally, metadata features can be statistical values, characteristic values, etc. For example, metadata features include maximum value, minimum value, median value, maximum / minimum data length, etc. This application does not impose any restrictions on metadata features.

[0035] S113: Sampling the original metadata using a preset sampling strategy to obtain metadata samples.

[0036] A set number of metadata samples can be collected. Specifically, a preset sampling strategy can be used to sample the original metadata to obtain metadata samples. In this way, the meaning of the original fields of the original metadata can be better understood through the metadata samples. For example, collecting metadata samples can facilitate the standard generation model to better understand the meaning of the data in subsequent steps and generate more accurate data standards. The standard generation model can be a large model, a neural network model, etc. In this application, a large model is taken as an example for illustration, and this application does not limit this.

[0037] Exemplarily, by collecting metadata samples, when the data definitions of the relevant information of the metadata are not standardized and the meaning cannot be understood through the original fields (such as field names, descriptions, etc.), the standard generation model can infer and understand the meaning according to the actual values of the original fields. For example, when the original fields have values such as "ZheXXXX", "WangXXX", "XX Province XXX Number", etc., the standard generation model can understand that the original fields are "license plate number", "name of the target object", "address" respectively through the metadata samples. In this way, the standard generation model can generate more accurate data standards.

[0038] In some embodiments, the preset sampling strategy includes: after de-duplicating the original fields in the original metadata, collecting a number of metadata samples; and / or, in the original metadata, collecting the first number of non-empty original fields to obtain metadata samples; and / or, in the original metadata, collecting the second number of non-empty original fields in a preset random manner to obtain metadata samples. Optionally, several preset sampling strategies can be included. The preset sampling strategies can avoid collecting null values and can determine the preset sampling strategy according to the data volume. This application does not limit the preset sampling strategy.

[0039] Optionally, after de-duplicating the original fields in the original metadata, a number of metadata samples are collected. For example, in the original fields of the original metadata, for the string type, according to the total data volume of the table, the original fields (such as field names, etc.) can be de-duplicated and then all data or a small number of samples can be collected to obtain a number of metadata samples.

[0040] Exemplarily, taking the corresponding SQL (Structured Query Language) as an example, the sampling process can be expressed as:

[0041] SELECT DISTINCT FiledName FROM Tbl;

[0042] Where FieldName represents the field name and Tbl represents the table name. The above SQL statement means collecting all values of the original fields and de-duplicating them to obtain metadata samples.

[0043] Optionally, the original fields of the original metadata can be sorted according to a preset sorting rule, thereby collecting the first number of original fields with non-null values in the sorted original metadata to obtain a metadata sample. The preset sorting rules include: sorting by field name, sorting by the string of the original field, etc., which are not limited in this application.

[0044] For example, taking the corresponding SQL as an example, the sampling process can be expressed as:

[0045] SELECT DISTINCT FieldName FROM Tbl WHERE A IS NOT NULL LIMIT 10;

[0046] Where FieldName represents the field name, Tbl represents the table name, and A represents the original field. The above SQL statement collects the first 10 non-empty values of the FieldName field (sorted by string by default) as metadata samples.

[0047] Optionally, a second number of original fields with non-null values may be collected from the original metadata in a preset random manner to obtain a metadata sample.

[0048] For example, taking the corresponding SQL as an example, the sampling process can be expressed as:

[0049] SELECT DISTINCT A FROM Tbl WHERE A IS NOT NULL ORDER BY RAND()LIMIT10;

[0050] Tbl represents the table name, A represents the original field, and the above SQL statement means collecting 10 non-null values in random order as metadata samples.

[0051] It is understandable that the present application does not limit the execution order of the above steps S111 to S113. For example, at least some of the steps may be executed simultaneously or each step may be executed in sequence, etc. The execution order of each step is not limited to this.

[0052] Through the above method, relevant information of various types of metadata can be collected.

[0053] Continue reading Figure 1 After the above step S11, the following steps are also included:

[0054] S12: Generate a transition data standard using the relevant information of the metadata; wherein the transition data standard includes several standard fields.

[0055] Data standards can be defined based on the relevant information of the collected metadata to generate transitional data standards. The transitional data standards can include several standard fields, which are also data standards defined for the original fields. During the generation of transitional data standards, the metadata information can be processed using a standard generation model to generate transitional data standards.

[0056] In some embodiments, at least one of the above-mentioned original metadata, metadata samples, metadata features, etc. can be input into a standard generation model, wherein the standard generation model can be a large model, such as an AI (Artificial Intelligence) large model. The standard generation model can understand or infer the meaning of the fields in each relevant information, and regenerate and define a standard field that meets the specifications to obtain a transition data standard, which can include several standard fields (hereinafter referred to as standard field V1 or first standard field).

[0057] In some embodiments, see Figure 3 The step S12 of the above embodiment can be further expanded. By using the metadata related information to generate the transition data standard, this embodiment may include the following steps:

[0058] S121: Determine first prompt information based on relevant information of metadata.

[0059] Based on the relevant information of the above metadata, the first prompt information for the standard generation model can be determined. The first prompt information includes information such as metadata related information, input prompt information, output prompt information, execution task information, preset definition data and / or preset field rules. Exemplarily, the first prompt information is used to prompt the standard generation model with information such as the context of input information, parameter information input to the standard generation model, execution task of the standard generation model, output format information, etc., and may include content-related information of various types of information, data format-related information, etc. This application does not limit the first prompt information.

[0060] In some implementations, metadata-related information may be input into a preset prompt model (e.g., Prompt) to obtain the first prompt information. Specifically, at least one of the above-mentioned original metadata, metadata sample, metadata feature, etc. may be input into the preset prompt model (e.g., Prompt) to obtain the first prompt information.

[0061] For example, the collected original metadata, metadata samples, and metadata features are entered into the prompt, such as the original field name SFZ, the original field type String, and the original field sample data. Optionally, if the original field type is a numeric type, the collected metadata features may also be included, such as metadata features containing data feature information such as "maximum value," "minimum value," and "average value." This application does not impose any restrictions on this.

[0062] As an example, the first prompt information is provided below:

[0063] The user is an expert in data governance. Their task is to infer the meaning of the fields based on the original data's table names and comments, the Chinese and English field names and comments, and sample data. They then redefine field names that conform to naming standards, generate corresponding comments, and assign a confidence level (e.g., confidence levels range from 0 to 1). If the original field names and comments already conform to the standards, they can be used directly or partially referenced.

[0064] The naming convention is as follows:

[0065] 1. Use an English name;

[0066] 2. Use camel case between multiple words;

[0067] 3. Capitalize each first letter;

[0068] 4. The field type shall be based on the field type of the MySQL database;

[0069] 5. The field length can refer to the field length of the original table and be adjusted appropriately based on the data meaning and relevant experience.

[0070] The output format example is as follows:

[0071] Original field name: GongSiDiZhi

[0072] Standard field name: CompanyAddress

[0073] Standard field type: String

[0074] Standard field length: 128

[0075] Standard Field Notes: Work Address

[0076] Confidence level: 0.91

[0077] Please provide a final answer, no further explanation is required.

[0078] Start (the following are the original fields of the original metadata input):

[0079] Original table information:

[0080] Original table name: GMSFXX

[0081] Original table description: Target object information table

[0082] Field information:

[0083] {Original field name: SFZ

[0084] Original field type: String

[0085] Original field length: 18

[0086] Original field annotation:

[0087] Original field data example: 330102201706042336,330204198703125023},

[0088] {

[0089] Original field name: NL

[0090] Original field type: Int

[0091] Original field length: 2

[0092] Original field annotation:

[0093] Original field sample values: 36, 45, 87, 12

[0094] Original field average: 37

[0095] Original field maximum value: 104

[0096] Minimum value of the original field: 0

[0097] }.

[0098] For example, by processing the standard generation model, a generated standard field V1 can be obtained, and the output example is as follows:

[0099] {

[0100] Standard table name: CitizenInfo

[0101] Standard table description: records the attribute information of the target object, including the target object's ID number, name, category, anniversary, address and other basic information.

[0102] }

[0103] {

[0104] Original field name: SFZ

[0105] Standard field name: IdentityCardNumber

[0106] Standard field type: String

[0107] Standard field length: 18

[0108] Standard Field Notes: ID Number

[0109] Confidence: 1

[0110] };

[0111] {

[0112] Original field name: NL

[0113] Standard field name: Age

[0114] Standard field type: Int

[0115] Standard field length: 2

[0116] Standard field annotation: ID number of the target object

[0117] Confidence level: 0.98

[0118] }.

[0119] In some embodiments, the relevant information of the metadata is structurally converted to obtain relevant information after structural conversion; the relevant information after structural conversion is input into a preset prompt model to obtain first prompt information. For example, taking the original table of the original metadata as a unit, the original metadata of the collected original table (such as the name / type / description of the table and field, etc.) is converted and organized into relevant information of a preset structure, such as converted into Json, Xml, Markdown format or other structured / semi-structured input format that is easy for large models to understand, etc., to obtain relevant information after structural conversion, and all fields of a single table can be input at one time. Then, the relevant information after structural conversion is input into the preset prompt model to obtain the first prompt information.

[0120] In some embodiments, to further improve recognition accuracy, the first prompt information includes preset definition data and / or preset field rules. For example, some preset definition data and / or preset field rules can be predefined in the preset prompt model Prompt to obtain the first prompt information including the preset definition data and / or preset field rules. The preset field rules can include rules for information such as the type, value range, and data format of the original field. The preset definition data includes the data type of predefined metadata-related information, such as the data type of metadata features of different metadata original fields. This application does not impose any restrictions on this.

[0121] For example, some attribute information of target objects may be predefined. Examples of preset field rules for the attribute information are as follows:

[0122] Anniversary: 0~120 (anniversary);

[0123] Height: 40-240 (cm), or 0.4-2.4 (m);

[0124] Height: 2-150 (kg / kg);

[0125] For example, examples of predefined metadata (such as metadata features of a target object) may be predefined as follows:

[0126] The attribute characteristics of the target object include category, anniversary, height, altitude, appearance status, etc.;

[0127] The decoration features of the target object include the decoration status of multiple parts, such as the decoration status of the head, the decoration status of the upper part, the decoration status of the lower part, the decoration status of each branch, etc.

[0128] It is understandable that the first prompt information can be determined according to a specific application scenario, and this application does not impose any restrictions on the first prompt information.

[0129] S122: Process the first prompt information using a standard generation model to obtain a transition data standard.

[0130] The above-mentioned first prompt information is input into the standard generation model, and the first prompt information is processed by the standard generation model to obtain the transition data standard (standard field V1). In this way, when the field definition / description of the original metadata is unclear, the metadata features, preset field rules, preset definition data, etc., combined with the metadata sample, can be used to further understand and infer the data meaning of the field. For example, the aforementioned NL original field is difficult to understand because its value is a numeric type. However, by providing the original table information (target object information table), combined with the actual data value sample and data range (such as maximum, minimum, median, etc.) of the original table, and combined with the preset field rules given in the first prompt information Prompt (such as data rules for anniversary, height, weight, etc.), the standard generation model can more accurately infer the true meaning of the field (such as the anniversary of the target object) and generate a standard field definition for it.

[0131] In some embodiments, the first prompt information is processed using a standard generation model, and the original metadata can be processed to output a transition data standard, a mapping relationship between the original field and the standard field, and a confidence level between the original field and the standard field (for example, the confidence level can represent similarity).

[0132] In some embodiments, after generating the transition data standard using the relevant information of the metadata, the mapping relationship between the original field and each standard field in the transition data standard can be recorded. Among them, the original metadata can contain several original fields, the transition data standard can contain several standard fields, and each original field / standard field can contain several field definitions. In this way, the mapping relationship between the original field of the original metadata and the standard field of the transition data standard can be obtained. Optionally, the mapping relationship can also represent the corresponding relationship between the original field and the standard field. Optionally, the mapping relationship can include the relationship between the field definition, index, etc. corresponding to the original field and the standard field. This application does not limit the mapping relationship.

[0133] In some implementations, the mapping rules can be determined first. For example, direct mapping: when the original field and the standard field have the same meaning, they are directly mapped. Conversion mapping: the original field needs to be converted in format or unit to correspond to the standard field. Merge / split mapping: multiple original fields are merged or one original field is split into multiple standard fields. Default value mapping: when the original field is missing, a default value is set for the standard field. Then, the mapping relationship between the original field and the standard field is determined according to the mapping rules, and a mapping table is created. For example, a table or document is used to record the mapping relationship between each original field and the standard field, which can include its mapping rules / conversion rules.

[0134] For example, assuming that the original field SFZ comes from the table Table1 of the original metadata, obtaining the mapping relationship can record the following information: the field name and field description of the standard field V1 and the index established therefor, where the index includes a vector index and a text index.

[0135] For example, the record table corresponding to the standard field V1 of the transition data standard is as follows:

[0136]

[0137] For example, the mapping relationship table corresponding to the mapping relationship between the original field and the standard field

[0138] Table 2 is as follows:

[0139]

[0140] in:

[0141] The "Standard Field V1 Serial Number" in Table 2 corresponds to the "Serial Number" in Table 1 and is the primary key or unique identifier of Standard Field V1. The "Standard Field V1 Name" in Table 2 is optional and is used only for redundancy. The "Metadata Sample" field value in Table 1 is sample data collected from the original metadata. This field is used to facilitate user understanding of the standard field definitions of the data and to improve the accuracy of merging identical fields when subsequently generating target data standards.

[0142] In some embodiments, when inserting a record of the above mapping relationship into the mapping relationship table, it is first determined whether the field definition of the same standard field already exists. If it already exists, there is no need to add a duplicate record to the above Table 1, and only a corresponding relationship needs to be added to Table 2.

[0143] It should be noted that the above "Table 1" and "Table 2" are only used as examples. The mapping relationship between the standard fields and the original fields and the implementation method can be determined according to the specific application scenario. This application does not impose any restrictions on this.

[0144] In some implementations, the above-mentioned mapping relationship can be used for subsequent data processing, for example, for metadata governance, metadata update, data standard update, etc., and this application does not impose any restrictions on this.

[0145] Continue reading Figure 1 After the above step S12, the following steps are also included:

[0146] S13: Determine the standard fields that meet the preset screening conditions from the transition data standards to obtain the screening standard fields.

[0147] By generating standard field definitions from metadata-related information, a transitional data standard is obtained, and the target data standard can be determined based on the transitional data standard. Optionally, the transitional data standard can be used as the target data standard, or standard fields that meet preset screening conditions can be determined based on the transitional data standard to obtain screening standard fields, and then the target data standard can be determined based on the screening standard fields, thereby obtaining a more accurate standard field definition.

[0148] In some embodiments, due to constraints such as the context length of the standard generation model (such as the large model), the above steps are used to input the standard generation model to understand and define standard fields for all original fields at the granularity of a single original table. However, due to the randomness of the content generated by the large model (for example, the same or similar question input may result in similar but not completely identical results when asked multiple times), and the fields with the same meaning in different original tables have different original names and the characteristics of the extracted metadata samples are also different. Therefore, when processing original fields with similar meanings, the field definitions of the standard fields generated each time may not be exactly the same. For example, suppose there are two original tables, "target object information table" and "target object related information table", both of which contain the original field of the target object's ID number, but the names of the corresponding original fields in the two original tables are SFZ and Id respectively. According to the method of the above steps, it is necessary to call the standard generation model twice to define the corresponding standard field for each field of the two tables. In actual operation, it is possible that the standard generation model has certain differences in the standard field definition of "target object ID number", such as: IdentifyCardNumber and IdentifyCard respectively. This approach may result in the standard definitions of fields with the same meaning, although each conforms to the naming convention, still not being unified. Therefore, it is possible to filter out standard fields from the transitional data standards, that is, to determine the standard fields that meet the preset filtering conditions and obtain the filtering standard fields. In this way, standard fields with the same or similar meanings can be filtered out for further integration, so that the final standard fields can be regenerated for the standard fields with the same or similar meanings.

[0149] In some embodiments, a first similarity between several standard fields in the transition data standard can be obtained, and then, using the first similarity, a standard field that meets the preset screening conditions is selected to obtain a screening standard field. Each standard field contains several field definitions. Optionally, the first similarity is the sum of the sub-similarity between at least some field definitions or preset field definitions in several field definitions. Optionally, sub-similarity is obtained for the first field definition (such as field description, field name, etc.) and the second field definition (such as metadata sample, preset field rule, etc.), and then different weights are used for weighted synthesis to obtain the first similarity, wherein the weight of the first field definition is greater than or equal to the weight of the second field definition.

[0150] In some embodiments, the preset screening condition includes at least one of the following: the first similarity is greater than a first similarity threshold, and a preset number before the first similarity sorting.

[0151] Specifically, several standard fields in the transition data standard can be obtained, each standard field containing several field definitions. The similarity between any two standard fields can be obtained to obtain a first similarity. The first similarity can be obtained by summing the sub-similarities between the field definitions contained therein. Then, the fields are sorted according to the first similarity, and standard fields whose first similarity exceeds a first similarity threshold and / or a preset number before the first similarity sorting are selected to obtain the screening standard fields.

[0152] For example, a record of a standard field V1 can be retrieved from Table 1, which may include a field definition: field name, field meaning, metadata sample, etc. Then, based on the vector index and text index corresponding to each of the above field definitions, a first similarity is obtained to determine all other similar standard fields. The top N similarities are selected based on the first similarity to obtain the screening standard field, where N is an integer.

[0153] For example, regarding the first similarity and its ranking, if the standard field includes field definitions: field name, field meaning, etc., the cosine similarity (sub-similarity K1) between the "field name" and the cosine similarity (sub-similarity K2) between the "field meaning" in the two standard fields are calculated. The first similarity is obtained by combining sub-similarities K1 and K2. The top N standard fields are sorted according to the first similarity values to obtain the screening standard fields.

[0154] It is understandable that, according to the above-mentioned screening method, each standard field can be screened in turn, thereby screening out several groups of screening standard fields with the same or similar meanings.

[0155] Continue reading Figure 1 After the above step S13, the following steps are also included:

[0156] S14: Determine target data standards based on the screening criteria fields.

[0157] The screening criteria fields are processed to determine the target data standard. Optionally, screening criteria fields with the same or similar meanings can be merged to generate a final standard field to obtain the target data standard, which includes several standard fields (hereinafter referred to as standard field V2 or second standard field or new standard field).

[0158] In some implementations, the screening standard field (such as the definition of the top N fields in the screening ranking) can be input into the standard generation model to regenerate a final standard field to obtain the target data standard.

[0159] The above scheme obtains relevant information of metadata; uses the relevant information of metadata to generate transition data standards, wherein the transition data standards include several standard fields, which can obtain standard fields with relatively accurate metadata, and then, from the transition data standards, determine the standard fields that meet the preset filtering conditions to obtain the filtering standard fields, and determine the target data standards based on the filtering standard fields. The filtering standard fields can further determine more accurate target data standards, which can improve the accuracy of obtaining data standards, reduce the number of user-defined methods for determining data standards, and improve the efficiency of obtaining data standards.

[0160] In some embodiments, see Figure 4 , step S14 of the above embodiment can be further expanded. Based on the screening standard field, the target data standard is determined. This embodiment may include the following steps:

[0161] S141: Determine the second prompt information based on the screening standard field.

[0162] The second prompt information can be determined based on the filtering standard field, wherein the second prompt information can include relevant information of the filtering standard field, execution task information, similarity rules, etc. The method for obtaining at least part of the information of the second prompt information can specifically refer to the method for obtaining the first prompt information. This application does not limit the second prompt information.

[0163] Optionally, the screening standard field, execution task information, similarity rules, preset field rules, input prompt information, output prompt information, etc. can be input into the preset prompt model to obtain the second prompt information. Among them, the execution task information includes selecting a screening standard field as a new standard field or regenerating a new standard field based on the screening standard field and other information, and obtaining the similarity between the original screening standard field and the new standard field to obtain the second similarity. The similarity rule includes a rule for obtaining the second similarity, combining the sum of the sub-similarity between the original screening field and at least part of the field definition or preset field definition (such as field description, field name, etc.) in the new screening field to obtain the second similarity. Alternatively, the sub-similarity is obtained for the first field definition (such as field description, field name, etc.) and the second field definition (such as metadata sample, preset field rule, etc.), and then different weights are used for weighted synthesis to obtain the second similarity, wherein the weight of the first field definition is greater than or equal to the weight of the second field definition. Among them, the preset field rule can include information such as field value range, field value rule, enumeration value definition, etc., and this application does not limit the preset field rule.

[0164] For example, taking the preset prompt model Prompt as an example, the second prompt information is as follows:

[0165] The user is a data governance expert, and the user's task is to define a set of data field specifications for the enterprise.

[0166] For the filter criteria fields, there are now a number of field definitions with similar meanings and standardized naming (the structure of each piece of information is the same).

[0167] Please select or regenerate the best field definition based on the field name, field description, metadata sample, preset field rules, and other information to obtain a new standard field (still maintaining the original information structure). Then, obtain the second similarity between each original screening standard field and the new standard field. When performing the comparison, the sub-similarity of the field description and field name in the field definition is prioritized. Then, the corresponding meaning is inferred by referring to the metadata sample and preset field rules to determine whether they can be summarized as the same standard field and the second similarity is given.

[0168] For example, the input prompt information is as follows:

[0169] enter:

[0170] Field name | Field description | Field type | Field length | Metadata sample | Field value rules

[0171] Name|Name of the target object|String|10|XX1|None

[0172] CitizenName|Name of the target object|String|10|XX2|None

[0173] For example, the output prompt information is as follows:

[0174] Standard field definitions:

[0175] Field name: CitizenName

[0176] Field Description: Name of the target object

[0177] Field type: String

[0178] Field length: 10

[0179] Metadata sample: XX2

[0180] Field value rules: None

[0181] Field similarity:

[0182] Original field: Name

[0183] Standard field: CitizenName

[0184] Similarity: 1

[0185] Original field: CitizenName

[0186] Standard field: CitizenName

[0187] Similarity: 1

[0188] It is understandable that the second prompt information can be determined according to a specific application scenario, and the present application is not limited to this.

[0189] S142: Process the second prompt information using the standard generation model to obtain a reference data standard and a second similarity; wherein the second similarity is the similarity between the screening standard field and the standard field of the reference data standard.

[0190] The second prompt information can be input into the standard generation model and processed using the standard generation model to obtain the reference data standard and the second similarity; wherein the second similarity is the similarity between the screening standard field and the standard field of the reference data standard.

[0191] In some embodiments, the second prompt information is processed using a standard generation model, and the screening standard field can be processed to output the reference data standard, the mapping relationship between the screening standard field and the standard field of the reference data standard, and the similarity between the screening standard field and the standard field of the reference data standard.

[0192] Exemplarily, the second prompt information is input into the standard generation model, wherein the input second prompt information includes a screening standard field (standard field V1), as follows:

[0193] Field name | Field description | Field type | Field length | Metadata sample | Field value rules

[0194] IdentityCardNumber|Target object's ID number|String|18|220103201507263226|None

[0195] IdentityCard|Target object's ID number|String|18|220223198508236253|None Using the standard generation model to process the above content, we can obtain the reference data standard (including several new standard fields V2) and the second similarity, as follows:

[0196]

[0197] S143: Merging standard fields in the reference data standard that meet preset merging conditions to obtain the target data standard.

[0198] The preset merging conditions include at least one of the following: a second similarity greater than a preset threshold, identical standard fields (e.g., identical field names, identical field meanings, etc.), etc. For example, a second similarity greater than 0.9 indicates that the V1 version's screening standard fields / reference data standard's standard fields meet the preset merging conditions. This allows the determination of the standard fields whose second similarity meets the preset merging conditions, and the merging of these standard fields into the reference data standard to obtain the target data standard. Optionally, if the transitional data standard contains unfiltered standard fields, the reference data standard and the transitional data standard after merging can be combined to obtain the target data standard.

[0199] In some embodiments, based on the second similarity corresponding to the standard field in the reference data standard, the standard field whose second similarity is greater than a preset threshold can be determined, and then the screening standard fields corresponding to their mapping relationships can be merged to obtain a merged reference data standard as the target data standard.

[0200] For example, for the reference data standard and the second similarity shown in the above example, the V1 version's screening standard fields IdentityCardNumber and IdentityCard can be merged into the V2 version's standard field IdentityCardNumber. That is, the V2 version's standard field IdentityCardNumber can be used as the final standard field.

[0201] In some embodiments, the target data standard, that is, the standard fields of version V2, can be saved. For example, Table 3 of the target data standard can be obtained by referring to Table 1 generated by the transition data standard, and the structure of the corresponding tables can be the same.

[0202] Alternatively, the merged standard fields can be removed from the transition data standard (e.g., Table 1) to avoid subsequent repeated processing. For example, if the standard field IdentityCard is merged into IdentityCardNumber, IdentityCardNumber needs to be inserted into Table 3 and the standard field IdentityCard needs to be deleted from Table 1.

[0203] Optionally, for each set of screening standard fields, the above steps may be repeated until all standard fields in the above transitional data standard (Table 1) have been processed to obtain the target data standard.

[0204] Optionally, after the merging of the pending standard fields is completed, further screening and confirmation can be performed by the user through interaction with a graphical interface; or after the target data standards are obtained by traversing Table 1, batch confirmation can be performed by the user. In response to the user's confirmation operation, the final target data standards are determined.

[0205] In some embodiments, after obtaining the target data standard, the mapping relationship can be updated based on the target data standard, wherein the standard field V1 of the transition data standard involved in the mapping relationship is updated to the standard field V2 of the target data standard to obtain an updated mapping relationship.

[0206] Optionally, the mapping relationship between the transition data standard and the original metadata obtained above can be traversed, that is, the relationship table corresponding to the mapping relationship (such as Table 2) can be traversed to update the standard field (V1 version) of the transition data standard to the standard field (V2 version) of the target data standard.

[0207] For example, the field definitions of the standard field IdentityCard originally associated with the V1 version in Table 2 are updated to the standard field IdentityCardNumber associated with the V2 version. For example, this process can be executed using SQL statements, as shown below:

[0208] Update table 2 set "name of standard field" = "IdentityCardNumber" where "name of standard field" = "IdentityCard".

[0209] In some embodiments, after obtaining the updated mapping relationship, the updated mapping relationship can be used to perform preset data processing, wherein the preset data processing includes data governance, data update, etc., and further processing can be performed based on the mapping relationship and the target data standard.

[0210] In some embodiments, multiple rounds of data standard definition can be performed to generate different versions of data standards. After determining the target data standard based on the screening standard fields, the target data standard can be used as the current transitional data standard to continue the steps of determining standard fields that meet the preset screening conditions from the transitional data standard to obtain the screening standard fields; and then determining the target data standard based on the screening standard fields to obtain a new target data standard.

[0211] Optionally, the new target data standard may be used as the final target data standard.

[0212] Optionally, the new target data standard may be used as the final target data standard until it meets a preset standard condition. The preset standard condition includes at least one of the following: obtaining the new target data standard a preset number of times, the standard field included in the new target data standard not meeting a preset screening condition, or the new target data standard meeting data standard requirements, etc. This application does not impose any restrictions on this.

[0213] For example, in the process of obtaining the target data standard, the field definitions of versions V1 and V2 can be structurally identical. Therefore, the transitional data standard of version V1 can be used as an intermediate transitional standard state. For example, if there are 100 original fields with the same meaning in the original table, generating a data standard definition that can cover these 100 original fields may be inaccurate or complex. By introducing a transitional data standard of version V1, the 100 original fields with similar meanings in the original table can be initially integrated into 20 standard fields with relatively similar and standardized naming in version V1. Then, from these 20 similar fields, the standard generation model's capabilities are again invoked to integrate and define the final standard fields, which can improve the accuracy of the data standard. Furthermore, the same method can be used to further obtain new target data standards such as V3 or V4 through the above process, so that the number of repeated similar fields is sufficiently small, and the model can accurately identify and summarize the standard field definitions in one go, thereby obtaining the final target data standard.

[0214] In some embodiments, before the above-mentioned step S13, the data volume of the original metadata can be obtained. Optionally, in response to the data volume being less than the first quantity threshold, the transition data standard is determined as the target data standard. Optionally, in response to the data volume being greater than the second quantity threshold and / or not less than the first quantity threshold, step S13 is executed to obtain the target data standard of the V2 version. Optionally, in response to the data volume being greater than the second quantity threshold, the target data standard is executed as the current transition data standard, and the steps of determining the standard fields that meet the preset filtering conditions from the transition data standard to obtain the filtering standard fields; and, based on the filtering standard fields, determining the target data standard to obtain a new target data standard, to obtain the data standard of the V3 / V4 version, etc., are executed. Among them, the first quantity threshold is less than the second quantity threshold, and the second quantity threshold is less than the third quantity threshold.

[0215] In the above embodiment, by utilizing the large model capability of the standard generation model, the data standard definition can be directly formed based on the original metadata, which can avoid or reduce user operations in the process of defining the data standard fields. In addition, by utilizing the capability of the standard generation model, while generating the data standard definition, the mapping relationship between the original field and the standard field can also be directly generated, which is convenient for the data governance required later. In addition, by first obtaining the V1 version of the data standard, and then comparing and merging the V1 version of the data standard to generate the V2 version of the data standard, it is possible to gradually filter and merge, narrow the scope, and improve the accuracy, thereby solving the limitations of the large model and the long context, and ensuring the accuracy of the data standard definition. By using the above method of generating the V1 version and V2 version of the data standard, the data standard definition can be quickly obtained when the amount of original metadata is small and relatively clear; it can also be further introduced according to the same idea when the amount of original metadata is very large and mixed, to further improve the accuracy of the data standard definition.

[0216] By collecting metadata samples, the standard generation model can further understand the inferred meaning based on the data content, improving the accuracy of obtaining standard field definitions. In addition, by collecting metadata characteristics, preset definition data, preset field rules, etc., the standard generation model can combine metadata characteristics, preset definition data, preset field rules, etc. to improve the accuracy of understanding fields, thereby improving the accuracy of obtaining data standards.

[0217] It is understandable that in the above method of the specific implementation method, the writing order of each step does not mean a strict execution order and constitutes any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0218] In some embodiments, the aforementioned data standard acquisition method can be implemented by a data processing system. The data processing system can be used to generate metadata data standards and perform at least a portion of subsequent pre-defined data processing. For example, the data processing system can perform at least a portion of data governance or a metadata management platform. This application is not limited to this.

[0219] See also Figure 5 , Figure 5 2 is a schematic diagram of the structure of an embodiment of the data processing system of the present application. The data processing system 20 includes a metadata management module 21, a persistence module 22, a database module 23 and a standard generation model 24. The modules are interconnected.

[0220] The metadata management module 21 is used to implement various processes involved in the above-mentioned data standard acquisition method, including calling interfaces of the persistence module 22 , the database module 23 and the standard generation model 24 .

[0221] The persistence module 22 is used to store various data involved in the above-mentioned data standard acquisition method, such as metadata information / metadata-related information, etc., for example, including the original database, charts corresponding to the original table, data samples (metadata samples), data fields, field definitions, mapping relationships, etc. This application does not impose any restrictions on this.

[0222] The database module 23 may be an index database for storing indexes of the above data as needed, including but not limited to a vector database, a text retrieval database, etc.

[0223] The standard generation model 24 can be an AI big model, for example, it can be a big model service deployed locally and privately, for example, it can be a direct call to the big model service interface on the Internet. This application does not impose any restrictions on this. The standard generation model 24 is used to generate data standards.

[0224] In some embodiments, the present application further provides a data standard acquisition device for implementing the data standard acquisition method of any of the above embodiments.

[0225] See also Figure 6 , Figure 6 1 is a schematic diagram of an embodiment of a data standard acquisition device of the present invention. The data standard acquisition device 30 includes an acquisition module 31, a generation module 32, a screening module 33 and a target module 34. The modules are interconnected.

[0226] The acquisition module 31 is used to acquire metadata related information.

[0227] The generation module 32 is used to generate a transition data standard using the relevant information of the metadata; wherein the transition data standard includes a plurality of standard fields.

[0228] The screening module 33 is used to determine the standard fields that meet the preset screening conditions from the transition data standards, and obtain the screening standard fields.

[0229] The target module 34 is used to determine target data standards based on the screening standard fields.

[0230] It should be noted that the data standard acquisition device provided in the above embodiment and the data standard acquisition method provided in the above embodiment are based on the same concept. The specific manner in which each module and unit performs operations has been described in detail in the method embodiment and will not be repeated here. In actual applications, the data standard acquisition device provided in the above embodiment can allocate the above functions to different functional modules as needed, that is, divide the internal structure of the device into different functional modules to complete all or part of the functions described above. This application is not limited to this.

[0231] It is understood that the data standard acquisition method in this application can be executed by a computer device. The computer device can be any device with processing capabilities, such as a mobile device, a computer, a server, etc., and this application does not limit this. In some possible implementations, the data standard acquisition method can be implemented by a processor calling program data stored in a memory.

[0232] For the above embodiment, this application provides a computer device, see Figure 7 , Figure 7 1 is a schematic diagram of the structure of an embodiment of a computer device of the present application. The computer device 40 includes a memory 41 and a processor 42, wherein the memory 41 and the processor 42 are coupled to each other, the memory 41 stores program data, and the processor 42 is configured to execute the program data to implement the steps of any embodiment of the method for obtaining data standards described above.

[0233] In this embodiment, the processor 42 may also be referred to as a CPU (Central Processing Unit). The processor 42 may be an integrated circuit chip having signal processing capabilities. The processor 42 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor, or the processor 42 may be any conventional processor.

[0234] The method of the above embodiment can be implemented in the form of a computer program, so this application proposes a computer readable storage medium, please refer to Figure 8 , Figure 8 The computer-readable storage medium 50 stores program data 51 that can be executed by a processor, and the program data 51 can be executed by the processor to implement the steps of any embodiment of the method for obtaining data standards described above.

[0235] The computer-readable storage medium 50 in this embodiment can be a medium that can store program data 51, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, or it can also be a server that stores the program data 51. The server can send the stored program data 51 to other devices for execution, or it can also execute the stored program data 51 itself.

[0236] In some embodiments, the functions or modules included in the device provided in the above embodiments of the present application can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, this application will not go into details here.

[0237] The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other. For the sake of brevity, this application will not go into details here.

[0238] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation methods described above are only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.

[0239] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0240] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0241] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium, which is a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the various embodiments of the present application.

[0242] Obviously, those skilled in the art should understand that the modules or steps of the present application described above can be implemented using a general-purpose computing device. They can be concentrated on a single computing device or distributed across a network consisting of multiple computing devices. Alternatively, they can be implemented using program code executable by the computing device, so that they can be stored in a computer-readable storage medium and executed by the computing device, or they can be fabricated into separate integrated circuit modules, or multiple modules or steps can be fabricated into a single integrated circuit module for implementation. Thus, the present application is not limited to any specific combination of hardware and software.

[0243] The above description is merely an embodiment of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A method for obtaining data standards, characterized in that: include: Get metadata related information; Generate a transition data standard using the metadata related information; wherein the transition data standard includes a plurality of standard fields; Determining, from the transition data standards, standard fields that meet preset screening conditions to obtain screening standard fields; Based on the screening criteria fields, target data criteria are determined.

2. The method according to claim 1, characterized in that The step of generating a transition data standard by utilizing the relevant information of the metadata includes: Determining first prompt information based on the relevant information of the metadata; The first prompt information is processed using a standard generation model to obtain the transition data standard.

3. The method according to claim 2, characterized in that The determining of the first prompt information based on the relevant information of the metadata includes: Inputting the metadata-related information into a preset prompt model to obtain the first prompt information; or, Performing structural transformation on the relevant information of the metadata to obtain relevant information after structural transformation; inputting the relevant information after structural transformation into a preset prompt model to obtain the first prompt information; And / or, the first prompt information includes preset definition data and / or preset field rules.

4. The method according to claim 1, wherein The step of determining a transition standard field that meets a preset screening condition from the transition data standard to obtain a screening standard field includes: Obtaining a first similarity between a plurality of standard fields in the transition data standard; Using the first similarity, selecting a standard field that meets a preset screening condition to obtain a screening standard field; The preset screening condition includes at least one of the following: the first similarity is greater than a first similarity threshold, and the first similarity is ranked before a preset number.

5. The method according to claim 1, characterized in that Determining the target data standard based on the screening standard field includes: Determining second prompt information based on the screening standard field; Processing the second prompt information using a standard generation model to obtain a reference data standard and a second similarity; wherein the second similarity is the similarity between the screening standard field and the standard field in the reference data standard; The standard fields in the reference data standard that meet the preset merging conditions are merged to obtain the target data standard.

6. The method according to claim 5, characterized in that After generating the transition data standard using the metadata related information, the method further includes: Obtaining a mapping relationship between original fields and each standard field in the transition data standard; After merging the standard fields that meet the preset merging conditions in the reference data standard to obtain the target data standard, the method further includes: Updating the mapping relationship based on the target data standard, wherein the standard fields of the transition data standard in the mapping relationship are updated to the standard fields of the target data standard to obtain an updated mapping relationship; The updated mapping relationship is used to perform preset data processing.

7. The method according to claim 1, characterized in that After determining the target data standard based on the screening standard field, the method includes: Taking the target data standard as the current transition data standard, continuing to perform the steps of determining a standard field that meets a preset screening condition from the transition data standard to obtain a screening standard field; and determining the target data standard based on the screening standard field to obtain a new target data standard; The new target data standard is used as the final target data standard, or until the new target data standard meets the preset standard condition, the new target data standard is used as the final target data standard.

8. The method according to claim 1, characterized in that The metadata-related information includes: at least one of original metadata, metadata samples, and metadata features; and obtaining the metadata-related information includes at least one of the following steps: Collect data from the original database to obtain original metadata; Perform feature extraction on the original metadata to obtain metadata features; The original metadata is sampled using a preset sampling strategy to obtain metadata samples; Among them, the preset sampling strategy includes: collecting a number of metadata samples after deduplicating the original fields in the original metadata; and / or, collecting a first number of original fields with non-null values in the original metadata to obtain metadata samples; and / or, collecting a second number of original fields with non-null values in the original metadata according to a preset random method to obtain metadata samples.

9. A computer device, characterized in that: The method comprises a memory and a processor coupled to each other, wherein the memory stores program data, and the processor is configured to execute the program data to implement the steps of the method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that Program data capable of being executed by a processor is stored, and the program data is used to implement the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method for mapping metadata through data standard and related device

    CN117370356A

  • Data standard benchmarking method and device, electronic equipment and readable storage medium

    CN117390170A

  • System for checking the parking position of a vehicle and a control method therefor

    KR1020240000869A

  • Identifying similar field sets using related source types

    US20200042626A1

  • Data standardization method and apparatus, computer device, and storage medium

    WO2020034873A1