Government affair data quality check rule generation method and system based on large language model
Through the method of generating government data quality verification rules based on large language models, using metadata information and government knowledge graphs to retrieve knowledge fragments, dynamically construct prompt words and generate data quality verification rules, solving the problem of inefficient rule generation in the existing technology, and achieving automated and adaptive data quality assurance.
Patent Information
- Application Number
- CN202510086114.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-20
AI Technical Summary
Existing data quality verification rules require the participation of technicians, and cannot be generated automatically and quickly and adapted, resulting in inefficient data processing and analysis.
The government data quality verification rule generation method based on the large language model is adopted. By obtaining the metadata information of the target table and fields, and combining the government knowledge graph to retrieve the knowledge fragments of the historical field with the largest similarity, the prompt words are dynamically constructed and the large language model is passed to generate data quality verification rules.
Automatic generation and adaptive adjustment of data quality verification rules are realized, improving the efficiency of data quality inspection and improvement, and ensuring data reliability.
Smart Images

Figure CN120012756A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data quality verification, and in particular to a method and system for generating government data quality verification rules based on a large language model. Background Art
[0002] In the context of the digital economy, government data has the remarkable characteristics of extensiveness, authority and real-time nature. Ensuring the quality of government data is crucial to improving decision-making accuracy, enhancing the level of public services, promoting the digital transformation of government affairs, and promoting the development of the data economy.
[0003] Traditional data quality check rules are mainly written manually, which is inefficient and prone to errors. Data quality check rules are usually configured as templates to complete some automated processing, but if the requirements change, technical personnel still need to have a deep understanding of the table structure and the relationship between fields in the data warehouse or data lake in order to write accurate check rules to check data quality. This leads to a high technical threshold for the entire process and requires a lot of time and manpower, resulting in low efficiency in data processing and analysis.
[0004] As large language models have shown powerful language understanding and generation capabilities, technical solutions have gradually emerged that use large language models to convert natural language descriptions into structured SQL retrieval statements to provide users with the required data records. These existing solutions focus on understanding natural language intent and do not take into account the standardization of database table structures and the accuracy of data. Therefore, there is still a lack of solutions that use large language models to generate SQL statements to verify data quality. Summary of the invention
[0005] In view of the above analysis, an embodiment of the present invention aims to provide a method and system for generating government data quality verification rules based on a large language model, so as to solve the problem that existing data quality verification rules require the participation of technical personnel and cannot be automatically and quickly generated and adaptively adjusted.
[0006] On the one hand, an embodiment of the present invention provides a method for generating government data quality check rules based on a large language model, comprising the following steps:
[0007] Obtain the metadata information of the target table to be checked and each target field therein;
[0008] Take out each target field in turn, obtain the field semantic vector, joint semantic vector, field structure vector and / or enumeration value vector according to the metadata information of the target table and the target field, and then retrieve the knowledge fragment of the historical field with the greatest similarity from the government knowledge graph; the knowledge fragment includes: metadata information of the historical field, data quality inspection rules, data standards and data quality inspection templates associated with the historical field;
[0009] When the knowledge fragment is not empty, a dynamic prompt word is constructed based on the metadata information of the target table and target field and the knowledge fragment, and is passed into the large language model to generate data quality check rules for the target field.
[0010] Based on the further improvement of the above method, the field semantic vector is obtained by using the embedding model to obtain the embedding vector of the target field annotation; the joint semantic vector is obtained by using the embedding model to obtain the embedding vector of the target table annotation and the target field annotation; the field structure vector is obtained by using the embedding model to obtain the embedding vector of multiple metadata of the target field; the enumeration value vector is obtained by using the embedding model to obtain the embedding vector of the enumeration value list if the target field is an enumeration type.
[0011] Based on the further improvement of the above method, the metadata information of the target table and each target field therein includes: target table name, target table annotation, each target field name, each target field annotation, basic attributes and constraints of each target field; basic attributes include: data type, length, precision and enumeration value list; constraints include: whether it is a primary key, whether it is a foreign key, whether it is allowed to be empty, unique constraint and value range constraint.
[0012] Based on the further improvement of the above method, each historical field in the government knowledge graph has multiple indexes, where the first index is a weighted fusion vector calculated based on the joint semantic vector, field structure vector and / or enumeration value vector of the historical field, and their respective weights, the second index is the joint semantic vector of the historical field, and the third index is the field semantic vector of the historical field.
[0013] Based on the further improvement of the above method, the knowledge fragments of the historical fields with the greatest similarity are retrieved from the government knowledge graph, including:
[0014] Calculate a weighted fusion vector of the target field according to the joint semantic vector, the field structure vector and / or the enumeration value vector of the target field and their respective weights;
[0015] The first similarity, second similarity and third similarity between the weighted fusion vector, joint semantic vector and field semantic vector of the target field and the first index, second index and third index of each historical field in the government knowledge graph are calculated in sequence. As long as the maximum value of the calculated similarity is greater than the similarity threshold, the corresponding knowledge fragment is obtained and the retrieval is exited.
[0016] Based on the further improvement of the above method, the weight of the field structure vector and the initial weight of the enumeration value vector are dynamically adjusted according to the metadata information of the field. The formula is as follows:
[0017]
[0018] Among them, W s and Respectively represent the adjusted weight and initial weight of the field structure vector, W e and They represent the adjusted weight and initial weight of the enumeration value vector respectively, k represents the enumeration coefficient; r represents the weight growth coefficient, and m represents the number of non-zero values in the field metadata information.
[0019] Based on the further improvement of the above method, if the maximum value of the first similarity is greater than the similarity threshold, the corresponding knowledge fragment obtained includes: the metadata information of the historical field corresponding to the first similarity maximum value, and its associated historical data quality verification rules; if the maximum value of the second similarity is greater than the similarity threshold, the corresponding knowledge fragment obtained includes: the metadata information of the historical field corresponding to the second similarity maximum value, and the field attribute consistency verification template in the data quality verification template; if the maximum value of the third similarity is greater than the similarity threshold, the corresponding knowledge fragment obtained includes: the data standard followed by the historical field corresponding to the third similarity maximum value, and the data quality verification template related to the data standard.
[0020] Based on the further improvement of the above method, dynamic prompt words are constructed according to the metadata information of the target table and target field and the knowledge fragments, including:
[0021] Set the role as data quality check rule generation expert;
[0022] If there are historical data quality check rules in the knowledge fragment, the task goal set for the role is: refer to the historical fields and historical data quality check rules to generate data quality check rules for the current fields in the table to be checked; if there are data standards in the knowledge fragment, the task goal set for the role is: refer to the data standards and their related data quality check templates to generate check rules for whether the current fields in the table to be checked are consistent with the data standards; otherwise, the task goal set for the role is: refer to the field attribute consistency check template to generate check rules for the consistency of the current fields in the table to be checked with the historical field attributes;
[0023] The to-be-checked table and current field in the prompt word are set according to the metadata information of the target table and the target field; and the reference information required for the task target in the prompt word is set according to the knowledge fragment.
[0024] Based on the further improvement of the above method, the large language model is obtained by supervised fine-tuning training on the basis of the base large model; the loss function of the large language model includes SQL generation loss and SQL syntax loss; the SQL generation loss is calculated by using the cross entropy loss function, and the SQL syntax loss is calculated by performing syntax error detection on the data quality check rules output by the large language model, according to the number and weight of each type of syntax errors in the syntax detection results.
[0025] On the other hand, an embodiment of the present invention provides a system for generating government data quality verification rules based on a large language model, including:
[0026] The metadata extraction module is used to obtain the metadata information of the target table to be checked and each target field therein;
[0027] The knowledge fragment retrieval module is used to retrieve each target field in turn, obtain the field semantic vector, joint semantic vector, field structure vector and / or enumeration value vector according to the metadata information of the target table and the target field, and then retrieve the knowledge fragment of the historical field with the greatest similarity from the government knowledge graph; the knowledge fragment includes: metadata information of the historical field, data quality inspection rules, data standards and data quality inspection templates associated with the historical field;
[0028] The verification rule generation module is used to construct dynamic prompt words based on the metadata information of the target table and target field and the knowledge fragment when the knowledge fragment is not empty, and pass them into the large language model to generate data quality verification rules for the target field.
[0029] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects:
[0030] 1. First, based on metadata information, feature vectors of fields are extracted from different dimensions and levels. Then, relevant knowledge fragments are adaptively retrieved in combination with the government knowledge graph, realizing automatic adjustment of retrieval strategies and flexible selection of knowledge fragments. Finally, a large language model is used to quickly generate data quality verification rules, forming a complete and automated data quality assurance system that can comprehensively and deeply detect and improve data quality, effectively solve problems such as data consistency, accuracy and completeness, and ensure data reliability.
[0031] 2. Dynamic adjustment of the weights of different feature vectors is achieved based on the metadata information of the field, and the semantic, structural and enumeration value features of the field are weighted and integrated, which improves the ability to deeply understand and accurately analyze the field, ensures that the generated statements are closely matched with the actual structure of the database, and improves the accuracy of data quality rules.
[0032] 3. Based on metadata information and adaptively retrieved knowledge fragments, dynamic and efficient prompt words are constructed to improve the generation ability of large language models. The grammatical detection results of data quality verification rules are introduced into the loss function to guide the model to continuously optimize the recognition and correction capabilities of grammatical errors during the learning process, thereby improving the robustness and adaptability of the model when facing various complex grammatical errors.
[0033] In the present invention, the above-mentioned technical solutions can also be combined with each other to achieve more preferred combination solutions. Other features and advantages of the present invention will be described in the subsequent description, and some advantages can become obvious from the description, or can be understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained through the contents particularly pointed out in the description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] The accompanying drawings are only used for the purpose of illustrating specific embodiments and are not to be considered as limiting the present invention. In the entire drawings, the same reference symbols represent the same components;
[0035] Figure 1 This is a flow chart of a method for generating government data quality verification rules based on a large language model in Example 1 of the present invention;
[0036] Figure 2 This is a schematic diagram of the structure of the government data quality verification rule generation system based on the large language model in Example 2 of the present invention. DETAILED DESCRIPTION
[0037] The preferred embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings, wherein the accompanying drawings constitute a part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not used to limit the scope of the present invention.
[0038] Example 1
[0039] A specific embodiment of the present invention discloses a method for generating government data quality inspection rules based on a large language model, such as Figure 1 As shown, the following steps are included:
[0040] S1. Obtain the metadata information of the target table to be checked and each target field therein.
[0041] This embodiment provides a target table selection operation through a visual interface, and obtains metadata information of the target table and each target field therein according to the selected target table.
[0042] This embodiment does not limit the method for obtaining metadata information, and the metadata information can be obtained by querying the database system tables or views. For example, for a SQL Server database, query the system tables such as sys.tables, sys.columns, and sys.foreign_keys; for a MySQL database, query the views such as information_schema.tables, information_schema.columns, and information_schema.table_constraints; various metadata information can also be obtained through the DatabaseMetaData interface of JDBC; if a metadata management platform is built, it can also be obtained directly from the metadata management platform.
[0043] Furthermore, the metadata information of the target table and each target field therein obtained includes: target table name, target table annotation, target field name, target field annotation, basic attributes and constraints of each target field; basic attributes include: data type, length, precision and enumeration value list; constraints include: whether it is a primary key, whether it is a foreign key, whether it is allowed to be empty, unique constraint and value range constraint. Among them, the target table annotation indicates the Chinese name of the target table, and each target field annotation indicates the Chinese name of each target field; when the target field is a string type, the length indicates the maximum number of characters that the field can store, when the target field is a numeric type, the precision indicates the number of digits after the decimal point, and when the target field is an enumeration type, the obtained data type includes an enumeration value list. For example, gender is an enumeration type field, and its corresponding field type is: ENUM ('male', 'female', 'unknown'), where "ENUM" indicates an enumeration type, and the information in brackets is an enumeration value list.
[0044] S2. Take out each target field in turn, obtain the field semantic vector, joint semantic vector, field structure vector and / or enumeration value vector according to the metadata information of the target table and the target field, and then retrieve the knowledge fragment of the historical field with the greatest similarity from the government knowledge graph; the knowledge fragment includes: metadata information of the historical field, data quality verification rules, data standards and data quality verification templates associated with the historical field.
[0045] It should be noted that since there are multiple target fields in the target table, this step is based on the government knowledge graph that has been constructed, and obtains the knowledge fragments that can be referenced for each target field, thereby making full use of historical experience and enriching the context information of the target field, so that the large language model can more comprehensively understand the business meaning and usage scenarios of the field, and generate data quality verification rules that are more in line with actual needs. The data quality verification rules in this embodiment are SQL statements (Structured Query Language) of the database, which are used to verify the standardization, accuracy and consistency of the database table structure and database table records.
[0046] It should be noted that the data in the government knowledge graph comes from: metadata of database tables and fields of various business systems, historical data quality inspection rules and instructions, data standard files, and the constructed entities include: business systems, historical tables, historical fields, and data standards; the relationships between entities include: inclusion, association, and compliance. For example, the CRM system includes an order table and a customer table. The order table is associated with the customer table. The customer table includes a customer name field, a customer ID field, and a customer age field. The customer ID field complies with the ID card data standard.
[0047] Furthermore, different types of entities include multiple attributes, among which the attributes of the business system include but are not limited to: business system name, business system description and business system creation time; the attributes of the history table include: table comments and table creation time; the attributes of the history field include but are not limited to: data type, length, precision, enumeration value list, whether it is a primary key, whether it is a foreign key, whether it is allowed to be empty, whether it is unique, and whether there is a value range constraint; the attributes of the data standard include but are not limited to: standard data type, standard data format, standard value range, standard basis, data definer, data manager and data user.
[0048] In order to obtain accurate knowledge fragments, the entities corresponding to each historical field in the government knowledge graph are set as indexes based on multiple semantic vectors to capture the characteristics of multiple dimensions of historical fields.
[0049] It should be noted that the first index is a weighted fusion vector calculated based on the joint semantic vector, field structure vector and / or enumeration value vector of the historical field, the second index is the joint semantic vector of the historical field, and the third index is the field semantic vector of the historical field.
[0050] Specifically, the field semantic vector is obtained by using the embedding model to obtain the embedding vector of the field annotation, which is only used to represent the business semantic features of the field itself; the joint semantic vector is obtained by using the embedding model to obtain the embedding vector of the target table annotation and the target field annotation, which is used to represent the business semantic features of the field in a specific table.
[0051] The field structure vector is obtained by using the embedding model to obtain the embedding vector of multiple metadata of the target field, which is used to represent the structural semantic characteristics of the field; since the primary key in the data table is generally realized by the database mechanism to achieve the uniqueness of the primary key, it usually does not have business meaning, and the foreign key is used to establish the association relationship between tables. Generally, the data quality of the values of the primary key and the foreign key is not checked. Therefore, this embodiment generates the field structure vector based on the metadata such as the data type, length, precision, whether it is allowed to be empty, unique constraint and value range constraint of the target field. Among them, the data type is standardized, that is, the same type is represented by the same English characters, for example, the string type is Varchar; the length and precision are represented by the actual numerical value, if it is a string type, the precision is represented by 0; whether it is allowed to be empty is represented by 0 and 1 respectively to indicate that it is allowed to be empty and not allowed to be empty, the unique constraint is represented by 1 and 0 respectively to indicate uniqueness and non-uniqueness, if there is a value range constraint, it is represented by the upper and lower limits of the range, otherwise it is represented by 0 to indicate that there is no value range constraint. The metadata after the above processing is spliced and passed into the embedding model to obtain the field structure vector.
[0052] The enumeration value vector is obtained by using the embedding model to obtain the embedding vector of the enumeration value information when the target field is of enumeration type. That is to say, only fields of enumeration type have enumeration value vectors, while fields of other types do not.
[0053] It should be noted that this embodiment does not limit the type of embedding model, including but not limited to models trained based on word2vec, Bert, etc.
[0054] Furthermore, a weighted fusion vector calculated based on the joint semantic vector, the field structure vector and / or the enumeration value vector, and their respective weights is used as the first index to represent comprehensive data features of the historical field.
[0055] It should be noted that this example is based on the initial weights of the joint semantic vector, field structure vector and enumeration value vector. After dynamically adjusting the initial weights of the field structure vector and the enumeration value vector according to the metadata information of the field, the three weights are normalized to obtain the final weights of each of the three vectors.
[0056] Specifically, when the field is identified as an enumeration type according to the metadata information of the field, it is not necessary to consider the length of the field, whether it is allowed to be empty, the unique constraint and the value range constraint. Therefore, the weight of the structural semantic vector is smaller than the weight of the enumeration value vector. At this time, based on the initial weight, this embodiment increases the enumeration value vector by setting the enumeration coefficient and reduces the weight of the field structure vector. The formula is as follows:
[0057]
[0058] Among them, Ws and Respectively represent the adjusted weight and initial weight of the field structure vector, W e and They respectively represent the adjusted weight and initial weight of the enumeration value vector, and k represents the preset enumeration coefficient.
[0059] When the field is identified as an enumeration type based on the field metadata information, the weight W of the enumeration value vector e Set to 0, dynamically adjust the weight of the field structure vector according to the number of non-zero metadata in the metadata information of the calculated field structure vector. If the metadata information is not zero, data quality check rules may be generated for it, which further highlights the structural characteristics of the field. The formula is as follows:
[0060]
[0061] Among them, r represents the preset weight growth coefficient, and m represents the number of non-zero values in the field metadata information.
[0062] For example, the data type of field A is Varchar, the length is 50, the precision is 0, the "Is it allowed to be empty" is 0, that is, it is allowed to be empty, the "Unique constraint" is 0, that is, it is not unique, the "Value range constraint" is 0, that is, there is no value range constraint, then the number of metadata that is not 0 is 2; the data type of field B is Decimal, the length is 10, the precision is 2, the "Is it allowed to be empty" is 1, that is, it is not allowed to be empty, the "Unique constraint" is 0, that is, it is not unique, the "Value range constraint" is 0, that is, there is no value range constraint, then the number of metadata that is not 0 is 4.
[0063] For the target field currently taken out, the field semantic vector, joint semantic vector, field structure vector and / or enumeration value vector are also calculated according to the method of calculating multiple indexes of historical fields in the government knowledge spectrum. And based on the joint semantic vector, field structure vector and / or enumeration value vector of the target field, the weighted fusion vector of the target field is calculated according to the above-mentioned method of calculating weights.
[0064] In this embodiment, the weighted fusion vector, the joint semantic vector and the field semantic vector respectively represent feature representations of different depths. The historical fields with the greatest similarity in different situations are obtained from the government knowledge graph in sequence, and knowledge fragments of different dimensions are selected from the data associated with the historical fields. This method not only gives priority to historical fields that are highly matched with the target fields, but also selects the most relevant and applicable knowledge fragments in the matching historical fields. At the same time, it realizes the adaptive adjustment of retrieval strategies and knowledge fragments, thereby improving the accuracy and relevance of retrieval results.
[0065] Specifically, the knowledge fragments of the historical fields with the greatest similarity are retrieved from the government knowledge graph, including:
[0066] The first similarity, second similarity and third similarity between the weighted fusion vector, joint semantic vector and field semantic vector of the target field and the first index, second index and third index of each historical field in the government knowledge graph are calculated in sequence. As long as the maximum value of the calculated similarity is greater than the similarity threshold, the corresponding knowledge fragment is obtained and the retrieval is exited.
[0067] That is to say, first calculate the first similarity between the weighted fusion vector of the target field and the first index of each historical field in the government knowledge graph. If the maximum value of the first similarity is greater than the similarity threshold, obtain the first knowledge fragment of the historical field corresponding to the maximum value of the first similarity, and no longer search; otherwise, continue to calculate the second similarity between the joint semantic vector of the target field and the second index of each historical field in the government knowledge graph. If the maximum value of the second similarity is greater than the similarity threshold, obtain the second knowledge fragment of the historical field corresponding to the second maximum value of the similarity, and no longer search; otherwise, calculate the third similarity between the field semantic vector of the target field and the third index of each historical field in the government knowledge graph. If the maximum value of the third similarity is greater than the similarity threshold, obtain the third knowledge fragment of the historical field corresponding to the third maximum value of the similarity.
[0068] It should be noted that the similarity in this embodiment is obtained by calculating the cosine similarity.
[0069] Furthermore, if the first maximum similarity value is greater than the similarity threshold, it means that each dimension of the target field is very similar to the historical field corresponding to the first maximum similarity value, then the data quality verification rules of the historical field can be directly referred to. Therefore, the first knowledge fragment obtained includes: the metadata information of the historical field, and its associated historical data quality verification rules.
[0070] If the maximum value of the second similarity is greater than the similarity threshold, it means that the table name and field name of the target field are very similar to the historical field corresponding to the second maximum value of similarity, then it is necessary to further check whether the target field is consistent with other field-level metadata of the historical field to ensure that the fields with the same business meaning are structurally unified with the existing historical fields. Therefore, the second knowledge fragment obtained includes: the metadata information of the historical field, and the field attribute consistency check template in the data quality check template; the field attribute consistency check template includes: data type consistency check template, length consistency check template, value range consistency check template and enumeration value consistency check template. Preferably, the relevant field attribute consistency check template is selected according to the metadata information of the target field.
[0071] If the maximum value of the third similarity is greater than the similarity threshold, it means that only the field name of the target field is similar to the historical field corresponding to the maximum value of the third similarity, then consider whether the historical field complies with the data standard. If so, the third knowledge fragment obtained includes: the data standard followed by the historical field, and the data quality verification template related to the data standard. Exemplarily, the field annotation of the historical field is ID card, which complies with the ID card data standard, and the data quality verification template related to the data standard is the ID card normative verification template.
[0072] Furthermore, if the maximum value of the third similarity is greater than the similarity threshold, but the historical field corresponding to the maximum value of the third similarity does not comply with any data standard, or the maximum value of the third similarity is less than or equal to the similarity threshold, in these two cases, the knowledge fragments obtained from the government knowledge graph are empty. Then, according to the metadata information of the target field, the data quality check rules are generated using the corresponding data quality check template, and step S3 does not need to be executed.
[0073] Exemplarily, if the "Is it allowed to be empty" in the metadata information of the target field is not allowed to be empty, the table name and field name of the target field are replaced by the variables in the non-empty check template to generate a non-empty check rule; if the "unique constraint" is required to be unique, the table name and field name of the target field are replaced by the variables in the uniqueness check template to generate a uniqueness check rule.
[0074] S3. When the knowledge fragment is not empty, a dynamic prompt word is constructed according to the metadata information of the target table and the target field and the knowledge fragment, and the dynamic prompt word is passed into the large language model to generate the data quality check rules for the target field.
[0075] It should be noted that efficient prompt words are conducive to improving the generation ability of large language models. By setting roles through prompt words, we can focus on knowledge and skills in specific fields, thereby generating more professional and accurate content. By clarifying the task objectives of the role, the large language model can perform tasks quickly and accurately, thereby improving interaction efficiency.
[0076] This embodiment constructs dynamic prompt words according to metadata information of the target table and target field and knowledge fragments, including:
[0077] ① Set the role to Data Quality Check Rule Generation Expert; for example, describe in the prompt: You are now a Data Quality Check SQL Generation Expert.
[0078] ② If there are historical data quality check rules in the knowledge fragment, the task goal set for the role is: refer to the historical fields and historical data quality check rules to generate data quality check rules for the current fields in the table to be checked; if there are data standards in the knowledge fragment, the task goal set for the role is: refer to the data standards and their related data quality check templates to generate check rules for whether the current fields in the table to be checked are consistent with the data standards; otherwise, the task goal set for the role is: refer to the field attribute consistency check template to generate check rules for the consistency of the current fields in the table to be checked with the historical field attributes;
[0079] ③ Set the to-be-checked table and current field in the prompt word according to the metadata information of the target table and target field; set the reference information required for the task target in the prompt word according to the knowledge fragment.
[0080] Specifically, after the description of the role and its task objectives, the reference information corresponding to the to-be-checked table, current field, and task objectives is concatenated with a line break: the target table name is taken out and placed after "to-be-checked table:", the metadata information of the target field is taken out and sequentially concatenated with the preset first connector after "current field:"; if the information referenced by the set task objective is the historical field and the historical data quality verification rule, the metadata information of the historical field is taken out from the knowledge fragment, sequentially concatenated with the preset connector after "historical field:", the historical data quality verification rule is taken out and sequentially concatenated with "historical data quality verification rule:", and finally the complete prompt word information is obtained. Exemplarily, the connector is set to "|".
[0081] Preferably, SQL generation rules are set in the prompt words to constrain the reasoning direction of the large language model.
[0082] The large language model of this embodiment is supervised fine-tuned based on the base large model. In order to ensure that the generated data quality check SQL can be correctly executed, the data quality check rules output by the large language model are checked for grammatical errors, and the grammar check results are introduced into the loss function. Therefore, the loss function of the large language model of this embodiment includes: SQL generation loss L SQL and SQL syntax loss L Syntax There are two parts, among which, the SQL generation loss is obtained by calculating the difference between the generated data quality check rule SQL and the actual data quality check rule SQL, and the SQL syntax loss is obtained by calculating the syntax error score. The formula is as follows:
[0083] L total =L SQL +λ·L Syntax ,
[0084] Among them, λ represents the influence coefficient, which is a hyperparameter used to control the influence of grammatical error loss on the total loss. The larger λ is, the more the model will pay attention to the correction of grammatical errors.
[0085] Furthermore, the SQL generation loss L SQL The cross entropy loss function is used, and the formula is as follows:
[0086]
[0087] Where N is the total number of all parts in SQL, y i represents the true label of the i-th part in SQL. If the part is correct, then y i is 1, otherwise it is 0; Represents the probability of the i-th part predicted by the model.
[0088] SQL syntax loss L Syntax It is calculated based on the number and weight of each type of grammatical error in the grammar detection results. The formula is as follows:
[0089]
[0090] Where M represents the total number of grammatical error types, W j Indicates the weight of the j-th grammatical error, Num j In this embodiment, the output data quality check rule SQL is constructed into an abstract syntax tree, and then the abstract syntax data is analyzed to obtain the syntax detection result. The syntax error types include: missing keywords, mismatched brackets, incorrect clause order, incorrect table name, and incorrect column name.
[0091] During the training process, the stochastic gradient descent method is used to update the model parameters to minimize the loss function value. After multiple rounds of training, the performance of the model in the government data quality verification task is gradually improved, and a trained large language model is obtained.
[0092] During implementation, the target table to be subject to data quality check is selected through step S1, the knowledge fragments of each target field are obtained one by one through step S2, and the prompt words in text form generated by dynamically splicing the role, task goal, target table, target field, and reference information are passed into the trained large language model through step S3, and the data quality check rules for the target field are output.
[0093] It should be noted that this embodiment does not limit the execution method of step S2 and step S3. In step S2, the knowledge fragment of the current target field can be extracted, and then step S3 can be executed to obtain the data quality verification rules of the current target field, and then the next target field can be extracted after returning to step S2. Alternatively, in step S2, the knowledge fragment of each target field can be obtained in turn, and then step S3 can be executed in turn to obtain the data quality verification rules of each target field.
[0094] Compared with the prior art, the method for generating government data quality check rules based on a large language model provided in this embodiment first extracts feature vectors of fields from different dimensions and levels based on metadata information, then adaptively retrieves related knowledge fragments in combination with the government knowledge graph, realizes automatic adjustment of retrieval strategies and flexible selection of knowledge fragments, and finally uses the large language model to quickly generate data quality check rules, forming a complete and automated data quality assurance system, which can comprehensively and deeply detect and improve data quality, effectively solve problems such as data consistency, accuracy and completeness, and ensure data reliability. According to the metadata information of the field, the dynamic adjustment of the weights of different feature vectors is realized, and the semantic, structural and enumerated value features of the field are weighted and fused, which improves the ability of deep understanding and precise analysis of the field, ensures that the generated statements are closely matched with the actual structure of the database, and improves the accuracy of the data quality rules. According to the metadata information and the knowledge fragments of adaptive retrieval, dynamic and efficient prompt words are constructed to improve the generation ability of the large language model; the grammar detection results of the data quality check rules are introduced into the loss function, guiding the model to continuously optimize the recognition and correction capabilities of grammar errors during the learning process, thereby improving the robustness and adaptability of the model in the face of various complex grammar errors.
[0095] Example 2
[0096] Another embodiment of the present invention discloses a system for generating government data quality verification rules based on a large language model, thereby realizing a method for generating government data quality verification rules based on a large language model in Embodiment 1. The specific implementation of each module refers to the corresponding description in Embodiment 1. Figure 2 As shown, the system includes:
[0097] The metadata extraction module 101 is used to obtain metadata information of the target table to be checked and each target field therein;
[0098] The knowledge fragment retrieval module 102 is used to retrieve each target field in turn, obtain the field semantic vector, joint semantic vector, field structure vector and / or enumeration value vector according to the metadata information of the target table and the target field, and then retrieve the knowledge fragment of the historical field with the greatest similarity from the government knowledge graph; the knowledge fragment includes: metadata information of the historical field, data quality inspection rules, data standards and data quality inspection templates associated with the historical field;
[0099] The verification rule generation module 103 is used to construct dynamic prompt words according to the metadata information of the target table and the target field and the knowledge fragment when the knowledge fragment is not empty, and pass them into the large language model to generate data quality verification rules for the target field.
[0100] Since the system embodiment and the aforementioned method for generating government data quality verification rules based on a large language model are related and can be mutually referenced, this is a repeated description and will not be repeated here. Since the system embodiment and the aforementioned method embodiment have the same principle, the system embodiment also has the corresponding technical effects of the aforementioned method embodiment.
[0101] Those skilled in the art will appreciate that all or part of the processes of the above-mentioned embodiments can be implemented by instructing related hardware through a computer program, and the program can be stored in a computer-readable storage medium, wherein the computer-readable storage medium is a disk, an optical disk, a read-only storage memory, or a random access memory, etc.
[0102] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by any technician familiar with the technical field within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. A method for generating government data quality verification rules based on a large language model, characterized in that: The following steps are involved: Obtain the metadata information of the target table to be checked and each target field therein; Take out each target field in turn, obtain the field semantic vector, joint semantic vector, field structure vector and / or enumeration value vector according to the metadata information of the target table and the target field, and then retrieve the knowledge fragment of the historical field with the greatest similarity from the government knowledge graph; the knowledge fragment includes: metadata information of the historical field, data quality inspection rules, data standards and data quality inspection templates associated with the historical field; When the knowledge fragment is not empty, a dynamic prompt word is constructed based on the metadata information of the target table and target field and the knowledge fragment, and is passed into the large language model to generate data quality check rules for the target field.
2. The method for generating government data quality check rules based on a large language model according to claim 1 is characterized in that: The field semantic vector is obtained by using the embedding model to obtain the embedding vector of the target field annotation; the joint semantic vector is obtained by using the embedding model to obtain the embedding vector of the target table annotation and the target field annotation; the field structure vector is obtained by using the embedding model to obtain the embedding vector of multiple metadata of the target field; the enumeration value vector is obtained by using the embedding model to obtain the embedding vector of the enumeration value list if the target field is an enumeration type.
3. The method for generating government data quality verification rules based on a large language model according to claim 1 or 2, characterized in that: The metadata information of the target table and each target field therein includes: target table name, target table annotation, target field names, target field annotations, basic attributes and constraints of each target field; the basic attributes include: data type, length, precision and enumeration value list; the constraints include: whether it is a primary key, whether it is a foreign key, whether it is allowed to be empty, unique constraint and value range constraint.
4. The method for generating government data quality check rules based on a large language model according to claim 1 is characterized in that: Each historical field in the government knowledge graph has multiple indexes, wherein the first index is a weighted fusion vector calculated based on the joint semantic vector, field structure vector and / or enumeration value vector of the historical field, and their respective weights, the second index is the joint semantic vector of the historical field, and the third index is the field semantic vector of the historical field.
5. The method for generating government data quality check rules based on a large language model according to claim 4 is characterized in that: The step of retrieving the knowledge fragment of the historical field with the greatest similarity from the government affairs knowledge graph includes: Calculate a weighted fusion vector of the target field according to the joint semantic vector, the field structure vector and / or the enumeration value vector of the target field and their respective weights; The first similarity, second similarity and third similarity between the weighted fusion vector, joint semantic vector and field semantic vector of the target field and the first index, second index and third index of each historical field in the government knowledge graph are calculated in sequence. As long as the maximum value of the calculated similarity is greater than the similarity threshold, the corresponding knowledge fragment is obtained and the retrieval is exited.
6. The method for generating government data quality verification rules based on a large language model according to claim 4 or 5, characterized in that: The weight of the field structure vector and the initial weight of the enumeration value vector are dynamically adjusted according to the metadata information of the field. The formula is as follows: Among them, W s and Respectively represent the adjusted weight and initial weight of the field structure vector, W e and They represent the adjusted weight and initial weight of the enumeration value vector respectively, k represents the enumeration coefficient; r represents the weight growth coefficient, and m represents the number of non-zero values in the field metadata information.
7. The method for generating government data quality check rules based on a large language model according to claim 5 is characterized in that: If the maximum value of the first similarity is greater than the similarity threshold, the corresponding knowledge fragment obtained includes: the metadata information of the historical field corresponding to the first similarity maximum value, and its associated historical data quality verification rules; if the maximum value of the second similarity is greater than the similarity threshold, the corresponding knowledge fragment obtained includes: the metadata information of the historical field corresponding to the second similarity maximum value, and the field attribute consistency verification template in the data quality verification template; if the maximum value of the third similarity is greater than the similarity threshold, the corresponding knowledge fragment obtained includes: the data standard followed by the historical field corresponding to the third similarity maximum value, and the data quality verification template related to the data standard.
8. The method for generating government data quality check rules based on a large language model according to claim 7 is characterized in that: The step of constructing dynamic prompt words according to metadata information of the target table and target field and knowledge fragments includes: Set the role as data quality check rule generation expert; If there are historical data quality check rules in the knowledge fragment, the task goal set for the role is: refer to the historical fields and historical data quality check rules to generate data quality check rules for the current fields in the table to be checked; if there are data standards in the knowledge fragment, the task goal set for the role is: refer to the data standards and their related data quality check templates to generate check rules for whether the current fields in the table to be checked are consistent with the data standards; otherwise, the task goal set for the role is: refer to the field attribute consistency check template to generate check rules for the consistency of the current fields in the table to be checked with the historical field attributes; The to-be-checked table and current field in the prompt word are set according to the metadata information of the target table and the target field; and the reference information required for the task target in the prompt word is set according to the knowledge fragment.
9. The method for generating government data quality check rules based on a large language model according to claim 1 is characterized in that: The large language model is obtained by supervised fine-tuning training based on the base large model; the loss function of the large language model includes SQL generation loss and SQL syntax loss; the SQL generation loss is calculated using the cross entropy loss function, and the SQL syntax loss is calculated by performing syntax error detection on the data quality check rules output by the large language model, based on the number and weight of various types of syntax errors in the syntax detection results.
10. A government data quality check rule generation system based on a large language model, characterized in that: include: The metadata extraction module is used to obtain the metadata information of the target table to be checked and each target field therein; The knowledge fragment retrieval module is used to retrieve each target field in turn, obtain the field semantic vector, joint semantic vector, field structure vector and / or enumeration value vector according to the metadata information of the target table and the target field, and then retrieve the knowledge fragment of the historical field with the greatest similarity from the government knowledge graph; the knowledge fragment includes: metadata information of the historical field, data quality inspection rules, data standards and data quality inspection templates associated with the historical field; The verification rule generation module is used to construct dynamic prompt words based on the metadata information of the target table and target field and the knowledge fragment when the knowledge fragment is not empty, and pass them into the large language model to generate data quality verification rules for the target field.
Citation Information
Patent Citations
Method and system for constructing data quality monitoring rule of data warehouse
CN115292297A
Data quality inspection rule construction method, storage medium and system
CN115357572A
Database query method and device, electronic equipment and nonvolatile storage medium
CN119226315A
Prompt word construction method and device
CN119312802A
Cited By
Data quality inspection method and system based on dynamic rule and computer equipment
CN120670417A
Data confusion relation evaluation method and device, equipment and storage medium
CN120723746A
Government affair material auditing method and system based on historical case analysis
CN121279955A
A Method and System for Reviewing Government Documents Based on Historical Case Analysis
CN121279955B
Semiconductor factory fault tracing method and system based on LLM and retrieval enhancement
CN121501845A