Logic model management method and device, electronic equipment, storage medium and product

CN122759331APending Publication Date: 2026-09-15BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610813144.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-05
Publication Date
2026-09-15

AI Technical Summary

Technical Problem

[0003]然而,随着业务人员频繁在物理表中新增字段、修改类型或调整命名,物理表的结构会动态变化,而逻辑模型往往仍停留在上一次人工维护的状态,导致两者逐渐偏离

Benefits of technology

[0008] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are configured to cause the computer to perform the methods described in embodiments of this disclosure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122759331A_ABST
    Figure CN122759331A_ABST
Patent Text Reader

Abstract

The present disclosure provides a logic model management method and device, electronic equipment, storage medium and product, relates to the technical field of computers and large models, in particular to the technical field of data management. The specific implementation scheme is: in response to detecting a change event for changing a physical table, generating a reverse engineering request based on information of the change event; obtaining a revised draft of a logic model of the changed physical table based on the reverse engineering request, the revised draft being obtained by reverse engineering processing based on the reverse engineering request; calling a large language model, generating an audit report for the revised draft based on at least one of a preset compliance rule and a data standard; and in response to a confirmation operation on the revised draft based on the audit report, publishing the revised draft. The present scheme can automatically obtain the change event of the changed physical table to realize the process of automatically publishing the revised draft, save human resources, and improve the consistency of the physical table and its logic model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to computer technology, large model technology, and especially to data management technology. Specifically, this disclosure relates to a logical model management method, apparatus, electronic device, storage medium, and product. Background Technology

[0002] In the field of data management technology, physical tables reside in heterogeneous data stores, such as traditional relational databases (e.g., MySQL, Doris) or data lakes (e.g., Iceberg). Data modeling platforms are used to maintain the logical models of these physical tables, defining their field names, data types, constraints, and column categories. To ensure the accuracy of data management, the physical tables and logical models need to remain consistent in real time.

[0003] However, as business personnel frequently add fields, modify types, or adjust names in physical tables, the structure of physical tables changes dynamically, while the logical model often remains in the state of the last manual maintenance, causing the two to gradually deviate. Summary of the Invention

[0004] This disclosure provides a method, apparatus, electronic device, storage medium, and product for managing logical models, enabling an automated process for publishing revision drafts to improve the consistency between physical tables and logical models.

[0005] According to one aspect of this disclosure, a logical model management method is provided, comprising: In response to the detection of a change event targeting a physical table, a reverse engineering request is generated based on information from the change event; A revised draft of the logical model of the modified physical table is obtained based on the reverse engineering request, and the revised draft is obtained by reverse engineering based on the reverse engineering request. The large language model is invoked to generate an audit report for the revised draft based on at least one of the preset compliance rules and data standards. In response to the confirmation of the revised draft based on the audit report, the revised draft is published.

[0006] According to another aspect of this disclosure, a logic model management apparatus is provided, comprising: An event response unit is configured to generate a reverse engineering request based on information from the detected change event for a modified physical table. The draft acquisition unit is configured to acquire a revised draft of the logical model of the changed physical table based on the reverse engineering request, wherein the revised draft is obtained by reverse engineering based on the reverse engineering request; The report generation unit is configured to invoke a large language model and generate an audit report for the revised draft based on at least one of preset compliance rules and data standards. The draft publishing unit is configured to publish the revised draft in response to the confirmation operation of the revised draft based on the review report.

[0007] According to another aspect of this disclosure, an electronic device is provided, comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the methods described in the embodiments of this disclosure.

[0008] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are configured to cause the computer to perform the methods described in embodiments of this disclosure.

[0009] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the methods described in the embodiments of this disclosure.

[0010] This disclosure can respond to the detection of a change event affecting a physical table, generate a reverse engineering request based on the change event information, obtain a revised draft of the logical model of the changed physical table based on the reverse engineering request, invoke a large language model, generate an audit report based on at least one of preset compliance rules and data standards, and, in response to the confirmation operation of the revised draft based on the audit report, publish the revised draft. By automatically converting the perceived change event information of the physical table into a revised draft and publishing it, the process from physical table change to logical model generation is automated, thereby avoiding information attenuation and omissions caused by manual transmission, saving human resources, effectively preventing the structural differences between the logical model and the physical table from continuously expanding over time, and improving the consistency between the logical model and the physical table.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0012] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure.

[0013] Figure 1This is the system architecture diagram to which this disclosure applies.

[0014] Figure 2 This is a flowchart of the logical model method provided in this disclosure.

[0015] Figure 3 This is a schematic diagram of a change column information provided in this disclosure.

[0016] Figure 4 This is a schematic diagram of a second prompt instruction provided in this disclosure.

[0017] Figure 5 This is a schematic diagram of a data standard matching result provided in this disclosure.

[0018] Figure 6 This is a schematic diagram of a fourth prompt instruction provided in this disclosure.

[0019] Figure 7 This is a schematic diagram of the output of a third-largest language model provided in this publication.

[0020] Figure 8 This is a schematic diagram of the overall process for publishing the revised draft.

[0021] Figure 9 This is a schematic block diagram of the logic model management device provided in this disclosure.

[0022] Figure 10 This is a block diagram of an electronic device used to implement the logic model management method of the embodiments of this disclosure. Detailed Implementation

[0023] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0024] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The singular forms “a,” “the,” and “the” as used in the embodiments of this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0025] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0026] Depending on the context, the word "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."

[0027] In the field of data management technology, physical tables reside in heterogeneous data stores, such as traditional relational databases (e.g., MySQL, Doris) or data lakes (e.g., Iceberg). Data modeling platforms are used to maintain the logical models of physical tables, which define the field names, data types, constraints, and column categories. To ensure the accuracy of data management, physical tables and logical models need to remain consistent in real time; that is, the logical model should accurately reflect the latest metadata of the physical tables. However, as business users frequently add fields, modify types, or adjust names in physical tables, the structure of the physical tables dynamically changes, while the logical model often remains in the state of the last manual maintenance, causing the two to gradually deviate.

[0028] To maintain consistency between the logical model and physical tables in a data modeling platform, the commonly used technical solution is to rely on manual reverse engineering. Specifically, developers periodically perform reverse engineering from the physical tables to the logical model within the data modeling platform to obtain the latest logical model of the physical tables. For example, developers first confirm changes to the physical table structure using the database client, then locate the corresponding physical table in the data modeling platform, and finally click the reverse engineering button. The data modeling platform then generates a new logical model based on the structure snapshot of the selected physical table.

[0029] However, when this solution is applied to daily operations and maintenance scenarios involving multiple data sources and frequent changes at the enterprise level, its performance is not ideal. Specifically, to achieve the accuracy of logical model updates, this solution relies entirely on manual intervention and judgment. However, the frequency and timeliness of manual intervention are low, resulting in long-term inconsistencies between the logical model and the physical table structure, leading to the negative consequence of chaotic data management.

[0030] In view of this, this disclosure provides a new approach. To facilitate understanding of this disclosure, the system architecture on which this disclosure is based will first be described. Figure 1 Exemplary system architectures that can be applied to embodiments of this disclosure are shown, such as Figure 1 As shown, the system architecture may include: a first type of data source, a second type of data source, a logical model management device, and a reverse engineering system, wherein the logical model management device and the reverse engineering system may be located in the data modeling platform.

[0031] It should be understood that Figure 1 The number of the first type of data source, the second type of data source, the logical model management device, and the reverse engineering system shown is merely illustrative. Depending on implementation needs, any number of the first type of data source, the second type of data source, the logical model management device, and the reverse engineering system can be included.

[0032] The logical model management device and the reverse engineering system can be partially coupled or completely decoupled. As a specific embodiment, in response to detecting a change event for a modified physical table from a first data source and a second data source, the logical model management device generates a reverse engineering request based on the information of the change event and sends the reverse engineering request to the reverse engineering system. Then, the reverse engineering system generates a revised draft of the logical model of the modified physical table based on the reverse engineering request, and invokes a large language model to generate an audit report for the revised draft based on at least one of preset compliance rules and data standards. In response to the confirmation operation of the revised draft based on the audit report, the revised draft is published.

[0033] In another specific embodiment, the logical model management device, in response to detecting a change event for a modified physical table from a first data source and a second data source, generates a reverse engineering request based on the information of the change event and sends the reverse engineering request to the reverse engineering system. The reverse engineering system then generates a revised draft of the logical model of the modified physical table based on the reverse engineering request and returns the revised draft to the logical model management device. The logical model management device then invokes a large language model and, based on at least one of preset compliance rules and data standards, generates an audit report for the revised draft. In response to the confirmation operation of the revised draft based on the audit report, the revised draft is published.

[0034] Figure 2 A flowchart of a logic model management method provided in this disclosure embodiment, the method can be... Figure 1 The logical model management device in the system shown is executed. For example... Figure 2 As shown, the method may include the following steps: Step 201: In response to the detection of a change event for a modified physical table, a reverse engineering request is generated based on the information of the change event.

[0035] Step 202: Obtain a revised draft of the logical model of the changed physical table based on the reverse engineering request. The revised draft is obtained by reverse engineering based on the reverse engineering request.

[0036] Step 203: Invoke the large language model and generate an audit report for the revised draft based on at least one of the preset compliance rules and data standards.

[0037] Step 204: In response to the confirmation of the revised draft based on the review report, the revised draft is published.

[0038] As can be seen from the above process, this disclosure can, in response to the detection of a change event affecting a physical table, generate a reverse engineering request based on the information of the change event, obtain a revised draft of the logical model of the changed physical table based on the reverse engineering request, call a large language model, generate an audit report based on at least one of preset compliance rules and data standards, and, in response to the confirmation operation of the revised draft based on the audit report, publish the revised draft. By automatically converting the information of the perceived change event affecting the physical table into a revised draft and publishing it, the process from physical table change to logical model generation is automated, thereby avoiding information attenuation and omissions caused by manual transmission, saving human resources, effectively preventing the structural differences between the logical model and the physical table from continuously expanding over time, and improving the consistency between the logical model and the physical table.

[0039] The following describes in detail each step of the above process and the effects that can be further produced, with reference to the embodiments.

[0040] First, the above step 201, namely "in response to detecting a change event for a modified physical table, generating a reverse engineering request based on the information of the change event", will be described in detail with reference to the embodiments.

[0041] In embodiments of this disclosure, in response to detecting a change event targeting a physical table, a reverse engineering request is generated based on the information of the change event. That is, the logical model management device continuously monitors for possible change events on one or more heterogeneous data sources. For example, a change event is triggered when the structure of a user profile table or order details table changes. Based on the information of the change event, a reverse engineering request is generated, which is used to subsequently trigger the reverse engineering process.

[0042] As one possible approach, heterogeneous data sources can include a first type of data source, which can refer to a data lake (such as Iceberg) with unified metadata management capabilities. In this first type of data source, all modifications to the structure of the physical table (such as adding columns, renaming columns, etc.) are executed through a unified metadata operation interface (such as the updateTable method).

[0043] For the first type of data source, a detection logic can be embedded within its metadata operation interface. When an operation representing a structural change of a physical table is detected in the call to the metadata operation interface, a metadata change event is generated based on this operation. This metadata change event is the change event, and it carries the location information of the physical table where the structural change occurred. Based on the data source identifier corresponding to the metadata operation interface and the location information of the physical table, the interface parameters required for the reverse engineering request are generated. These request parameters are encapsulated into a request body (i.e., a reverse engineering request) that conforms to the interface specification of the reverse engineering system, thus achieving zero-latency proactive detection of structural changes of physical tables in the first type of data source.

[0044] As another feasible approach, heterogeneous data sources can include a second type of data source, which can refer to a traditional relational database (such as MySQL, PostgreSQL, etc.). This second type of data source lacks a unified metadata manipulation interface, and changes to its physical table structure are typically recorded in its own audit log files.

[0045] For the second type of data source, a scheduled task can periodically scan the audit log files of this data source and extract Data Definition Language (DDL) change statements. As a specific example, a pluggable log parser is first used to extract database operation statements from the scanned audit log files. These database operation statements refer to Structured Query Language (SQL) statements. Then, regular expressions are used to filter candidate DDL change statements containing preset DDL keywords. Next, preset non-structure change keywords are used to identify non-table structure change statements from the candidate DDL change statements. These preset non-structure change keywords can refer to at least one of "ALTER TABLE…ADD ROLLUP" (add a materialized view) or "ALTER TABLE…SET" (modify table attributes). Candidate DDL change statements excluding non-table structure change statements are then identified as DDL change statements.

[0046] Furthermore, a reverse engineering request is generated based on the DDL change statements. Specifically, based on the DDL change statements, the first location information of the changed physical table is determined. Based on the data source identifier corresponding to the audit log file and the first location information of the changed physical table, a reverse engineering request is generated. Although the reverse engineering request generated in this way has a delay of minutes, it can effectively cover the perception of changes to the second type of data source.

[0047] In the reverse engineering request generated based on the data source identifier corresponding to the audit log file and the first location information of the changed physical table, as a specific implementation, the statement type of the DDL change statement is determined. The statement type may include a first statement type and a second statement type, and the complexity of the first statement type is less than that of the second statement type.

[0048] In response to a DDL change statement being of type 1, a reverse engineering request is generated based on the data source identifier corresponding to the audit log file and the first location information of the changed physical table. In other words, for DDL change statements of type 1, relatively accurate first location information of the changed physical table can be extracted from the DDL change statement based on preset rules. The data source identifier corresponding to the audit log file and the first location information of the changed physical table are then encapsulated into a request body (i.e., a reverse engineering request) that conforms to the interface specifications of the reverse engineering system, without needing to call a large language model.

[0049] In response to a DDL change statement being of type 2, the large language model is used to generate the interface parameters required for the reverse engineering request based on the DDL change statement, the data source identifier corresponding to the audit log file, and the first location information of the changed physical table. Based on these interface parameters, the reverse engineering request is then generated. In other words, for DDL change statements of type 2, which may contain semantics difficult to parse using preset rules (such as multi-table renaming, mixed syntax, non-standard extensions, etc.), the first location information of the changed physical table extracted solely based on preset rules may not be accurate enough. In this case, the DDL change statement, the data source identifier corresponding to the audit log file, and the first location information of the changed physical table can be input into the large language model. This allows the large language model to extract the second location information of the changed physical table from the DDL change statement, and based on the second and first location information, generate the final location information of the changed physical table. Then, based on the final location information of the changed physical table and the data source identifier corresponding to the audit log file, the interface parameters required for the reverse engineering request are generated. These request parameters are then encapsulated into a request body (i.e., the reverse engineering request) that conforms to the interface specifications of the reverse engineering system.

[0050] By transforming unstructured DDL change statements extracted from audit log files into structured first location information, it is ensured that change events from the second type of data source can generate reverse engineering requests with a unified format, enabling subsequent reverse engineering processing to access change events from all sources without discrimination.

[0051] It should be noted that the location information of the changed physical table is used to uniquely identify the physical table that has undergone structural changes. In a specific embodiment, the location information includes at least the database name and the table name, and may also include a catalog (applicable to the three-segment naming of Doris in the second type of data source or the first type of data source, which can be empty), a schema (applicable to the two-segment naming of PostgreSQL or the first type of data source, which can be empty), and a prefix table name (i.e., the table name with a logical model type prefix, which can be empty).

[0052] For the first type of data source, the database name, table name, directory, and schema can be extracted from metadata change events (if not extracted, use default values ​​or set to an empty string). For the second type of data source, the database name, table name, directory, and schema can be extracted from DDL change statements (if not extracted, use default values ​​or set to an empty string).

[0053] Additionally, the purpose of the prefix table name (i.e., prefixTable) is to provide a table name that includes the logical model type for reverse engineering, making it easier for downstream businesses to identify which layer the changed physical table belongs to (such as the source layer, summary layer, application layer, etc.). The prefix table name usually cannot be directly extracted from DDL change statements or metadata change events. Therefore, when determining the prefix table name, it is first checked whether the table name already contains the prefix information of the preset logical model type. The preset logical model types and their meanings are shown in Table 1. The prefix information of the preset logical model type can refer to "ODS_", "DIM_", "FACT_", "SUM_", "APP_", etc. If the table name already contains the prefix information of the above logical model type, then the table name is directly used as the prefix table name. If the table name does not contain the prefix information of the above logical model type, then the prefix information of the logical model type corresponding to the table name is appended to the beginning of the table name, and the table name after appending the prefix information of the logical model type is used as the prefix table name.

[0054] Table 1 The logical model type corresponding to the table name can be determined in the following way: Based on the correspondence between the prefix information in the table name and the preset logical model type, the logical model type corresponding to the table name is determined. The correspondence between the prefix information of the table name and the logical model type can be shown in Table 2. For example, the prefix information "ods_" of the table name has a correspondence with "ODS_TABLE".

[0055] Table 2 Regardless of whether it is the first type of data source or the second type of data source, the interface parameters required for the final reverse engineering request should remain consistent. That is, in the above embodiment, the interface parameters may include data source identifier, directory, schema, database name, table name and prefix table name, or may include data source identifier, directory, schema, database name, table name, prefix table name and logical model type.

[0056] It should also be noted that a change event may include only one of the two types of events mentioned above, or it may include both types of events simultaneously. For example, in a medical data governance scenario, there is both a health record data lake managed through a unified metadata service (a first type of data source) and business databases from multiple hospitals (a second type of data source). This disclosure can embed detection logic within the metadata operation interface of the first type of data source while simultaneously obtaining the audit log file of the second data source and extracting DDL change statements from it. This ensures that regardless of the type of data source on which the change event occurs, a response can be made to the change event, providing a unified reverse engineering request for subsequent reverse engineering processes.

[0057] Through the change awareness mechanisms described above for the two types of data sources, comprehensive coverage of physical table changes in the two mainstream heterogeneous data sources, namely the first type of data source (such as data lake) and the second type of data source (such as database), is achieved. Zero-latency awareness is achieved for the first type of data source, and minute-level timed awareness is achieved for the second type of data source, ensuring that any structural changes to physical tables can be captured in a timely manner and drive the subsequent reverse engineering process.

[0058] The following describes in detail step 202, namely, "obtaining a revised draft of the logical model of the changed physical table based on the reverse engineering request, wherein the revised draft is obtained by reverse engineering based on the reverse engineering request," with reference to the embodiments.

[0059] In embodiments of this disclosure, a reverse engineering request may be sent to a reverse engineering system, enabling the system to perform reverse engineering processing based on the request and generate a revised draft of the logical model for the modified physical table. Specifically, upon receiving the reverse engineering request, the system can connect to the corresponding data source based on the data source identifier carried in the request. Then, based on the location information of the modified physical table, it obtains the latest metadata of the modified physical table. This latest metadata may include column information (such as column name, physical type, comments, etc.) and table attributes (such as storage engine, partition information, table comments, primary key, index, creation time, update time, etc.). Subsequently, based on the latest metadata of the modified physical table, a revised draft of the logical model for the modified physical table is generated.

[0060] In the aforementioned stage of generating revised drafts, in order to address the potential version conflicts that may arise from various methods of triggering reverse engineering, this application introduces a multi-path conflict resolution strategy based on the lifecycle state of the logical model.

[0061] The lifecycle states of the logical model include, but are not limited to: There is no published main version and no revision draft (here, revision draft refers to the draft version of the logical model, that is, the logical model has not yet been generated). There is no published main version but there is a revision draft (here, revision draft refers to the draft version of the logical model, that is, the logical model is still in the draft state and has not been released. In other words, the logical model only exists in the draft version). There is a published main version but no revision draft (here, revision draft refers to the revision version of the logical model, that is, at this time, the logical model has a published main version, but there is no revision version for the main version). There is a published main version and a revision draft (here, the revision draft refers to the revised version of the logical model, that is, at this time, the logical model has a published main version and a revision version for the main version). The state of obsolescence (i.e., the logical model is no longer in use).

[0062] As a specific implementation, in response to the lifecycle state of the logical model of the changed physical table being in a state where there is no published major version and no revision draft (i.e., the logical model corresponding to the changed physical table has not been created in any version), the reverse engineering system obtains the revision draft of the logical model of the changed physical table based on the reverse engineering request, and identifies this revision draft as the target revision draft of the logical model of the changed physical table. At this time, the revision draft exists as a draft version of the logical model of the changed physical table. For example, in a medical scenario, a new table A is added to the data source. When a change event for table A is detected, the reverse engineering system finds that there is no logical model with the same name as table A, and then generates a revision draft of the logical model of table A based on the latest metadata of table A obtained, and identifies this revision draft as the draft version of the logical model of table A.

[0063] When the logical model of a changed physical table is in a lifecycle state where there is no published master version but a draft revision exists (i.e., the logical model has not been published, but a draft version already exists), the reverse engineering system retrieves the draft revision of the logical model of the changed physical table based on the reverse engineering request and uses this draft revision to overwrite the existing draft version. For example, if table B is in the design phase and is modified multiple times during the design process, each modification triggers reverse engineering. Therefore, the logical model of table B does not have a published master version, but a draft version exists. To ensure that the draft version is consistent with the latest table B, the latest draft revision is directly used to overwrite the draft version.

[0064] When the logical model of a changed physical table is in a lifecycle state where a published master version exists but no revision draft exists (i.e., the logical model has a publicly available master version, but no revisions are currently being edited), the reverse engineering system retrieves the revision draft of the logical model of the changed physical table based on a reverse engineering request. This revision draft is then identified as the target revision draft for the logical model of the changed physical table. At this point, the revision draft exists as a revision version of the logical model of the changed physical table. In other words, the generated revision draft does not directly overwrite the published master version; instead, it exists as a revision version, and the master version continues to provide services unaffected. For example, there is a published master version of the logical model for table C, which is referenced by multiple downstream reports and Extract Transform Load (ETL) tasks. One day, an administrator adds a column to table C in the data source. After detecting the change event for table C, the reverse engineering system generates a reverse engineering request. The system finds that the logical model has been published but has no revision version, and automatically creates a revision draft as the revision version of the logical model.

[0065] In response to a change in the logical model of a physical table, where the lifecycle state includes a published master version and a revision draft (i.e., a revision of the logical model is currently being edited or reviewed), the reverse engineering system retrieves the revision draft of the logical model of the changed physical table based on the reverse engineering request and uses this revision draft to overwrite the existing revision. This approach ensures that for multiple concurrent changes to the same physical table, only the latest revision is maintained, avoiding the chaos caused by multiple revisions coexisting. For example, if there is a published master version of the logical model for table D, which is referenced by multiple downstream reports and ETL tasks, and during the review process of the revision of the logical model of table D, the administrator makes another structural change to table D in the data source, the reverse engineering system detects that a revision of the logical model of table D already exists. Therefore, it directly uses the revision draft to overwrite the existing revision, and subsequently only the latest revision needs to be reviewed.

[0066] If the logical model of a modified physical table is in a deprecated state (i.e., the logical model has been manually marked as deprecated or expired), the reverse engineering system will not execute the step of obtaining a draft revision of the logical model of the modified physical table based on the reverse engineering request; instead, it will directly reject the operation and return an error message. For example, if table E and its logical model have both been marked as deprecated, but an administrator accidentally modifies table E, the reverse engineering system will detect that the logical model of table E is in a deprecated state, stop the subsequent processing flow, and avoid generating an invalid draft revision on a deprecated logical model, thus avoiding wasting computing resources.

[0067] By employing the aforementioned multi-path conflict resolution strategy based on lifecycle states, this solution ensures the isolation between released main versions and revisions, guaranteeing uninterrupted online services. Furthermore, by overwriting draft or revision versions that are currently being edited rather than creating new ones, the uniqueness of draft and revision versions is ensured. This resolves the version confusion and conflict issues that may arise from various methods of triggering reverse engineering, thereby improving the stability and consistency of the logical model.

[0068] The following describes in detail step 203, namely, "calling the large language model and generating an audit report for the revised draft based on at least one of the preset compliance rules and data standards," with reference to the embodiments.

[0069] As one possible approach, a large language model can be invoked to generate an audit report for the revised draft based on at least one of the preset compliance rules and data standards.

[0070] The preset compliance rules can include sensitive data classification rules. For example, mobile phone numbers, email addresses, names, addresses, and bank card numbers are considered sensitive personal identification information; medical record numbers and diagnosis information are considered sensitive personal health information; and credit card numbers and expiration dates are considered sensitive payment card information. Preset compliance rules can also include column naming conventions, such as "must use snake_case format," "must include business semantics, and prohibit meaningless names such as coll and tmp," and "prefix conventions: 'is_' represents a boolean value, 'cnt_' represents a count, and 'amt_' represents an amount." Preset compliance rules may also include physical type change security policies. For example, widening conversion is low-risk, such as converting an integer to a large integer (i.e., INT→BIGINT) and converting a variable-length string (which can store up to 50 characters) to a variable-length string (which can store up to 200 characters) is low-risk (i.e., VARCHAR(50)→VARCHAR(200)); narrowing conversion is high-risk, which may result in data truncation, such as converting a large integer to an integer (i.e., BIGINT→INT) and converting a variable-length string (which can store up to 200 characters) to a variable-length string (which can store up to 50 characters) is high-risk (i.e., VARCHAR(200)→VARCHAR(50)); type conversion is high-risk, which may result in data loss, such as converting a string to an integer (i.e., STRING→INT).

[0071] Data standards refer to a set of common, reusable rules and specifications for a specific column in a physical table. Their purpose is to ensure that the column is understood, defined, named, and represented uniformly throughout the enterprise, thereby eliminating ambiguity and improving data quality. A data standard typically includes: a standard identifier (e.g., "S001"), a standard name (e.g., "mobile number"), a standard column name (e.g., "mobile_phone"), a physical data type (e.g., "STRING"), a data length (e.g., 11 characters), and a data format (e.g., 1[3-9]xxxxxxxxx).

[0072] Specifically, after generating the revised draft, based on at least one of the above compliance rules (such as sensitive data classification rules, column naming conventions, physical type change security policies, etc.) and data standards (which may refer to multiple preset data standards or the data standards corresponding to columns that match the data standards in the main version of the logical model that has been released for changing the physical table), a prompt statement is constructed and provided to the large language model so that the large language model can generate an audit report from the perspective of compliance or data standard matching.

[0073] As another feasible approach, the first major language model is invoked to generate a pre-audit result for the revised draft based on at least one of the preset compliance rules and data standards. Then, the second major language model is invoked to generate an audit report for the revised draft based on the pre-audit result and the change column information. The change column information is derived from the differences between the revised draft and the published master version of the logical model of the changed physical table. For example, it includes information on newly added columns, deleted columns, columns with changed physical types, and renamed columns. Note that if there is no published master version of the logical model of the changed physical table, the columns in the revised draft can be treated as newly added columns, meaning the change column information only includes information on newly added columns. The first and second major language models can refer to the same major language model or two different major language models.

[0074] Specifically, if the pre-audit report is generated based on preset compliance rules, the pre-audit results may include compliance pre-inspection results; if the pre-audit report is generated based on data standards, the pre-audit results may include data standard matching results; and if the pre-audit report is generated based on both preset compliance rules and data standards, the pre-audit results may include both compliance pre-inspection results and data standard matching results.

[0075] When generating a review report for the revised draft based on the pre-review results and change column information, the review report may include the pre-review results and a summary of the change columns. As a specific example, the change column information may be as follows: Figure 3 As shown, Figure 3 The change column summary corresponding to the change column information in the table can be "This change is that the physical table has added two new columns, column name 1 and column name 2, renamed column name 3 to column name 4, changed the physical type of column name 5 from string to integer, and deleted column name 6".

[0076] By splitting the process of generating audit reports into a pre-audit stage and a final audit stage, and calling a large language model twice to perform complex semantic analysis on the revised draft, the pre-audit stage focuses on rule pre-checking or data standard matching, while the final audit report can make comprehensive decisions based on the high-quality pre-audit results. This solves the problem that a single large language model is difficult to guarantee accuracy when handling complex audit tasks, and significantly improves the accuracy of audit reports.

[0077] Next, the process of calling the first language model when the pre-audit results include compliance pre-inspection results will be explained. Specifically, the first language model is called to generate compliance pre-inspection results based on the change column information and compliance rules. The compliance pre-inspection results include at least one of the following: the sensitive column identification results of each change column included in the change column information; the column name standardization verification results of each change column; and the physical type conversion security assessment results of each change column.

[0078] As a specific implementation, based on the changed column information and compliance rules (including sensitive data classification rules, column naming conventions, and physical type change security policies), a first prompt instruction is generated. This first prompt instruction instructs the first language model to generate a compliance pre-check result based on the changed column information and compliance rules. The first prompt instruction is then provided to the first language model to obtain the compliance pre-check result.

[0079] In other words, the first prompt instruction here can be used to instruct the first major language model to evaluate each changed column (including added columns, deleted columns, columns with changed physical types, and renamed columns) using sensitive data classification rules, column naming specifications, and physical type change security policies. This results in the following assessments for each changed column: sensitive column identification (whether the changed column is sensitive), column name specification verification (whether the changed column name conforms to the column naming specifications), and physical type conversion security assessment (whether the physical type change of the changed column is secure). Furthermore, the compliance pre-inspection results can also include the change risk level (low, medium, high) for the changed column and assessment suggestions for the changed column (e.g., if the changed column name does not conform to the column naming specifications, modification suggestions can be given; if a newly added column is a sensitive column, it can be suggested that the data in the newly added column be encrypted and stored, etc.).

[0080] For example, in a shopping scenario, the order table undergoes a structural change, adding a column named "home_Address" (address). Subsequently, the changed column information and compliance rules are submitted to the First Language Model. The First Language Model's compliance pre-check results include the following sensitive column identification results: "The newly added column 'home_Address' is a medium-risk sensitive column; it is recommended to encrypt its storage or configure a dynamic desensitization strategy." The column name specification verification results include: "The newly added column 'home_Address' does not conform to the column naming specification and can be changed to 'home_address'." It should be noted that this structural change to the order table does not include a physical type change; therefore, the First Language Model does not need to perform a physical type conversion security assessment on the changed column information, meaning the compliance pre-check results do not include a physical type conversion security assessment result.

[0081] By clearly defining the content included in the compliance pre-inspection results, the first language model can conduct compliance pre-inspections from three dimensions: sensitive column identification, naming convention verification, and physical type conversion security assessment. The resulting compliance pre-inspection results are more comprehensive and accurate, providing a more reliable basis for subsequent audit reports.

[0082] As another specific embodiment, based on the changed column information, the changed physical table name, the changed logical model type of the physical table's logical model, and compliance rules, a first prompt instruction is generated. This first prompt instruction is used to prompt the first major language model to generate a compliance pre-check result based on the changed column information, the changed physical table name, the changed logical model type of the physical table's logical model, and the aforementioned compliance rules. The first prompt instruction is then provided to the first major language model to obtain the compliance pre-check result.

[0083] In some embodiments, the information on changing column types, changing the name of the physical table, changing the logical model type of the physical table's logical model, and compliance rules can be filled into a preset first prompt instruction template. This first prompt instruction template includes: role setting, used to position the first major language model as an expert in a specific domain (e.g., "You are an enterprise data compliance expert"); changing the name of the physical table (e.g., fact_demo_ecommerce_order_detail); changing the logical model type of the physical table's logical model (e.g., fact table, source layer table, dimension table, etc.); changing column information, including information on newly added columns, deleted columns, columns with changed physical types, and renamed columns; compliance rules, including sensitive data classification rules, column naming conventions, and physical type change security policies; and output format requirements, returning the compliance pre-inspection results in a preset JSON format.

[0084] The ability to change the physical table name and the logical model type of the physical table's logical model can be extracted from reverse engineering requests. By adding these information to the first prompt instruction, the problem of potential misjudgments by the first language model when lacking business context is resolved. For example, if an address column is added to a physical table with a corresponding logical model type of FACT_TABLE, the first language model, combined with the logical model type, can more accurately determine that directly storing an explicit address in the FACT_TABLE violates the data minimization principle and alert administrators in the compliance pre-check results. Similarly, if a column named "Total Price" is added to a table named "Health Information Table," the first language model, combined with the table name, can determine that the total price is unrelated to the health information table. Therefore, the compliance pre-check results can alert administrators to the possibility of an incorrect column name.

[0085] By incorporating the change column information, table names, logical model types, and compliance rules required for compliance pre-inspection into the first prompt instruction, the first major language model can obtain a complete evaluation basis, effectively improving the accuracy and consistency of compliance pre-inspection results.

[0086] It should be noted that the compliance pre-check results do not necessarily need to include sensitive column identification results, column name specification verification results, and physical type conversion security assessment results every time. The First Language Model can perform the corresponding compliance pre-check based on the actual compliance pre-check requirements corresponding to the changed column information. For example, if the changed column information only includes information on physical type changed columns, then the First Language Model only needs to perform a physical type conversion security assessment on the changed columns. In this case, the compliance pre-check result only includes the physical type conversion security assessment result. This approach can effectively control the resource consumption and response latency of calling the First Language Model.

[0087] During the generation of the revision draft, the reverse engineering system can obtain the columns in the published master version of the logical model of the changed physical tables that match the data standards. By performing precise string comparison between the column names in the revision draft and the column names in the master version that match the data standards, the system matches the data standards corresponding to the columns in the master version that match the data standards to the columns in the revision draft with the same column names. For example, if the revision draft includes a column named "mobile" and the master version also includes a column named "mobile," and this column in the master version matches the data standard identified as "std-001," then because the "mobile" column in the revision draft has the same column name as the "mobile" column in the master version, the "mobile" column in the revision draft can be directly matched to the data standard identified as "std-001." In other words, when generating the revision draft, the data standards corresponding to each column in the master version can be inherited into the revision draft through precise string comparison.

[0088] However, the column names in the revision draft may change compared to the main version, making it impossible to match the data standards in the revision draft using the above method. Therefore, this disclosure provides a method for performing data standard matching by calling the first major language model. In this case, the pre-audit result includes the data standard matching result. Specifically, the column names of the columns in the revision draft that do not match the data standards (i.e., the column names of the columns in the revision draft that do not match the data standards through the above precise string comparison method) are determined. The first major language model is then called to generate the data standard matching result based on the column names of the columns in the revision draft that do not match the data standards, the column names of the columns in the main version that match the data standards, and the data standards corresponding to the columns in the main version that match the data standards. The data standard matching result includes the matching relationship between the columns in the revision draft that do not match the data standards and the columns in the main version that match the data standards.

[0089] By introducing a semantic matching mechanism based on the first major language model after a failure of precise string comparison, the business meaning of column names can be understood, and semantically equivalent columns such as "mobile phone number" and "telephone number" can be automatically discovered. This solves the problem of data standard inheritance failure caused by the inability of traditional precise string comparison to handle column renaming scenarios, achieves accurate matching of data standards, and significantly reduces manual maintenance costs.

[0090] As a specific implementation, based on the column names of columns that do not match the data standard in the draft revision, the column names of columns that match the data standard in the main version, and the data standard corresponding to the columns that match the data standard in the main version, a second prompt instruction is generated. This second prompt instruction instructs the first language model to generate a data standard matching result based on the column names of columns that do not match the data standard in the draft revision, the column names of columns that match the data standard in the main version, and the data standard corresponding to the columns that match the data standard in the main version. The second prompt instruction is then provided to the first language model to obtain the data standard matching result.

[0091] For example, in a change event, the column name of a column in the physical table is changed from "phone" to "mobile". In this case, the primary version includes a column named "phone", and this column matches the data standard identified by "std-001". The revision draft includes a column named "mobile". Because the "mobile" column in the revision draft is inconsistent with the "phone" column in the primary version, it is impossible to match the data standard through precise string comparison. This solution uses the first major language model to achieve semantic matching. That is, the first major language model performs a semantic comparison between the "mobile" column in the revision draft and the "phone" column in the primary version. It finds that the semantic similarity of the column names of these two columns meets the similarity threshold, meaning there is a matching relationship between the two columns. Therefore, the data standard "std-001" corresponding to the "phone" column in the primary version can be directly matched to the "mobile" column in the revision draft.

[0092] To make the matching process more accurate, this scheme can also add physical type to the second prompt instruction, so that the first language model can not only make judgments based on the semantic similarity of column names, but also evaluate the reliability of the matching by combining the consistency of physical type.

[0093] As another specific embodiment, based on the column names of columns that do not match the data standard in the draft revision, the physical types corresponding to the columns that do not match the data standard in the draft revision, the column names of columns that match the data standard in the main version, and the data standard and physical type corresponding to the columns that match the data standard in the main version, a second prompt instruction is generated. This second prompt instruction instructs the first language model to generate a data standard matching result based on the column names of columns that do not match the data standard in the draft revision, the physical types corresponding to the columns that do not match the data standard in the draft revision, the column names of columns that match the data standard in the main version, and the data standard and physical type corresponding to the columns that match the data standard in the main version. The second prompt instruction is then provided to the first language model to obtain the data standard matching result.

[0094] Using the previous example, the first language model performs a semantic comparison between the "mobile" column in the revised draft and the "phone" column in the main version. It finds that the semantic similarity of the column names of the two columns meets the similarity threshold. If the physical types of the two columns are the same, it can be considered that there is a matching relationship between the two columns. The data standard "std-001" corresponding to the "phone" column in the main version can be directly matched to the "mobile" column in the revised draft. If the physical types of the two columns are different, it can be considered that there is no matching relationship between the two columns. The "mobile" column in the revised draft cannot directly inherit the data standard "std-001" corresponding to the "phone" column in the main version.

[0095] As a more complete embodiment, the second prompt instruction can be as follows: Figure 4 As shown, Figure 4 The second prompt instruction is input into the first large language model, and the resulting data standard matching results can be as follows: Figure 5 As shown, in Figure 5 The data standard matching results also include confidence levels. Different inheritance suggestions are given for matching relationships with different confidence levels. For example, for a confidence level greater than or equal to 0.9, columns in the draft revision that do not match the data standard can directly inherit the data standard corresponding to columns in the main version that match the data standard. For a confidence level between 0.7 and 0.9, manual confirmation is required before matching columns in the draft revision that do not match the data standard with columns in the main version that match the data standard. For a confidence level less than or equal to 0.7, there is no matching relationship between columns in the draft revision that do not match the data standard and columns in the main version that match the data standard, so a new data standard needs to be selected for association.

[0096] By adding the constraint of the physical type of the column to the second prompt instruction, the first language model can make a comprehensive judgment on whether the physical types are compatible when performing semantic matching, thereby avoiding the problem of matching columns of different physical types and improving the accuracy and reliability of semantic matching.

[0097] It should be noted that the prerequisite for performing the above precise string comparison and semantic matching is that the logical model of the physical table being modified has a published main version. If the logical model of the physical table being modified does not have a published main version, a second hint instruction can be generated based on the column names and physical types of all columns in the revision draft, the standard name and physical type of the preset data standard. This second hint instruction is then provided to the first language model, so that the first language model can perform semantic matching between the columns in the revision draft and the preset data standard based on the column names and physical types of all columns in the revision draft, the standard name and physical type of the preset data standard. If a match is found, the columns in the revision draft can be directly associated with the preset data standard.

[0098] Before invoking the second largest language model and generating a final review report for the revised draft based on the pre-review results and change column information, this disclosure also provides a method for optimizing the revised draft.

[0099] First, based on the revised draft, the column information for each column in the physical table is determined. The column information includes the column name, physical type, and first comment. The first comment refers to the comment information (i.e., business semantics) for each column obtained from the revised draft, and this first comment may be empty.

[0100] As one feasible approach, the third language model is invoked for the first time. Based on column information, the changed physical table name, and the logical model type of the changed physical table's logical model, a second comment is determined for each column. This second comment is obtained by completing the first comment. Then, the third language model is invoked a second time. Based on column information, the changed physical table name, and the logical model type of the changed physical table's logical model, the logical type of each column is determined. The logical type refers to the business role each column plays in the logical model, such as a dimension, measure, or partitioning field. A dimension refers to a column that is a descriptive attribute, category identifier, or foreign key; a measure refers to a column that is an aggregatable numerical indicator; and a partitioning field refers to a column used for data partitioning (usually time or region).

[0101] As another possible approach, since the inputs of the two aforementioned task calls are highly consistent, the two calls can be merged into one call to save resources. Specifically, the third language model is called, and based on column information, the name of the physical table is changed, and the logical model type of the logical model of the physical table is changed, the logical type of each column and the second comment of each column are determined. The second comment is obtained by completing the first comment.

[0102] Specifically, a fourth prompt instruction can be generated based on column information, the name of the physical table being changed, and the logical model type of the logical model of the physical table being changed. This fourth prompt instruction instructs the third language model to determine the logical type of each column and the second comment for each column based on the column information, the name of the physical table being changed, and the logical model type of the logical model of the physical table being changed. The fourth prompt instruction is then provided to the third language model to obtain the logical type of each column and the second comment for each column.

[0103] For example, the fourth prompt instruction can be as follows: Figure 6 As shown, Figure 6 The fourth prompt instruction shown is input into the third language model, resulting in the following: Figure 7 The output shown includes the logical type of each column and a second comment for each column.

[0104] Existing technologies typically determine the logical type of each column in a draft revision based solely on the physical type. For example, the logical type of a column with a physical type of string is determined as a dimension. However, this solution utilizes the semantic understanding capabilities of the third language model to accurately determine the logical type based on column information, table name, and logical model type. It also automatically generates comments with clear business meaning for columns that lack comments, thereby significantly improving the quality of the draft revision and saving human resources.

[0105] The above content mentions that, by invoking the second language model and generating an audit report for the revised draft based on the pre-audit results (including compliance pre-check results and / or data standard matching results) and change column information, one feasible approach is to generate a third prompt instruction based on the change column information and pre-audit results. This third prompt instruction prompts the second language model to generate an audit report based on the change column information and pre-audit results. The audit report must at least include a summary of the change columns and the pre-audit results. Providing the third prompt instruction to the second language model yields the audit report.

[0106] As another feasible approach, a third prompt instruction is generated based on the changed column information, downstream dependency information, and pre-audit results. This third prompt instruction prompts the second language model to generate an audit report based on the changed column information, downstream dependency impact analysis, and pre-audit results. The audit report includes at least a summary of the changed columns, a downstream dependency impact analysis, and the pre-audit results. The third prompt instruction is then provided to the second language model to obtain the audit report.

[0107] Downstream dependency information refers to the dependency information of downstream dependent items on the changed physical table. Specifically, the reference relationship table can be queried to obtain which downstream dependent items depend on the changed physical table, including reports, ETL tasks, and data metrics, and the specific column names referenced by each downstream dependent item are recorded. Downstream dependency impact analysis can include downstream dependent items, the specific column names referenced, and recommended information (such as recommending that the administrators of downstream dependent items confirm the compatibility of column name changes).

[0108] By inputting change column information, downstream dependency information, and pre-audit results into the second language model, the generated audit report is ensured to be comprehensive and of decision-making reference value. This solves the problems of high audit knowledge threshold, low efficiency, and inconsistent conclusions, and achieves standardization and efficiency of the audit process.

[0109] To further optimize the data standard matching process, as another feasible approach, in response to situations where there is a mismatch between columns in the draft revision that do not match data standards and columns in the main version that do match data standards, candidate data standards are retrieved from the data standard library. The physical type of the candidate data standard must be the same as the physical type of the column in the draft revision that does not match data standards, and the business domain (e.g., e-commerce) of the candidate data standard must be the same as the business domain of the draft revision. Furthermore, to ensure the response speed of the second largest language model, the number of candidate data standards can be limited. For example, they can be sorted by the number of times they are associated, from highest to lowest, and the top 50 can be determined as the final candidate data standards.

[0110] It should be noted that the above-mentioned method for screening candidate data standards is not fixed. In some embodiments, the business domain to which the candidate data standards belong is the same as the business domain to which the revised draft belongs. In other embodiments, if there are few candidate data standards that can be obtained, candidate data standards can also be obtained from other business domains.

[0111] Specifically, before constructing the third-party prompt instruction, the data standard matching results in the pre-audit results can be analyzed. If it is found that a column in the revision draft fails to establish an association with the data standard corresponding to the column that matches the data standard in the main version or the preset data standard through precise string comparison or semantic matching (e.g., confidence level below 0.7 or no matching relationship), then candidate data standards are obtained from the data standard library. It should be noted that for each column in the revision draft that does not match a data standard, its corresponding candidate data standard needs to be obtained.

[0112] Then, based on the change column information, downstream dependency information, candidate data standards, and pre-audit results, a third prompt instruction is generated. The change column information, downstream dependency information, candidate data standards, and pre-audit results are collected in advance by the logic model management device. That is, the audit context is collected according to the preset query logic. The audit context includes change column information, downstream dependency information, compliance pre-inspection results, data standard matching results, candidate data standards, etc. The third prompt instruction is obtained based on the audit context. For example, the change column information, downstream dependency information, candidate data standards, and pre-audit results in the audit context are assembled into a preset third prompt instruction template to obtain the third prompt instruction.

[0113] The third prompt instruction is provided to the second language model to obtain the review report. In other words, the second language model can select candidate data standards that are associated with columns that do not match the data standards in the revision draft.

[0114] The audit report at this stage may include a summary of the change columns, downstream dependency impact analysis, compliance pre-inspection results (if any), and a data standard matching summary. The data standard matching summary is determined based on the data standard matching results generated by calling the first major language model and the data standard matching results obtained when calling the second major language model. Additionally, the audit may include the overall change risk level (e.g., low, medium, high) based on change column information, downstream dependency information, candidate data standards, and pre-audit results.

[0115] It should be noted that during the process of generating the review report using the second language model, the output of the second language model can be verified first based on the preset output structure. If the output of the second language model conforms to the preset output structure, then the output of the second language model can be used as the review report.

[0116] Specifically, the output of the second language model is parsed, and structured content is extracted from the preset start and end marks in the output of the second language model. If the structured content conforms to the preset output structure, the output of the second language model is used as an audit report. If the structured content does not conform to the preset output structure, the second language model is called again or the failure event of this call is written to the exception queue (the exception queue can be analyzed manually later).

[0117] By providing a filtered candidate data standard to the second language model in the third prompt instruction and explicitly requiring it to make recommendations only from the candidate data standard, the second language model is prevented from freely outputting incorrect data standards due to a lack of supporting information. Furthermore, by pre-screening the data standards in the data standard library, useless interference items are eliminated, making the matching recommendations of the second language model more accurate and improving the reliability of the review report.

[0118] The following describes in detail step 204, namely "publishing the revised draft in response to the confirmation operation of the revised draft based on the review report," with reference to the embodiments.

[0119] After generating the audit report, it is displayed on the audit interface, allowing auditors to confirm the revised draft based on the report. Auditors can quickly make decisions based on the report; for example, they can click the "Approve" button to directly publish the revised draft, or select "Conditional Approve," which means adopting the recommended data standards from the audit report for columns in the revised draft that did not match the data standards before publishing.

[0120] The publication process of this disclosure includes, but is not limited to, the following sub-steps.

[0121] The first sub-step: Obtain and store the materialization configuration information and materialization status of the published master version of the logical model for changing the physical table. Materialization configuration information refers to the relevant parameters required to convert the logical model into a physical table (i.e., synchronizing the physical table using the logical model), such as data source type (e.g., Doris, Iceberg), table name, partitioning strategy, etc.; materialization status refers to the synchronization status of the current master version relative to the physical table, such as synchronized, requiring synchronization, etc.

[0122] The second sub-step: Delete the published major version.

[0123] The third sub-step: finalize the revised draft as the new master version.

[0124] The fourth sub-step: Generate a change log based on the released major version and the new major version. The change log records all column-level differences between the released major version and the new major version, including any additions, deletions, modifications, or renamings.

[0125] The fifth sub-step: Based on the reference relationship information of the published master version, update the reference relationship information of the new master version to ensure that all downstream dependencies pointing to the published master version (such as reports, ETL tasks, data metrics) can automatically switch to the new master version.

[0126] It should be noted that, to ensure the integrity of version switching and data consistency, the above five sub-steps are executed atomically within the same release transaction. This means that if any sub-step fails (e.g., a lock conflict occurs when deleting a released major version), the entire release transaction will be rolled back, preventing intermediate states of partial success and partial failure. This approach avoids problems such as version conflicts, data loss, or broken references caused by exceptions, ensuring the safety and reliability of version switching in the logical model.

[0127] Additionally, to help non-technical managers understand changes to physical tables and reduce communication costs, the fourth language model can be invoked. Based on the change log generated in the fourth sub-step and the changed physical table name, a change summary can be generated, and the change summary and change log can be stored together. As a specific example, the changed physical table name is "User Profile Table." The change log may include: "Added column: home_address (string); Renamed column: phone → mobile." The change summary is: "This revision adds an address field to the User Profile Table, unifying the mobile number column name from 'phone' to 'mobile' to conform to naming conventions. The newly added address field involves sensitive personal information; it is recommended to pay attention to data anonymization configuration."

[0128] In the above process, the revision draft is obtained in response to the detection of a change event for the physical table. That is, the physical table undergoes a structural change first. After this solution detects the structural change of the physical table, it generates a revision draft and uses the revision draft to update the published main version of the logical model of the physical table. At this time, when the revision draft is published, there is no need to perform the step of synchronizing the physical table.

[0129] In some embodiments, administrators can directly modify the logical model of a physical table to generate a revision draft. When publishing the revision draft, it is necessary to perform the step of synchronizing the physical table.

[0130] This solution assigns a source tag to a revision draft, distinguishing whether the step of synchronizing physical tables is required. Specifically, the source tag for the revision draft is determined, including a first tag and a second tag. The first tag indicates that the revision draft was generated in response to the detection of a change event targeting a modified physical table. For example, when an administrator executes "ALTER TABLE" in the database, and this change event is obtained through the database's audit log file, the revision draft generated by the reverse engineering system driven by this event will have its source tag set to the first tag. The second tag indicates that the revision draft was generated in response to an event that triggers manual editing of the logical model. That is, a revision draft saved after an administrator directly modifies the definition of the logical model on the data modeling platform will have its source tag set to the second tag.

[0131] Based on this, the above-mentioned release process for the revised draft also includes the following sub-steps: in response to the source marker of the revised draft being marked as the first marker, the materialization status of the logical model of the changed physical table is set to synchronized; in response to the source marker of the revised draft being marked as the second marker, the materialization status of the logical model of the changed physical table is set to synchronized.

[0132] After executing the above publishing process, in response to the materialized state of the logical model of the changed physical table being synchronized, the changed physical table is updated. For example, if the changed physical table to be updated belongs to the second type of data source, DDL statements are executed through Java Database Connectivity (JDBC); if the changed physical table to be updated belongs to the first type of data source, incremental column updates are performed through its metadata operation interface.

[0133] In other words, when the revision draft originates from a change event affecting the physical table, the physical table itself is already up-to-date, and the logical model has only undergone reverse engineering. Therefore, there is no need to re-synchronize the logical model to the physical table. Conversely, when the revision draft originates from manually edited logical models, the logical model contains changes that the physical table does not yet have. Therefore, the materialization process must be triggered to synchronize the logical model to the physical table.

[0134] By differentiating the sources of revision drafts and implementing differentiated materialization status settings accordingly, accurate judgment of physical table synchronization requirements is achieved. For changes generated in reverse from physical tables, unnecessary materialization steps can be skipped, avoiding unnecessary computational overhead; for manually initiated logical model changes, materialization synchronization is triggered, ensuring the consistency between the logical model and the physical table and improving the reliability of automated synchronization.

[0135] It should be noted that the compliance pre-inspection results, data standard matching results, and audit reports in this solution can also be generated by rules or manually entered. Therefore, upon obtaining the compliance pre-inspection results, data standard matching results, and audit reports, generation source tags can be applied to these results. These generation source tags include a first generation tag, a second generation tag, and a third generation tag. The first generation tag indicates that the compliance pre-inspection results, data standard matching results, and audit reports are generated based on the corresponding rules. The second generation tag indicates that the compliance pre-inspection results, data standard matching results, and audit reports are manually entered. The third generation tag indicates that the compliance pre-inspection results, data standard matching results, and audit reports are generated by calling the large language model. In response to receiving manual modifications, the results of the manual modifications are used to overwrite the results generated by the large language model or the results generated based on the rules, and the corresponding generation source tag is updated to the second generation tag.

[0136] It should also be noted that this solution provides a degradation method for large language model call failure. Specifically, when the call to the large language model fails, times out, or the output result does not meet the preset output structure, degradation processing can be performed. Degradation processing includes at least one of the following: retaining the compliance pre-inspection results, data standard matching results, audit reports, etc. generated by the corresponding rules; setting the fields in the results that do not meet the preset requirements to empty; marking the corresponding results as pending manual confirmation; and writing the call failure event to the exception queue.

[0137] It should also be noted that in the above-mentioned tasks that call the large language model multiple times, the same large language model or different large language models can be called. As a specific implementation, large language models with different capability levels can be selected based on the task complexity, response latency, resource consumption, etc. of each call task. For example, the third large language model can refer to the large language model with the first capability level, and the first and second large language models can refer to the large language models with the second capability level. The semantic reasoning ability of the large language model with the second capability level is higher than that of the large language model with the first capability level.

[0138] Next, to explain the logical model management method in this disclosure in more detail, a more complete embodiment is given below, such as... Figure 8 As shown.

[0139] In response to the detection of a change event targeting a physical table (either an event that obtains metadata change information through the metadata operation interface of the first type of data source or an event that obtains the audit log file of the second type of data source and extracts the DDL change statement from the audit log file), a reverse engineering request is generated based on the information of the change event.

[0140] A reverse engineering request is sent to the reverse engineering system. Upon receiving the request, the system connects to the corresponding data source based on the data source identifier carried in the request. Then, based on the location information of the changed physical table, it obtains the latest metadata of that table. This latest metadata may include column information (such as column name, physical type, comments, etc.) and table attributes (such as storage engine, partition information, table comments, primary key, indexes, creation time, update time, etc.). Subsequently, based on the latest metadata of the changed physical table, a revised draft of the logical model of the changed physical table is generated.

[0141] Based on the lifecycle status of the logical model that changes the physical table, the revision draft is used as either a draft version or a revision version. For example, if the logical model that changes the physical table has a published main version and a revision version of the main version, the revision draft is used to overwrite the revision version, that is, the revision draft is used as the only revision version of the logical model, and the revision version can be published later.

[0142] Then, since there are no input-output dependencies between the following invocation tasks, the following three invocation tasks can be executed in parallel: Invoking the first major language model to generate compliance pre-check results based on changed column information, changed physical table names, changed logical model types of the physical table's logical model, and compliance rules; Invoking the first major language model to generate data standard matching results based on column names of columns that do not match data standards in the revision draft, the physical types corresponding to columns that do not match data standards in the revision draft, column names of columns that match data standards in the main version, and the data standards and physical types corresponding to columns that match data standards in the main version; Invoking the third major language model to determine the logical types of each column and the second comments for each column based on column information (including column names, physical types, and first comments), changed physical table names, and changed logical model types of the physical table's logical model. The second comments are obtained by completing the first comments. Based on the logical types of each column and the second comments for each column, the revision draft is updated. The first, second, and third major language models can refer to the same major language model or different major language models.

[0143] Furthermore, in response to a situation where there is a mismatch between columns in the draft revision that do not match the data standard and columns in the main version that do match the data standard, candidate data standards are obtained from the data standard library. The physical type corresponding to the candidate data standard is the same as the physical type corresponding to the columns in the draft revision that do not match the data standard, and the business domain to which the candidate data standard belongs is the same as the business domain to which the draft revision belongs.

[0144] The second language model is invoked to generate audit results based on change column information, downstream dependency information, candidate data standards, compliance pre-inspection results, and data standard matching results. The audit results include a change column summary, downstream dependency impact analysis, compliance pre-inspection results, and data standard matching summary.

[0145] The audit report is presented to the administrator. In response to the administrator's confirmation of the revised draft based on the audit report, the following release process is performed on the revised draft: Obtain and store the materialized configuration information and materialized status of the published master version of the logical model of the changed physical table; delete the published master version; identify the revised draft as the new master version; generate a change log based on the published master version and the new master version; update the reference relationship information of the new master version based on the reference relationship information of the published master version; in response to the source tag of the revised draft being marked as the first tag, set the materialized status of the logical model of the changed physical table to synchronized; in response to the source tag being marked as the second tag, set the materialized status of the logical model of the changed physical table to synchronized.

[0146] Subsequently, in response to the materialization status of the logical model of the changed physical table being required to be synchronized, the changed physical table is updated, and the materialization status of the logical model of the changed physical table is updated to synchronized.

[0147] It should be noted that in the above process, the large language model is used as an intelligent processing node in the logic model deployment process. The logic model management device predetermines the timing of the large language model's invocation, input context, and output structure. The large language model is not responsible for determining the flow of the process, nor does it need to independently select external tools. Before invoking the large language model, the logic model management device collects the required context through deterministic code and populates the context into a predefined prompt instruction template. After invoking the large language model, the output results of the large language model are subjected to structured parsing and validation. After successful validation, the output results are written to the corresponding business fields.

[0148] In some embodiments, the output of the large language model is constrained to a preset structured format, such as JSON, key-value pairs, table structure, or other structures that can be parsed by the program. After obtaining the output of the large language model, the structured content is extracted first. For example, when the output of the large language model is in JSON format, the result can be directly parsed first. If direct parsing fails, the structured content is extracted from the corresponding code block; if it still fails, the structured content is extracted from the preset start and end characters in the output result.

[0149] Then, the structured content is validated to ensure it contains preset fields, that the field types meet the requirements, and that the enumerated values ​​are within the preset range. If the validation fails, the large language model can be called again; if it still fails after retrying, the corresponding call failure event is written to the exception queue, and degradation processing is performed.

[0150] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0151] The foregoing has described specific embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0152] According to another embodiment, a logic model management device is provided. Figure 9 A schematic block diagram of the logic model management device according to one embodiment is shown. Figure 9 As shown, the device 900 includes an event response unit 901, a draft acquisition unit 902, a report generation unit 903, and a draft publishing unit 904, and further includes a materialization synchronization unit 905. The main functions of each component are as follows: Event response unit 901 is configured to generate a reverse engineering request based on information from a detected change event targeting a physical table.

[0153] The draft acquisition unit 902 is configured to acquire a revised draft of the logical model of the changed physical table based on a reverse engineering request. The revised draft is obtained by reverse engineering based on the reverse engineering request.

[0154] The report generation unit 903 is configured to invoke a large language model to generate an audit report for the revised draft based on at least one of preset compliance rules and data standards.

[0155] Draft release unit 904 is configured to release the revised draft in response to the confirmation operation of the revised draft based on the review report.

[0156] As one possible implementation method, when the report generation unit 903 calls the large language model to generate an audit report for the revised draft based on at least one of the preset compliance rules and data standards, it can be specifically configured as follows: calling the first large language model to generate a pre-audit result for the revised draft based on at least one of the preset compliance rules and data standards; calling the second large language model to generate an audit report for the revised draft based on the pre-audit result and the change column information, wherein the change column information is obtained based on the differences between the revised draft and the main version of the logical model of the changed physical table that has been released.

[0157] As one possible approach, pre-audit results include compliance pre-inspection results.

[0158] The report generation unit 903, when calling the first language model and generating a pre-audit result for the revised draft based on at least one of the preset compliance rules and data standards, can be specifically configured to: call the first language model and generate a compliance pre-inspection result based on the change column information and compliance rules; wherein the compliance pre-inspection result includes at least one of the following: the sensitive column identification result of each change column included in the change column information; the column name specification verification result of each change column; and the physical type conversion security assessment result of each change column.

[0159] As one possible implementation method, when the report generation unit 903 calls the first language model to generate compliance pre-inspection results based on the changed column information and compliance rules, it can be specifically configured to: generate a first prompt instruction based on the changed column information, the name of the changed physical table, the logical model type of the changed physical table's logical model, and compliance rules. The first prompt instruction is used to prompt the first language model to generate compliance pre-inspection results based on the changed column information, the name of the changed physical table, the logical model type of the changed physical table's logical model, and compliance rules; and provide the first prompt instruction to the first language model to obtain the compliance pre-inspection results.

[0160] As one possible approach, pre-audit results include data standard matching results.

[0161] The report generation unit 903, when calling the first major language model to generate a pre-audit result for the revised draft based on at least one of the preset compliance rules and data standards, can be specifically configured to: determine the column names of columns in the revised draft that do not match the data standards; call the first major language model to generate a data standard matching result based on the column names of columns in the revised draft that do not match the data standards, the column names of columns in the main version that match the data standards, and the data standards corresponding to the columns in the main version that match the data standards; wherein, the data standard matching result includes the matching relationship between the columns in the revised draft that do not match the data standards and the columns in the main version that match the data standards.

[0162] As one possible implementation method, when the report generation unit 903 calls the first language model to generate data standard matching results based on the column names of columns that do not match the data standard in the revision draft, the column names of columns that match the data standard in the main version, and the data standard corresponding to the columns that match the data standard in the main version, it can be specifically configured to: generate a second prompt instruction based on the column names of columns that do not match the data standard in the revision draft, the physical type corresponding to the columns that do not match the data standard in the revision draft, the column names of columns that match the data standard in the main version, and the data standard and physical type corresponding to the columns that match the data standard in the main version. The second prompt instruction is used to instruct the first language model to generate data standard matching results based on the column names of columns that do not match the data standard in the revision draft, the physical type corresponding to the columns that do not match the data standard in the revision draft, the column names of columns that match the data standard in the main version, and the data standard and physical type corresponding to the columns that match the data standard in the main version; and provide the second prompt instruction to the first language model to obtain the data standard matching results.

[0163] As one possible implementation method, before generating an audit report for the revised draft by calling the second language model based on the pre-audit results and the changed column information, the report generation unit 903 can also be configured to: determine the column information of each column in the changed physical table based on the revised draft, including the column name, physical type, and first comment; call the third language model to determine the logical type of each column and the second comment of each column based on the column information, the table name of the changed physical table, and the logical model type of the logical model of the changed physical table, the second comment being obtained by completing the first comment; and update the revised draft based on the logical type of each column and the second comment of each column.

[0164] As one possible implementation method, when the report generation unit 903 calls the second language model to generate an audit report for the revised draft based on the pre-audit results and change column information, it can be specifically configured to: generate a third prompt instruction based on the change column information, downstream dependency information, and pre-audit results. The third prompt instruction is used to prompt the second language model to generate an audit report based on the change column information, downstream dependency information, and pre-audit results. The audit report includes at least a change column summary, downstream dependency impact analysis, and pre-audit results. The third prompt instruction is then provided to the second language model to obtain the audit report.

[0165] As one possible implementation method, the report generation unit 903, when generating a third prompt instruction based on change column information, downstream dependency information, and pre-audit results, can be specifically configured as follows: In response to a mismatch between columns in the draft revision that do not match data standards and columns in the main version that do match data standards, it retrieves candidate data standards from the data standard library. The physical type corresponding to the candidate data standard is the same as the physical type corresponding to the columns in the draft revision that do not match data standards, and the business domain to which the candidate data standard belongs is the same as the business domain to which the draft revision belongs; and generates a third prompt instruction based on the change column information, downstream dependency information, candidate data standards, and pre-audit results.

[0166] As one possible approach, change events include at least one of the following: obtaining metadata change events through the metadata operation interface of a first type of data source; obtaining audit log files from a second type of data source and extracting data definition language (DDL) change statements from the audit log files.

[0167] As one possible approach, the second type of data source is a database.

[0168] The event response unit 901, when generating a reverse engineering request based on information about a change event, can be specifically configured to: determine the first location information of the changed physical table based on the DDL change statement; and generate a reverse engineering request based on the data source identifier corresponding to the audit log file and the first location information of the changed physical table.

[0169] As one possible implementation, the draft acquisition unit 902, when acquiring a revision draft of the logical model of the changed physical table based on a reverse engineering request, can be specifically configured as follows: In response to the lifecycle state of the logical model of the changed physical table being in a state where there is no published main version and no revision draft, it acquires a revision draft of the logical model of the changed physical table based on a reverse engineering request, and identifies the acquired revision draft as the target revision draft of the logical model of the changed physical table; In response to the lifecycle state of the logical model of the changed physical table being in a state where there is a published main version but no revision draft, it acquires a revision draft of the logical model of the changed physical table based on a reverse engineering request, and identifies the acquired revision draft as the target revision draft of the logical model of the changed physical table. The following steps are performed: 1. If the logical model of the changed physical table is in a lifecycle state where a released major version exists and a revision draft exists, the revision draft of the logical model of the changed physical table is obtained through reverse engineering, and the obtained revision draft is used to overwrite the existing revision draft. 2. If the logical model of the changed physical table is in a lifecycle state where a released major version does not exist but a revision draft exists, the revision draft of the logical model of the changed physical table is obtained through reverse engineering, and the obtained revision draft is used to overwrite the existing revision draft. 3. If the logical model of the changed physical table is in a lifecycle state that is obsolete, the step of obtaining the revision draft of the logical model of the changed physical table through reverse engineering is not executed.

[0170] As one possible implementation method, the draft release unit 904, when releasing the revised draft, can be specifically configured to: obtain and store the materialized configuration information and materialized status of the published main version of the logical model of the changed physical table; delete the published main version; determine the revised draft as the new main version; generate a change log based on the published main version and the new main version; and update the reference relationship information of the new main version based on the reference relationship information of the published main version.

[0171] As one possible implementation, the draft publication unit 904 can also be configured to: determine the source marker of the revised draft, the source marker including a first marker and a second marker, the first marker being used to characterize that the revised draft is obtained in response to the detection of a change event for a changed physical table, and the second marker being used to characterize that the revised draft is obtained in response to an event that triggers a manual editing logic model.

[0172] The draft publishing unit 904 can also be configured to: in response to the source flag being the first flag, set the materialization state of the logical model of the changed physical table to synchronized; and in response to the source flag being the second flag, set the materialization state of the logical model of the changed physical table to synchronized.

[0173] Furthermore, the materialization synchronization unit 905, after the revised draft is published, can be specifically configured to update the changed physical table in response to the materialization state of the logical model of the changed physical table being required to be synchronized.

[0174] As one possible implementation method, the draft release unit 904 can also be configured to: call the fourth language model, generate a change summary based on the change log and the table name of the changed physical table, and store the change summary and change log together.

[0175] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0176] Figure 10 A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0177] like Figure 10 As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1002 or a computer program loaded into random access memory (RAM) 1003 from storage unit 1008. The RAM 1003 may also store various programs and data required for the operation of device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. Input / output (I / O) interface 1005 is also connected to bus 1004.

[0178] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of monitors, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0179] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as the logic model management method. For example, in some embodiments, the logic model management method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program may be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by the computing unit 1001, one or more steps of the logic model management method described above may be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured to execute a logic model management method by any other suitable means (e.g., by means of firmware).

[0180] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0181] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0182] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0183] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0184] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0185] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0186] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0187] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A logical model management method, comprising: In response to the detection of a change event targeting a physical table, a reverse engineering request is generated based on information from the change event; A revised draft of the logical model of the modified physical table is obtained based on the reverse engineering request, and the revised draft is obtained by reverse engineering based on the reverse engineering request. The large language model is invoked to generate an audit report for the revised draft based on at least one of the preset compliance rules and data standards. In response to the confirmation of the revised draft based on the audit report, the revised draft is published.

2. The method according to claim 1, wherein, The invoked large language model, based on at least one of preset compliance rules and data standards, generates an audit report for the revised draft, including: The first major language model is invoked, and a pre-review result for the revised draft is generated based on at least one of the preset compliance rules and data standards; The second language model is invoked, and an audit report for the revised draft is generated based on the pre-audit results and the change column information. The change column information is obtained based on the differences between the revised draft and the main version of the logical model of the changed physical table that has been released.

3. The method according to claim 2, wherein, The pre-audit results include compliance pre-inspection results. The step of calling the first major language model, based on at least one of preset compliance rules and data standards, to generate pre-audit results for the revised draft includes: The first large language model is invoked to generate a compliance pre-inspection result based on the change column information and the compliance rules; The compliance pre-inspection results include at least one of the following: The change column information includes the sensitive column identification results of each change column; The column name standardization verification results for each changed column; The physical type conversion security assessment results for each of the change columns.

4. The method according to claim 3, wherein, The step of calling the first large language model, based on the change column information and the compliance rules, to generate a compliance pre-inspection result includes: Based on the changed column information, the table name of the changed physical table, the logical model type of the logical model of the changed physical table, and the compliance rules, a first prompt instruction is generated. The first prompt instruction is used to prompt the first large language model. Based on the changed column information, the table name of the changed physical table, the logical model type of the logical model of the changed physical table, and the compliance rules, a compliance pre-inspection result is generated. The first prompt instruction is provided to the first large language model to obtain the compliance pre-inspection result.

5. The method according to claim 2, wherein, The pre-audit results include data standard matching results. The step of calling the first major language model, based on at least one of preset compliance rules and data standards, to generate pre-audit results for the revised draft includes: Identify the column names of columns in the revised draft that do not match the data standard; The first large language model is invoked, and a data standard matching result is generated based on the column names of columns that do not match the data standard in the revised draft, the column names of columns that match the data standard in the main version, and the data standard corresponding to the columns that match the data standard in the main version. The data standard matching result includes the matching relationship between columns in the revised draft that do not match the data standard and columns in the main version that do match the data standard.

6. The method according to claim 5, wherein, The step of calling the first large language model, based on the column names of columns that do not match the data standard in the revised draft, the column names of columns that match the data standard in the main version, and the data standard corresponding to the columns that match the data standard in the main version, generates a data standard matching result, including: Based on the column names of columns that do not match the data standard in the revised draft, the physical types corresponding to the columns that do not match the data standard in the revised draft, the column names of columns that match the data standard in the main version, and the data standard and physical type corresponding to the columns that match the data standard in the main version, a second prompt instruction is generated. The second prompt instruction is used to instruct the first large language model to generate the data standard matching result based on the column names of columns that do not match the data standard in the revised draft, the physical types corresponding to the columns that do not match the data standard in the revised draft, the column names of columns that match the data standard in the main version, and the data standard and physical type corresponding to the columns that match the data standard in the main version. The second prompt instruction is provided to the first large language model to obtain the data standard matching result.

7. The method according to claim 2, before calling the second major language model and generating an audit report for the revised draft based on the pre-audit results and change column information, further comprising: Based on the revised draft, the column information of each column in the changed physical table is determined, and the column information includes column name, physical type and first comment; The third language model is invoked to determine the logical type of each column and the second annotation of each column based on the column information, the table name of the changed physical table, and the logical model type of the logical model of the changed physical table. The second annotation is obtained by completing the first annotation. The revised draft is updated based on the logical type of each column and the second comment of each column.

8. The method according to claim 2, wherein, The process of calling the second major language model, based on the pre-review results and change column information, generates a review report for the revised draft, including: Based on the change column information, downstream dependency information, and the pre-audit result, a third prompt instruction is generated. The third prompt instruction is used to prompt the second language model to generate an audit report based on the change column information, downstream dependency information, and the pre-audit result. The audit report includes at least a change column summary, a downstream dependency impact analysis, and the pre-audit result. The third prompt instruction is provided to the second large language model to obtain the audit report.

9. The method according to claim 8, wherein, The third prompt instruction is generated based on the change column information, downstream dependency information, and the pre-audit result, including: In response to a mismatch between a column in the draft revision that does not match the data standard and a column in the main version that matches the data standard, a candidate data standard is obtained from the data standard library. The physical type of the candidate data standard is the same as the physical type of the column in the draft revision that does not match the data standard, and the business domain to which the candidate data standard belongs is the same as the business domain to which the draft revision belongs. Based on the changed column information, downstream dependency information, candidate data standards, and pre-audit results, a third prompt instruction is generated.

10. The method according to any one of claims 1 to 9, wherein, The change event includes at least one of the following: Obtain metadata change events through the metadata operation interface of the first type of data source; Obtain the audit log file of the second type of data source, and extract the events of data definition language (DDL) change statements in the audit log file.

11. The method according to claim 10, wherein, The second type of data source is a database. The step of generating a reverse engineering request based on the information from the change event includes: Based on the DDL change statement, determine the first location information of the physical table to be changed; A reverse engineering request is generated based on the data source identifier corresponding to the audit log file and the first location information of the changed physical table.

12. The method according to any one of claims 1 to 9, wherein, The revised draft of the logical model for obtaining the changed physical table based on the reverse engineering request includes: In response to the fact that the lifecycle state of the logical model of the changed physical table is that there is no published major version and no revision draft, the revision draft of the logical model of the changed physical table is obtained based on the reverse engineering request, and the obtained revision draft is determined as the target revision draft of the logical model of the changed physical table. In response to the fact that the lifecycle state of the logical model of the changed physical table is that there is a published major version but no revision draft, the revision draft of the logical model of the changed physical table is obtained based on the reverse engineering request, and the obtained revision draft is determined as the target revision draft of the logical model of the changed physical table. In response to the fact that the logical model of the changed physical table is in a lifecycle state where there is a published major version and a revision draft, the revision draft of the logical model of the changed physical table is obtained based on the reverse engineering request, and the obtained revision draft is used to overwrite the existing revision draft. In response to the fact that the logical model of the changed physical table is in a lifecycle state where there is no published major version but there is a revision draft, the revision draft of the logical model of the changed physical table is obtained based on the reverse engineering request, and the obtained revision draft is used to overwrite the existing revision draft. In response to the fact that the logical model of the modified physical table is in an obsolete state in its lifecycle, the step of obtaining the revised draft of the logical model of the modified physical table based on the reverse engineering request is not executed.

13. The method according to any one of claims 1 to 9, wherein, The publication of the revised draft includes: Obtain and store the materialization configuration information and materialization status of the published main version of the logical model of the modified physical table; Delete the aforementioned major release; The revised draft was designated as the new master version; Based on the released major version and the new major version, generate a change log; Based on the reference relationship information of the already released main version, update the reference relationship information of the new main version.

14. The method of claim 13, further comprising: The source markers of the revised draft are determined, including a first marker and a second marker. The first marker is used to characterize that the revised draft is obtained in response to the detection of a change event for a changed physical table, and the second marker is used to characterize that the revised draft is obtained in response to an event that triggers a manual editing logic model. The publication of the revised draft also includes: In response to the source being marked as the first mark, the materialization state of the logical model of the changed physical table is set to synchronized; in response to the source being marked as the second mark, the materialization state of the logical model of the changed physical table is set to synchronized. Following the publication of the revised draft, the following is also included: In response to the materialization state of the logical model of the changed physical table being required to be synchronized, the changed physical table is updated.

15. The method of claim 13, further comprising: The fourth language model is invoked to generate a change summary based on the change log and the table name of the changed physical table, and the change summary and the change log are stored together.

16. A logic model management device, comprising: An event response unit is configured to generate a reverse engineering request based on information from the detected change event for a modified physical table. The draft acquisition unit is configured to acquire a revised draft of the logical model of the changed physical table based on the reverse engineering request, wherein the revised draft is obtained by reverse engineering based on the reverse engineering request; The report generation unit is configured to invoke a large language model and generate an audit report for the revised draft based on at least one of preset compliance rules and data standards. The draft publishing unit is configured to publish the revised draft in response to the confirmation operation of the revised draft based on the review report.

17. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-15.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-15.

19. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-15.