Data warehouse intelligent construction method and device based on large model
By introducing large-scale model technology into the data warehouse, building multi-level task processing modules and parallel processing mechanisms, the entire data processing process is automated and intelligent, solving the problem of lack of automation and intelligence in the data processing process in the existing technology, and significantly improving processing efficiency and data quality.
Patent Information
- Application Number
- CN202510106818.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-23
AI Technical Summary
The existing technology lacks full-process automation and intelligent processing capabilities in the data processing process, resulting in cumbersome, complex and manual dependence in data collection, classification, standardization and integration.
Using the intelligent construction method of data warehouse based on large models, by building a multi-level task processing module and parallel processing mechanism, the large model is introduced for table feature analysis, parameter recommendation, field feature recognition, cleaning and transformation suggestions generation, intelligent field mapping and rule generation, to realize the automation and intelligence of the entire data processing process.
It significantly improves data processing efficiency and accuracy, reduces manual intervention costs, and realizes that the automation rate of the entire data processing process has been increased to more than 90%, the processing efficiency has been improved by 5-10 times, the data quality compliance rate has been increased to 90%, and the manual intervention has been reduced by 80%.
Smart Images

Figure CN120045545A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data processing, and particularly relates to a method and device for intelligently constructing a data warehouse based on a large model, which are used to realize intelligent collection, processing, and management of data. The present invention is particularly suitable for intelligent automatic processing of data in scenarios such as enterprise data middle platforms and data warehouses. Background Art
[0002] With the deepening of enterprise digital transformation, enterprises are facing the need to process a large amount of heterogeneous data. Although there are some data processing tools and platforms in the prior art, these solutions often only focus on a specific link and lack overall control and intelligent coordination of the entire data processing process, resulting in many problems in the data processing process: In the data collection stage, since it is necessary to manually judge whether a table needs to be collected, determine the collection period, select incremental fields, etc., the process of configuring data collection parameters is cumbersome and complex; In the data classification and standardization link, since the recognition of data table types highly depends on manual experience, and due to factors such as inaccurate understanding of field semantics and the need for professional knowledge for standard mapping, it is difficult to accurately establish data association relationships; In the data fusion stage, a large amount of manual participation is required for field mapping, and the fusion rules lack intelligent features, making it difficult to ensure data consistency and fusion effects.
[0003] More critically, the prior art generally lacks intelligent processing capabilities, cannot automatically identify data characteristics, lacks intelligent decision-making capabilities, and the processing rules are too rigid to adapt to the changing needs of different scenarios.
[0004] Therefore, there is an urgent need for a solution that can realize the automation of the entire data processing process. Summary of the Invention
[0005] In order to solve the problems of automation and intelligence in the data processing process, the object of the present invention is to provide a method for intelligently constructing a data warehouse based on a large model, which realizes the full process automation and intelligence of data from collection and processing in the business system to the theme layer of the data warehouse.
[0006] Another object of the present invention is to provide a device for realizing the above-mentioned method for intelligently constructing a data warehouse based on a large model. By constructing a multi-level task processing module and a parallel processing mechanism, a large model is introduced in the data collection link for table feature analysis and parameter recommendation, field feature recognition and cleaning and transformation suggestion generation are realized in the data exploration link, business rules are automatically bound and data standards are associated in the data transformation link, and intelligent field mapping and rule generation are realized in the data fusion link.
[0007] The object of the present invention and the technical problems to be solved are achieved by the following technical solutions. A method for intelligent construction of a data warehouse based on a large model proposed according to the present invention includes: users register the data source information of the business system to be collected through the data source registration entry and submit it as a data processing task. The data source information includes IP, port, username, password, and database, and it is submitted as a data processing task; the task executor receives the data processing task, obtains the data table information from the registered data source, calls the large model to perform feature analysis on the data table, and identifies the table type, business attributes, and collection parameters; performs data collection according to the analysis results, where the analysis results include identifying whether the data table is a temporary table, determining the data collection period, and selecting incremental fields; and performs intelligent exploration on the collected data, identifies field features and data quality, and generates exploration results, where the field features are field types, including numeric type, time type, and text type; performs data conversion based on the exploration results, automatically binds business rules, associates data standards, and generates conversion rules, submits and runs the data cleaning and conversion job; based on the relevant information of the standardized table and the target table in the subject layer, constructs a prompt text, inputs the prompt text into the large model, intelligently matches the target table and field mapping relationship, generates a fusion rule, and executes the fusion.
[0008] Furthermore, the device for the method of intelligent construction of a data warehouse based on a large model includes: a data source registration module, a change perception module, a large model intelligent analysis and decision-making module, a data collection module, a data exploration module, a data cleaning and conversion module, and a data multi-dimensional fusion module; the data source registration module is used to manage the data source connection information, monitor the changes in the data source status and metadata, and includes a connection configuration unit, a status monitoring unit, and a metadata change monitoring unit.
[0009] Furthermore, for the device of the method of intelligent construction of a data warehouse based on a large model, the connection configuration unit is used for the full life cycle management of the data source connection, obtains the text content input by the user, and passes the text content into the large model intelligent analysis and decision-making module, automatically analyzes the parameters required for the connection, completes the configuration and management of the data source, and uses a high-strength encryption algorithm to ensure the secure storage of the connection information. The text content includes IP, port, username, and password.
[0010] This unit supports the connection configuration of various types of data sources such as MySQL, Oracle, and PostgreSQL. Through a text input box, the user inputs the text containing the necessary connection information, including key information such as IP, port, username, password, and database, and passes the text content into the large model intelligent analysis and decision-making module, automatically analyzes the parameters required for the connection, completes the configuration and management of the data source, and uses a high-strength encryption algorithm to ensure the secure storage of the connection information.
[0011] The connection configuration unit is also used to implement network connectivity tests, database response time tests, permission verification tests, etc., and through the collection and analysis of performance metrics, timely detect and warn of potential connection problems; this unit also provides a full range of connection test verification functions, including network connectivity tests, database response time tests, permission verification tests, etc., and through the collection and analysis of performance metrics, timely detect and warn of potential connection problems to ensure the stability and reliability of the data source connection.
[0012] Furthermore, for the device of the intelligent construction method of the data warehouse based on the large model, the status monitoring unit is used to collect and track various operation metrics of the data source in real time, including core metrics such as connection response time, data throughput, error rate, as well as performance metrics at the system level such as CPU usage, memory occupancy, and IO performance; these metrics are pushed to the change perception module, and the parallelism of data collection is intelligently adjusted according to various operation metrics to reduce the impact on the normal operation of the business system; the metadata monitoring unit is used to obtain all-round metadata such as the table structure, field attributes, index information, and statistical information of the data source in real time through an automated real-time monitoring mechanism, and push the changed metadata information to the data change perception module to realize the configuration of data collection, data conversion, data fusion and other processing links to be adjusted in time after the metadata of the business system changes.
[0013] Furthermore, for the device of the intelligent construction method of the data warehouse based on the large model, the change perception module is used to listen to the status change event information of each module, track and distribute the status events, including a change listening unit, a status event tracking unit, and a status event distribution unit.
[0014] 1) The change listening unit is used to capture structure changes in the data source, changes in data source operation metrics, and changes in platform service status in real time through an event-driven listening mechanism, and implements a rule-based change filtering and aggregation mechanism. Through row-level refined tracking and intelligent event routing, it ensures that change events can be processed in a timely manner according to business importance.
[0015] 2) The status event tracking unit is used to track the entire process of status events from creation, distribution to completion in real time, and persistently store the status event information to support multi-dimensional status event query and statistical analysis.
[0016] 3) The status event distribution unit is used to distribute to the corresponding various modules according to the type of status event.
[0017] Furthermore, for the apparatus of the intelligent construction method of the data warehouse based on the large model, the large model intelligent analysis and decision-making module is used to receive the prompt words and large model call requests of each module and return the analysis results of the large model, including a prompt word construction unit, a large model call unit, a call result parsing unit, and an exception reflection unit;
[0018] 1) The prompt word construction unit is used for the intelligent construction of prompt words for data processing scenarios. Through a multi-dimensional information collection mechanism, it automatically obtains context information, including data table structures, field attributes, business rules, as well as domain knowledge and experience knowledge, and based on a scenario-based prompt word template library, fills and optimizes dynamic parameters in combination with specific task requirements.
[0019] 2) The large model call unit is used to send the prompt words of each data processing task type to the large model, and the large model returns results in a specified format, such as JSON, according to the context requirements of the prompt words.
[0020] 3) The call result parsing unit is used to convert the unstructured text returned by the model into a standard JSON or other structured format through a formatting engine, and ensure the legality and integrity of the results through multi-dimensional verification rules.
[0021] 4) The exception reflection unit is used to capture various exceptions of each task during the process of automatic collection and processing, conduct in-depth cause analysis in combination with task context information, and based on the self-optimization process of the reflection mechanism, automatically optimize the prompt word construction strategy and call parameter configuration by analyzing exception patterns and influencing factors.
[0022] Furthermore, for the apparatus of the intelligent construction method of the data warehouse based on the large model, the data collection module is used to collect parameters of the data tables in the business system through large model analysis, and complete the data collection through collection jobs to form original data tables, including a collection prompt word generation unit, a table feature analysis unit, a collection parameter configuration unit, a collection job generation unit, and a collection status monitoring unit.
[0023] 1) The collection prompt word generation unit is used to integrate the data table name, description, field information, sampled data content to be collected with the data collection parameter recommendation template, submit it to the intelligent analysis and decision-making module, and return the table feature analysis results through the large model.
[0024] 2) The table feature parsing unit is used to parse the table feature analysis results returned by the large model, including whether it is a temporary table, the type of data table, the data collection method, the incremental collection time field, the encoding identification field, the name field, the data table types include entity data, transaction data, and dimension data, and the data collection methods include full volume collection and incremental collection.
[0025] (3) The data collection job generation unit is used to automatically generate a complete data collection job configuration based on the table feature parsing results, including data source connection parameters, reading strategies, and concurrency levels, and implement intelligent allocation and dynamic adjustment of computing resources through the resource scheduling engine.
[0026] (4) The data collection status monitoring unit is used to detect the running status of the data collection job, as well as data integrity, accuracy, and consistency, and push the status information to the change perception module.
[0027] (5) The data exploration module
[0028] is used to generate an exploration result report for the collected raw data through statistical analysis algorithms, including a data distribution analysis unit, an exploration report generation unit, and a cleaning and transformation suggestion generation unit.
[0029] 1) The data distribution analysis unit is used to extract multi-dimensional features of the data through statistical analysis algorithms, including analyzing statistical metrics for numerical fields, where the statistical metrics include maximum value, minimum value, mean, median, and standard deviation; identifying periodic features for time fields, where the periodic features include time span, update frequency, and change pattern; and deeply analyzing the null value distribution of text fields to obtain the null value ratio.
[0030] 2) The exploration report generation unit is used to automatically generate an exploration report for the collected data table, and the exploration report includes data overview, quality score, and problem distribution.
[0031] 3) The cleaning and transformation suggestion generation unit is used to generate data transformation suggestions through a large model based on the characteristics of the field types in the exploration results, and the field types in the exploration results include text fields, time fields, and category fields.
[0032] (6) The data cleaning and transformation module
[0033] is used to complete the cleaning and transformation of the collected raw data according to the data exploration results, recommend business rules and transformation rules through a large model, and form a standardized result table, including a business rule binding unit, a data standard association unit, a transformation rule generation unit, and a transformation job execution unit.
[0034] 1) Business rule binding unit, which is used to construct a prompt text by combining field meanings with a rule knowledge base, submit it to a large model for intelligent matching, and return whether the field matches the business rule. If so, it automatically establishes the association relationship between the data field and the business rule, and evaluates the applicability and influence scope of the rule through a rule conflict detection mechanism and verification means to ensure the accuracy and effectiveness of rule execution. The rule knowledge base consists of rules written by business experts and requirements analysts based on domain business knowledge, including rule name, rule description, rule type, and regular expression.
[0035] 2) Data standard association unit, which is used to construct a prompt text by combining category field meanings with a data standard library, submit it to a large model for intelligent matching, and return whether the field matches the data standard. If so, it automatically establishes the association relationship between the data field and the data standard, and performs conversion according to the naming of the data standard and the standard dictionary code; the data standard library includes international standards, ministerial standards, and industry standards, which are periodically obtained from public sites or documents, and are solidified into the data standard library through manual entry or automated program parsing.
[0036] 3) Conversion rule generation unit, which is used to generate appropriate conversion rules based on the mapping relationship with the built-in conversion rule template according to the field type of the exploration result, such as removing special characters from the name field, converting the certificate code field to uppercase uniformly, and standardizing the time field format; the built-in conversion rule template is solidified into the program during the development stage and can quickly generate conversion rules according to the field during the running stage.
[0037] 4) Conversion job execution unit, which is used to create and run conversion jobs according to business rules and conversion rules to implement the processing of data from the collected state to the converted state, and form standardized result data.
[0038] 5) Conversion status monitoring unit, which is used to detect the running status of the conversion job and push the status information to the change perception module.
[0039] (7) Data multi-dimensional fusion module
[0040] It is used to complete data fusion by identifying the target tables of the theme layer to be fused from the standardized result data output by the data cleaning and conversion module, including a target table matching unit, a field mapping unit, a fusion rule configuration unit, and a fusion execution unit.
[0041] 1) Target table matching unit, which is used to construct a prompt text from the standardized table and the relevant information of the theme layer target, submit it to the intelligent analysis and decision-making module, and return the table feature analysis result through the large model; the prompt text is matched with each data table in the theme layer for semantic and structural similarity recognition, and the recognition result is returned. The returned recognition result includes whether the target table matches consistently, the field matching relationship, and the confidence level.
[0042] 2) The fusion rule configuration unit is used to generate a fusion rule based on the result of the target table matching unit and configure it on the fields of the data table in the theme layer.
[0043] 3) The fusion execution unit is used to generate a data fusion operation according to the fusion rule, execute the fusion operation, and realize the fusion of the standardized result data output by the data cleaning and transformation module into the target table in the theme layer.
[0044] 4) The fusion status monitoring unit is used to detect the running status of the fusion operation and push the status information to the change perception module.
[0045] Based on the above technical solutions, the device and method of the intelligent construction method of the data warehouse based on the large model of the present invention at least have the following advantages:
[0046] 1) The present invention realizes the intelligence of the whole data processing process by introducing the large model technology, significantly improves the data processing efficiency and accuracy, and reduces the cost of manual intervention;
[0047] 2) Verified by actual application, the present invention improves the automation rate of the whole data processing process to more than 90%, the processing efficiency is increased by 5-10 times, the data quality compliance rate is increased to 90%, the manual intervention is reduced by 80%, and a large amount of labor cost for data processing can be saved every year;
[0048] 3) At the same time, the system has the adaptability to mainstream data sources and the ability to handle more than 90% of abnormal scenarios, providing an efficient, accurate and economical solution for the digital transformation of enterprises. Description of the Drawings
[0049] Figure 1 Shows the overall structure diagram of the intelligent construction device of the data warehouse based on the large model of the present invention;
[0050] Figure 2 Shows the overall flow chart of the intelligent construction method of the data warehouse based on the large model of the present invention;
[0051] Figure 3 Shows an example diagram of the data table to be collected in the human resources business system of an enterprise. Detailed Embodiment
[0052] To make the purpose, technical solutions and advantages of the present invention clearer, the following examples are given in conjunction with the drawings to further describe the technical solutions provided by the present invention in detail. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0053] The present invention provides a method and apparatus for intelligent construction of a data warehouse based on a large model, belonging to the technical field of big data processing. The method includes: the user registers the data source information of the business system to be collected through the data source registration entry, and the data source information includes IP, port, username, password, database, etc., and the data source information is submitted as a data processing task; the task executor receives the data processing task, obtains the data table information from the registered data source, calls the large model to perform feature analysis on the data table, and identifies the table type, business attributes, and collection parameters; performs data collection according to the analysis results, where the analysis results include identifying whether the data table is a temporary table, determining the data collection period, and selecting incremental fields; and performs intelligent exploration on the collected data, identifies field features and data quality, and generates exploration results, where the field features are field types, including numeric type, time type, and text type; performs data conversion based on the exploration results, automatically binds business rules, associates data standards, and generates conversion rules, submits and runs a data cleaning and conversion job; based on the relevant information of the standardized table and the target table in the subject layer, constructs a prompt text, inputs the prompt text into the large model, intelligently matches the target table and field mapping relationship, generates a fusion rule and executes the fusion, where the prompt text includes the standardized table name, table description, field information, domain knowledge, and experience knowledge, as well as the relevant information of the target table in the subject layer, and performs semantic and structural similarity matching recognition between the prompt text and each data table in the subject layer, and returns the recognition results, and the returned recognition results include whether the target table matches, the field matching relationship, and the confidence level. If they match, a fusion rule is generated and configured on the fields of the data table in the subject layer, a data fusion job is generated according to the fusion rule, and the fusion job is executed to fuse the standardized result data output by the data cleaning and conversion module into the target table in the subject layer.
[0054] Based on the above method, the present invention provides an intelligent construction device for a data warehouse based on a large model, as Figure 1 shown, which mainly includes a data source registration module, a change perception module, a large model intelligent analysis and decision-making module, a data collection module, a data exploration module, a data cleaning and conversion module, and a data multi-dimensional fusion module.
[0055] The data source registration module 101 is used to manage the data source connection information, monitor the changes in the data source status and metadata, and includes a connection configuration unit, a status monitoring unit, and a metadata change monitoring unit.
[0056] The connection configuration unit is mainly responsible for the full - life - cycle management of data source connections. This unit supports the connection configuration of various types of data sources such as MySQL, Oracle, and PostgreSQL. Through a text input box, users input the text necessary for the connection, including key information such as IP, port, username, password, and database, and transmit the text content to the large - model intelligent analysis and decision - making module. The module automatically analyzes the parameters required for the connection, completes the configuration and management of the data source, and uses high - strength encryption algorithms to ensure the secure storage of connection information. At the same time, this unit also provides a comprehensive connection test and verification function, including network connectivity test, database response time test, permission verification test, etc., and through the collection and analysis of performance indicators, it can timely detect and warn potential connection problems to ensure the stability and reliability of the data source connection.
[0057] The status monitoring unit is used to collect and track various operation indicators of the data source in real - time, including core indicators such as connection response time, data throughput, error rate, and performance indicators at the system level such as CPU usage, memory occupancy, and IO performance. These indicators are pushed to the change perception module, which intelligently adjusts the parallelism of data collection according to various operation indicators to reduce the impact on the normal operation of the business system.
[0058] The metadata monitoring unit is used to obtain all - round metadata of the data source such as table structure, field attributes, index information, and statistical information in real - time through an automated real - time monitoring mechanism, and push the changed metadata information to the data change perception module. After the metadata of the business system changes, it can timely adjust the configurations of data collection, data conversion, data fusion, and other processing links.
[0059] The change perception module 102 is used to listen to the status change event information of each module, track and distribute the status events, including a change listening unit, a status event tracking unit, and a status event distribution unit.
[0060] 1) The change listening unit is used to capture in real - time the structural changes (such as table creation, field change, index modification) in the data source, the changes in the data source operation indicators, and the changes in the platform service status through an event - driven listening mechanism. It has implemented a rule - based change filtering and aggregation mechanism. Through row - level fine - grained tracking and intelligent event routing, it ensures that change events can be processed in a timely manner according to business importance.
[0061] 2) The status event tracking unit mainly tracks the whole process of status events from creation, distribution to completion in real - time, and persists the status event information for storage, supporting multi - dimensional query and statistical analysis of status events.
[0062] 3) The status event distribution unit is used to distribute status events to the corresponding various modules according to the type of status events.
[0063] The large model intelligent analysis and decision-making module 103 is used to receive the prompt words and large model call requests from each module and return the analysis results of the large model, including a prompt word construction unit, a large model call unit, a call result parsing unit, and an exception reflection unit.
[0064] 1) The prompt word construction unit is used for the intelligent construction of prompt words for data processing scenarios. Through a multi-dimensional information collection mechanism, it automatically obtains context information such as data table structures, field attributes, and business rules, as well as domain knowledge and experience knowledge. Based on a scenario-based prompt word template library (including various templates such as data collection parameter recommendations and field mapping relationship identifications), it dynamically fills and optimizes parameters in combination with specific task requirements.
[0065] 2) The large model call unit is used to send the prompt words of each data processing task type to the large model, and the large model returns results in a specified format, such as JSON, according to the requirements of the prompt word context.
[0066] 3) The call result parsing unit is used to convert the unstructured text returned by the model into a standard JSON or other structured format through a formatting engine, and ensure the legality and integrity of the results through multi-dimensional verification rules.
[0067] 4) The exception reflection unit is used to capture various exceptions of each task during the process of automatic collection and processing (such as call timeouts, result exceptions, parsing failures, collection task exceptions, etc.), conduct in-depth cause analysis in combination with task context information, and based on the self-optimization process of the reflection mechanism, automatically optimize the prompt word construction strategy and call parameter configuration by analyzing exception patterns and influencing factors. The exception reflection unit is used to maintain a knowledge base of manual processing opinions, continuously accumulate and summarize exception handling experience, and through the way of regularly fine-tuning prompt words, realize the automatic adjustment and optimization of exception handling strategies, and continuously improve the system's exception handling ability and service stability.
[0068] The data collection module 104 is used to analyze the data table collection parameters of the business system through the large model, complete the data collection through the collection job, and form an original data table, including a collection prompt word generation unit, a table feature analysis unit, a collection parameter configuration unit, a collection job generation unit, and a collection status monitoring unit.
[0069] 1) The collection prompt word generation unit is used to integrate the data table name, description, field information, and sampled data content to be collected with the data collection parameter recommendation template, submit it to the intelligent analysis and decision-making module, and return the table feature analysis result through the large model.
[0070] 2) The table feature parsing unit is used to parse the table feature analysis results returned by the large model, including whether it is a temporary table, the type of data table (entity data, transaction data, dimension data), the data collection method (full - volume collection, incremental collection), the incremental collection time field, the coding identification field, the name field, etc.
[0071] 3) The collection job generation unit is used to automatically generate a complete collection job configuration according to the table feature parsing results, including data source connection parameters, reading strategies, concurrency degrees, etc., and realizes the intelligent allocation and dynamic adjustment of computing resources through the resource scheduling engine.
[0072] 4) The collection status monitoring unit is used to detect the running status of the collection job, as well as indicators such as data integrity, accuracy, and consistency, and push the status information to the change perception module.
[0073] The data exploration module 105 is used to generate an exploration result report for the collected raw data through statistical analysis algorithms, including a data distribution analysis unit, an exploration report generation unit, and a cleaning and transformation suggestion generation unit.
[0074] 1) The data distribution analysis unit is used to extract multi - dimensional features of the data through statistical analysis algorithms, including statistical index analysis of numerical data (maximum value, minimum value, mean, median, standard deviation, etc.), periodic feature recognition of time - type data (time span, update frequency, change rule, etc.), and in - depth analysis of null value distribution (null value ratio, pattern recognition, cause analysis, etc.).
[0075] 2) The exploration report generation unit is used to automatically generate an exploration report for the collected data tables. The exploration report covers multi - dimensional statistical analyses such as data overview, quality score, and problem distribution.
[0076] 3) The cleaning and transformation suggestion generation unit is used to generate data transformation suggestions through the large model according to the characteristics of the field types (text fields, time fields, category fields) of the exploration results.
[0077] The data cleaning and transformation module 106 is used to complete the cleaning and transformation of the collected raw data according to the data exploration results, through the business rules and transformation rules recommended by the large model, and form a standardized result table, including a business rule binding unit, a data standard association unit, a transformation rule generation unit, and a transformation job execution unit.
[0078] 1) The business rule binding unit is used to construct a prompt text by combining the field meaning and the rule knowledge base, submit it to the large model for intelligent matching, and return whether the field and the business rule match. If so, it automatically establishes the association relationship between the data field and the business rule, and evaluates the applicability and influence scope of the rule through the rule conflict detection mechanism and verification means to ensure the accuracy and effectiveness of rule execution. The rule knowledge base consists of rules written by business experts and requirements analysts based on domain business knowledge, including rule name, rule description, rule type, and regular expression, and evaluates the applicability and influence scope of the rule through the rule conflict detection mechanism and verification means to ensure the accuracy and effectiveness of rule execution.
[0079] 2) The data standard association unit is used to construct a prompt text by combining the category field meaning and the data standard library, submit it to the large model for intelligent matching, and return whether the field and the data standard match. If so, it automatically establishes the association relationship between the data field and the data standard, and performs conversion according to the naming of the data standard and the standard dictionary code; the data standard library includes international standards, ministerial standards, and industry standards, which are periodically obtained from public sites or documents, and are solidified into the data standard library through manual input or automated program parsing.
[0080] 3) The conversion rule generation unit is used to generate appropriate conversion rules based on the mapping relationship with the built-in conversion rule template by exploring the field type of the exploration result, such as removing special characters from the name field, converting the certificate code field to uppercase uniformly, and standardizing the time field format.
[0081] 4) The conversion job execution unit is used to create and run conversion jobs according to business rules and conversion rules, and implement the processing of data from the collected state to the converted state to form standardized result data.
[0082] 5) The conversion status monitoring unit is used to detect the running status of the conversion job and push the status information to the change perception module.
[0083] The data multi-dimensional fusion module 107 is used to complete data fusion by identifying the target table of the theme layer to be fused through the large model for the standardized result data output by the data cleaning and conversion module, including a target table matching unit, a field mapping unit, a fusion rule configuration unit, and a fusion execution unit.
[0084] 1) The target table matching unit is used to construct a prompt text from the standardized table and the relevant information of the theme layer target, submit it to the intelligent analysis and decision-making module, and return the table feature analysis result through the large model; the prompt text is matched with each data table in the theme layer for semantic and structural similarity recognition, and the recognition result is returned. The returned recognition result includes whether the target table matches consistently, the field matching relationship, and the confidence level.
[0085] 2) The fusion rule configuration unit is used to generate a fusion rule based on the result of the target table matching unit and configure it on the fields of the data table in the theme layer.
[0086] 3) The fusion execution unit is used to generate a data fusion job according to the fusion rule, execute the fusion job, and integrate the standardized result data output by the data cleaning and transformation module into the target table in the theme layer.
[0087] 4) The fusion status monitoring unit is used to detect the running status of the fusion job and push the status information to the change perception module.
[0088] Next, in combination with Figures 2-3 Figure 1-3, the main implementation principles, specific implementation methods, and corresponding beneficial effects of the technical solutions in the embodiments of the present application will be elaborated in detail.
[0089] Example 1:
[0090] In this embodiment, taking the construction of a data warehouse for data collection, processing, and fusion in a certain enterprise's human resources business system as an example, the specific implementation process of the present invention will be described in detail.
[0091] The main method flow is as Figure 2 shown.
[0092] Step 201, intelligent registration of data sources. The user registers the data source information of the business system to be collected through the data source registration entry, including IP, port, username, password, database, etc., and submits it as a data processing task.
[0093] The registration information provided in this embodiment is shown in Table 1 below:
[0094] Table 1 Data source information of the business system to be collected
[0095]
[0096]
[0097] According to the provided registration information, combined with domain knowledge and experience knowledge to form a prompt, and submit it to the large model intelligent analysis and decision-making module 103, return the recognized JSON result data structure, and the data source registration module 101 automatically calls the data source registration API to complete the data source registration, avoiding the operation of manually inputting each item in the form of traditional data source registration.
[0098] Step 202, intelligent data collection. Obtain data table information from registered data sources, use a large model to analyze the characteristics of the data table, and identify the table type, business attributes, and collection parameters. Perform data collection according to the analysis results, including identifying whether it is a temporary table, determining the collection period, selecting incremental fields, etc., and submit and run the data collection job. Open-source large models are used for private deployment, such as QWen 2.5 and Llama 3.1. Commercial large models such as ChatGPT, DeepSeek, and Gemini can also be used. They all provide standard unified APIs that can receive prompt text and return the table feature analysis results. The analysis results include whether it is a temporary table, the type of data table (entity data, transaction data, dimension data), the data collection method (full-volume collection, incremental collection), the incremental collection time field, the coding identification field, the name field, etc.
[0099] In this embodiment, the metadata information of the employee entry registration form, employee file form, and gender form will be obtained, as Figure 3 shown.
[0100] Then, the metadata information of each table is combined with domain knowledge and experience knowledge into a prompt, which is submitted to the large model intelligent analysis and decision-making module 103, and the identified JSON result data structure is returned. Taking the employee entry registration form as an example,
[0101] The prompt is shown in Table 2 below:
[0102] Table 2 Prompts for analyzing the employee entry registration form to be collected
[0103]
[0104]
[0105] The return results are shown in Table 3 below:
[0106] Table 3 Identification results of the employee entry registration form to be collected
[0107]
[0108] The data collection module 104 automatically calls the collection job creation API to complete the creation and execution of the collection job, avoiding the manual analysis and input operations of traditional configuration of collection jobs. The change perception module 102 perceives the job running status in real time. When the job runs successfully, it enters Step 203.
[0109] Step 203, intelligent data exploration. The data exploration module 105 performs intelligent exploration on the collected data to identify field characteristics and data quality. According to the characteristics of text fields, time fields, and category fields, generate data conversion suggestions through a large model.
[0110] In this embodiment, taking the employee onboarding registration form as an example, data conversion suggestions are generated through statistical analysis algorithms and large models, and the results are shown in Table 4 below:
[0111] Table 4 Exploration Results and Conversion Suggestions for Employee Onboarding Registration Form
[0112]
[0113]
[0114] The change perception module 102 perceives the running status of the exploration job in real time. After the job runs successfully, it enters step 204.
[0115] Step 204, intelligent data cleaning and conversion. Based on the exploration results, data conversion is performed, including automatically binding business rules, associating data standards, and generating conversion rules, and submitting and running the data cleaning and conversion job. By constructing the field meaning and the rule knowledge base into a prompt text and submitting it to the large model for intelligent matching, it is returned whether the field matches the business rule. If so, the association relationship between the data field and the business rule (including format rules, value rules, logical rules, etc.) is automatically established, and through the rule conflict detection mechanism and verification means, the applicability and influence range of the rule are evaluated to ensure the accuracy and effectiveness of rule execution. The rule knowledge base is composed of rules written by business experts and requirements analysts based on domain business knowledge, including rule names, rule descriptions, rule types, regular expressions, etc.
[0116] Then, by constructing the category field meaning and the data standard library into a prompt text and submitting it to the large model for intelligent matching, it is returned whether the field matches the data standard. If so, the association relationship between the data field and the data standard is automatically established, and conversion is performed according to the naming of the data standard and the standard dictionary code. The data standard library includes international standards, ministerial standards, and industry standards, which are periodically obtained from public sites or documents, and are solidified into the data standard library through manual entry or automated program parsing. Then, through the field type of the exploration results and the mapping relationship with the built-in conversion rule template, appropriate conversion rules are generated, such as removing special characters from the name field, converting the certificate code field to uppercase uniformly, and standardizing the time field format. The built-in conversion rule template is solidified into the program during the development stage and can quickly generate conversion rules according to the fields during the running stage.
[0117] As shown in Table 5, in this embodiment, taking the employee onboarding registration form as an example, according to the exploration results and conversion suggestions, the data cleaning and conversion module 106 automatically binds the recommended rules to the fields of employee name, mobile phone number, gender, department name, and onboarding time through the rule binding API, and creates and runs the data cleaning and conversion job. The change perception module 102 perceives the running status of the conversion job in real time. After the job runs successfully, it enters step 205.
[0118] Step 205, intelligent multi-dimensional fusion processing of data. Finally, data fusion is performed. By using the large model to intelligently match the target table in the theme layer and the field mapping relationship, a fusion rule is generated and the data fusion operation is executed. In this embodiment, taking the employee onboarding registration form as an example, the standardized result data output by the data cleaning and transformation module 106 is combined with domain knowledge and experience knowledge into a prompt. The prompt includes the table name, table description, and field information, and semantic analysis and structural similarity matching recognition are respectively performed with each data table in the theme layer. Submitted to the large model intelligent analysis and decision-making module 103, the returned results include information such as whether the target table matches, the field matching relationship, and the confidence level. The results are shown in Table 5 below:
[0119] Table 5 Mapping relationship between the human resources library test system (employee onboarding registration form) and the target table in the theme layer (employee master table)
[0120]
[0121] The data multi-dimensional fusion module 107 automatically calls the fusion operation creation API according to the recognition result, completes the creation and execution of the fusion operation, and avoids the manual analysis and input operations of the traditional configuration of the fusion operation. The change perception module 102 perceives the operation status of the job in real time. When the job runs successfully, it completes the intelligent automatic acquisition, processing, and fusion of data from the business system to the target table in the theme layer, realizing the full automation of the data processing process.
[0122] The above is only a preferred embodiment of the present invention, and does not impose any form of limitation on the present invention. Therefore, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A method for intelligently constructing a data warehouse based on a large model, characterized in that: The method comprises the following steps: Step S1: The user registers the data source information of the business system to be collected through the data source registration portal and submits it as a data processing task; Step S2: Receive data processing tasks through the task executor, obtain data table information from the registered data source, call the big model to perform feature analysis on the data table, and identify the table type, business attributes and collection parameters; Step S3: Execute data collection according to the analysis results, wherein the analysis results include identifying whether the data table is a temporary table, determining the data collection cycle, and selecting the incremental field; and intelligently explore the collected data, identify field characteristics and data quality, and generate exploration results, wherein the field characteristics are field types, including numeric type, time type, and text type; Step S4: Perform data conversion based on the exploration results, automatically bind business rules, associate data standards and generate conversion rules, submit and run data cleaning and conversion jobs; construct prompt word text based on the standardized table and the relevant information of the subject layer target table, input the prompt word text into the big model, intelligently match the target table and field mapping relationship, generate fusion rules and execute fusion; Among them, the prompt word text includes the standardized table name, table description, field information, domain knowledge and experience knowledge, and relevant information of the target table of the subject layer. The prompt word text is matched and identified with each data table of the subject layer for semantic and structural similarity, and the recognition result is returned. The returned recognition result includes whether the target table matches consistently, the field matching relationship, and the confidence level. If it matches, a fusion rule is generated and configured to the field of the subject layer data table. A data fusion job is generated according to the fusion rule, and the fusion job is executed to fuse the standardized result data output by the data cleaning conversion module into the target table of the subject layer.
2. The method for intelligently constructing a data warehouse based on a large model as claimed in claim 1, characterized in that: The data source information includes IP, port, user name, password, and database.
3. The method for intelligently constructing a data warehouse based on a large model as claimed in claim 1, characterized in that: The steps of intelligently exploring the collected data to identify field characteristics and data quality include the following steps: Analyze the statistical indicators of numeric fields, including maximum value, minimum value, mean, median, and standard deviation; Identify the periodic features of the time type field, where the periodic features include time span, update frequency, and change pattern; Perform in-depth analysis on the null value distribution of text fields to obtain the null value ratio.
4. The method for intelligently constructing a data warehouse based on a large model as claimed in claim 1, characterized in that: The steps are to convert data based on the exploration results, automatically bind business rules, associate data standards and generate conversion rules, submit and run data cleaning and conversion jobs, including the following steps: The field meaning and rule knowledge base are constructed into prompt word text, which is input into the big model for intelligent matching and the matching results are returned; If the field matches the business rule, the association relationship between the data field and the business rule is automatically established, and the conversion is performed according to the naming of the data standard and the standard dictionary code. The data field and business rules include format rules, value rules, and logic rules. Evaluate the applicability and impact of data fields and business rules through rule conflict detection mechanisms and verification methods; By probing the field type of the result and mapping it with the built-in conversion rule template, conversion rules are generated. The conversion rules include removing special characters from the name field, converting the certificate code field to uppercase, and unifying the time field to a standard format.
5. A device for implementing the method for intelligently constructing a data warehouse based on a large model according to any one of claims 1 to 4, characterized in that: It includes data source registration module, change perception module, large model intelligent analysis and decision module, data collection module, data exploration module, data cleaning and conversion module, and data multi-dimensional fusion module; The data source registration module is used to manage data source connection information and monitor changes in data source status and metadata, and includes a connection configuration unit, a status monitoring unit, and a metadata change monitoring unit.
6. The data warehouse intelligent construction device based on a large model as claimed in claim 5, characterized in that: The connection configuration unit is used for the full life cycle management of the data source connection, obtains the text content input by the user, and transmits the text content to the large model intelligent analysis and decision module, automatically analyzes the parameters required for the connection, completes the configuration and management of the data source, and uses a high-intensity encryption algorithm to ensure the safe storage of the connection information. The text content includes IP, port, user name, and password; The connection configuration unit is also used to implement network connectivity testing, database response time testing, and authority verification testing, and to timely discover and warn of potential connection problems through performance indicator collection and analysis.
7. The data warehouse intelligent construction device based on a large model as claimed in claim 5, characterized in that: The state monitoring unit is used to collect and track various operating indicators of the data source in real time, including core indicators and system-level performance indicators. The core indicators include connection response time, data throughput, and error rate. The system-level performance indicators include CPU usage, memory usage, and IO performance; and push these indicators to the change perception module, which intelligently adjusts the parallelism of data collection according to various operating indicators; The metadata monitoring unit is used to obtain all-round metadata in real time through an automated real-time monitoring mechanism, including the table structure, field attributes, index information, and statistical information of the data source, and push the changed metadata information to the data change perception module to achieve timely adjustment of the configuration of processing links including data collection, data conversion, and data fusion after the metadata of the business system changes.
8. The data warehouse intelligent construction device based on a large model as claimed in claim 5, characterized in that: The change sensing module is used to monitor the state change event information of each module, track and distribute the state events, and includes a change monitoring unit, a state event tracking unit, and a state event distribution unit; 1) Change monitoring unit, which is used to capture structural changes, data source operation indicator changes, and platform service status changes in real time through event-driven monitoring mechanism, and implements rule-based change filtering and aggregation mechanism. Through refined row-level tracking and intelligent event routing, it ensures that change events can be processed in a timely manner according to business importance; 2) Status event tracking unit, which is used to track the entire process of status events from creation, distribution to completion in real time, store status event information persistently, and support multi-dimensional status event query and statistical analysis; 3) A status event distribution unit, used to distribute status events to corresponding modules according to their types.
9. The data warehouse intelligent construction device based on a large model as claimed in claim 5, characterized in that: The large model intelligent analysis and decision module is used to receive the prompt words and large model call requests of each module, and return the analysis results of the large model, including a prompt word construction unit, a large model call unit, a call result parsing unit, and an abnormal reflection unit; 1) The prompt word construction unit is used for intelligent prompt word construction for data processing scenarios. Through a multi-dimensional information collection mechanism, it automatically obtains context information, including data table structure, field attributes, business rules, as well as domain knowledge and experience knowledge, and dynamically fills and optimizes parameters based on the scenario-based prompt word template library and in combination with specific task requirements; 2) A large model calling unit, which is used to send prompt words of each data processing task type to the large model, and the large model returns the result in a specified format according to the context requirements of the prompt words; 3) Call the result parsing unit to convert the unstructured text returned by the model into standard JSON or other structured formats through the formatting engine, and ensure the legitimacy and integrity of the results through multi-dimensional verification rules; 4) The exception reflection unit is used to capture various exceptions of each task during the automatic collection and processing process, conduct in-depth cause analysis based on the task context information, and automatically optimize the prompt word construction strategy and call parameter configuration by analyzing the exception mode and influencing factors based on the self-optimization process of the reflection mechanism.
10. The data warehouse intelligent construction device based on a large model as claimed in claim 5, characterized in that: The data acquisition module is used to analyze the data table acquisition parameters of the business system through the large model, complete the data acquisition through the acquisition operation, and form the original data table, including the acquisition prompt word generation unit, the table feature analysis unit, the acquisition parameter configuration unit, the acquisition operation generation unit, and the acquisition status monitoring unit; 1) A collection prompt word generation unit is used to integrate the data table name, description, field information, sampled data content, and data collection parameter recommendation template to be collected, submit them to the intelligent analysis and decision module, and return the table feature analysis results through the big model; 2) Table feature parsing unit, used to parse the table feature analysis results returned by the large model, including whether it is a temporary table, data table type, data collection method, incremental collection time field, coding identification field, name field, data table types include entity data, transaction data, dimension data, data collection methods include full collection and incremental collection; 3) The collection job generation unit is used to automatically generate a complete collection job configuration based on the table feature analysis results, including data source connection parameters, reading strategy, and concurrency, and to achieve intelligent allocation and dynamic adjustment of computing resources through the resource scheduling engine; 4) Collection status monitoring unit, used to detect the running status of the collection operation, as well as the data integrity, accuracy, and consistency, and push the status information to the change perception module; (5) Data exploration module Used to generate a detection result report for the collected raw data through a statistical analysis algorithm, including a data distribution analysis unit, a detection report generation unit, and a cleaning conversion suggestion generation unit; 1) Data distribution analysis unit, used to extract multi-dimensional features of data through statistical analysis algorithms, including analyzing statistical indicators of numerical fields, where the statistical indicators include maximum value, minimum value, mean, median, and standard deviation; identifying periodic features of time fields, where the periodic features include time span, update frequency, and change pattern; and deeply analyzing the null value distribution of text fields to obtain the null value ratio; 2) A detection report generation unit, which is used to automatically generate a detection report of the collected data table, and the detection report includes a data overview, a quality score and a problem distribution; 3) a cleaning conversion suggestion generating unit, which is used to generate data conversion suggestions through a large model according to the characteristics of the field types of the exploration results, and the field types of the exploration results include text fields, time fields, and category fields; (6) Data cleaning and conversion module It is used to recommend business rules and conversion rules for the collected raw data through the big model according to the data exploration results, complete the data cleaning and conversion, and form a standardized result table. The data cleaning and conversion module includes a business rule binding unit, a data standard association unit, a conversion rule generation unit, and a conversion job execution unit; 1) Business rule binding unit, which is used to construct prompt word text by combining the meaning of the field with the rule knowledge base, submit it to the big model for intelligent matching, and return whether the field matches the business rule. If so, it automatically establishes the association between the data field and the business rule, and evaluates the applicability and scope of influence of the rule through the rule conflict detection mechanism and verification means to ensure the accuracy and effectiveness of rule execution. The rule knowledge base is composed of rules written by business experts and demand analysts based on domain business knowledge, including rule name, rule description, rule type, and regular expression; 2) Data standard association unit, which is used to construct the meaning of the category field and the data standard library into prompt word text, submit it to the big model for intelligent matching, and return whether the field matches the data standard. If so, it automatically establishes the association relationship between the data field and the data standard, and converts according to the naming of the data standard and the standard dictionary code; the data standard library includes international standards, ministerial standards, and industry standards, which are periodically obtained from public sites or documents, and are solidified into a data standard library through manual entry or automated program parsing; 3) A conversion rule generation unit, which is used to generate appropriate conversion rules by detecting the field type of the result and based on the mapping relationship with the built-in conversion rule template, such as removing special characters from the name field, converting the certificate code field to uppercase, and unifying the time field to a standard format; 4) A conversion job execution unit, which is used to create and run conversion jobs according to business rules and conversion rules, realize the processing of data from collected to converted, and form standardized result data; 5) A conversion status monitoring unit, which is used to detect the running status of the conversion job and push the status information to the change perception module; (7) Data multi-dimensional fusion module It is used to transform the standardized result data output by the data cleaning conversion module, identify the subject layer target table to be fused through the large model, and complete data fusion, including the target table matching unit, field mapping unit, fusion rule configuration unit, and fusion execution unit; 1) The target table matching unit is used to construct the prompt word text from the standardized table and the target related information of the subject layer, submit it to the intelligent analysis and decision module, and return the table feature analysis results through the large model; the prompt word text is matched and identified with each data table of the subject layer for semantic and structural similarity, and the recognition result is returned. The returned recognition result includes whether the target table matches consistently, the field matching relationship, and the confidence level; 2) A fusion rule configuration unit, which is used to generate fusion rules and configure them on the fields of the subject layer data table according to the results of the target table matching unit; 3) A fusion execution unit is used to generate data fusion jobs according to fusion rules and execute fusion jobs to achieve the fusion of standardized result data output by the data cleaning conversion module into the target table of the subject layer; 4) Fusion status monitoring unit, used to detect the running status of the fusion job and push the status information to the change perception module.
Citation Information
Patent Citations
Data fusion method and device, electronic equipment and storage medium
CN114386509A
Multi-source heterogeneous data mapping method of dynamic ontology semantic fusion model
CN115630066A
Method for providing high-quality data for multi-mode large model system
CN117743315A
Commercial intelligent decision-making question-answering system and method of knowledge graph-driven large model
CN118227767A
Metadata management system and method of integrated large model
CN118643071A
Cited By
Digital main line automatic modeling method and system based on large language model
CN120523828A
Interface data automatic integration processing method and device based on configuration
CN120892485A
Automatic data quality inspection system and method based on large model and data flow arrangement
CN120893585A
Automated Data Quality Inspection System and Method Based on Large Model and Data Flow Orchestration
CN120893585B
Method for automatically extracting ecological parameters and driving factors of large language model
CN120910562A