A large model-based data warehouse intelligent construction method and device
By adopting a data warehouse intelligent construction method based on a large model, the challenges of intelligentization and automation of the entire data processing process have been solved. This method automates and intelligentizes data collection, exploration, cleaning, transformation, and fusion, thereby improving data processing efficiency and accuracy and reducing the cost of manual intervention.
Patent Information
- Application Number
- CN202510106818.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-01-23
AI Technical Summary
Existing technologies lack the ability to be intelligent and automated throughout the entire data processing process, resulting in cumbersome configuration of data acquisition parameters, difficulty in accurately establishing data relationships, and a lack of intelligent features in fusion rules, making it difficult to adapt to the changing needs of different scenarios.
We adopt a data warehouse intelligent construction method based on a large model. Through multi-level task processing modules and parallel processing mechanisms, we introduce a large model to perform table feature analysis and parameter recommendation in the data acquisition stage, realize field feature identification and cleaning and transformation suggestions in the data exploration stage, and realize intelligent field mapping and rule generation in the data fusion stage.
It has achieved full automation and intelligence in the data processing process, significantly improving data processing efficiency and accuracy, reducing the cost of manual intervention, increasing the automation rate of the entire data processing process to over 90%, improving processing efficiency by 5-10 times, increasing the data quality compliance rate to 90%, and reducing manual intervention by 80%.
Smart Images

Figure CN120045545B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of big data processing technology, specifically to a method and apparatus for intelligent construction of a data warehouse based on a large model, used to realize intelligent data collection, processing, and management. This invention is particularly suitable for intelligent and automated data processing in scenarios such as enterprise data platforms and data warehouses. Background Technology
[0002] As enterprises deepen their digital transformation, they face a large demand for processing heterogeneous data. While some data processing tools and platforms exist, these solutions often focus only on a specific stage, lacking overall control and intelligent coordination of the entire data processing process. This leads to numerous problems: In the data acquisition phase, the process of configuring data acquisition parameters is cumbersome and complex due to the need for manual judgment on whether a table needs to be collected, determining the collection cycle, and selecting incremental fields; in the data classification and standardization phase, the identification of data table types relies heavily on human experience, coupled with inaccurate understanding of field semantics and the need for professional knowledge in standard mapping, making it difficult to accurately establish data relationships; in the data fusion phase, field mapping requires significant manual intervention, and the fusion rules lack intelligent features, making it difficult to guarantee data consistency and fusion effectiveness.
[0003] More importantly, existing technologies generally lack intelligent processing capabilities, cannot automatically identify data characteristics, lack intelligent decision-making capabilities, and have overly rigid processing rules that are difficult to adapt to changing needs in different scenarios.
[0004] Therefore, there is an urgent need for a solution that can automate the entire data processing process. Summary of the Invention
[0005] To address the challenges of automation and intelligence in data processing, the present invention aims to provide a data warehouse intelligent construction method based on a large model, thereby achieving full automation and intelligence of the data process from data collection and processing in business systems to the subject layer of the data warehouse.
[0006] Another objective of this invention is to provide an apparatus for implementing the above-mentioned intelligent data warehouse construction method for large models. By constructing multi-level task processing modules and parallel processing mechanisms, a large model is introduced in the data acquisition stage for table feature analysis and parameter recommendation. In the data exploration stage, field feature identification and cleaning and transformation suggestions are generated. In the data transformation stage, business rules and associated data standards are automatically bound. In the data fusion stage, intelligent field mapping and rule generation are achieved.
[0007] The objective of this invention and the technical problem it solves are achieved by the following technical solutions. According to the present invention, a data warehouse intelligent construction method based on a large model is proposed. The method includes: a user registering data source information of a business system to be collected through a data source registration portal and submitting it as a data processing task. The data source information includes IP address, port, username, password, and database. The task executor receives the data processing task, obtains data table information from the registered data source, calls the large model to perform feature analysis on the data table, and identifies the table type, business attributes, and collection parameters. Data collection is performed based on the analysis results, including identifying whether the data table is a temporary table, determining the data collection period, and selecting incremental fields. The collected data is intelligently explored to identify field characteristics and data quality, generating exploration results. Field characteristics are field types, including numeric, time-based, and text-based. Data transformation is performed based on the exploration results, automatically binding business rules, associating data standards, and generating transformation rules. The data cleaning and transformation job is submitted and run. Based on the standardized table and the relevant information of the target table at the theme layer, prompt text is constructed. The prompt text is input into the large model, intelligently matching the target table and field mapping relationship, generating fusion rules, and executing the fusion.
[0008] Furthermore, the apparatus for the intelligent construction method of data warehouse based on large model includes: a data source registration module, a change perception module, a large model intelligent analysis and decision-making module, a data acquisition module, a data exploration module, a data cleaning and transformation module, and a data multi-dimensional fusion module; the data source registration module is used to manage data source connection information and monitor changes in data source status and metadata, including a connection configuration unit, a status monitoring unit, and a metadata change monitoring unit.
[0009] Furthermore, in the apparatus of the intelligent data warehouse construction method based on the large model, the connection configuration unit is used for the full lifecycle management of the data source connection, obtains the text content input by the user, and transmits the text content to the intelligent analysis and decision module of the large model, automatically analyzes the parameters required for the connection, completes the configuration and management of the data source, and uses a high-strength encryption algorithm to ensure the secure storage of connection information. The text content includes IP, port, username, and password.
[0010] This unit supports connection configuration for various data sources such as MySQL, Oracle, and PostgreSQL. Through a text input box, the user enters text containing the necessary connection information, including IP address, port, username, password, database, and other key information. The text content is then passed to the large model intelligent analysis and decision-making module, which automatically analyzes the parameters required for the connection, completes the configuration and management of the data source, and uses a high-strength encryption algorithm to ensure the secure storage of connection information.
[0011] The connection configuration unit is also used to perform network connectivity testing, database response time testing, and permission verification testing. Through performance indicator collection and analysis, it can promptly identify and warn of potential connection problems. This unit also provides comprehensive connection testing and verification functions, including network connectivity testing, database response time testing, and permission verification testing. Through performance indicator collection and analysis, it can promptly identify and warn of potential connection problems, ensuring the stability and reliability of data source connections.
[0012] Furthermore, in the apparatus for the intelligent data warehouse construction method based on large models, the status monitoring unit is used to collect and track various operational indicators of the data source in real time, including core indicators such as connection response time, data throughput, and error rate, as well as system-level performance indicators such as CPU utilization, memory usage, and IO performance. These indicators are pushed to the change perception module, which intelligently adjusts the parallelism of data collection based on these operational indicators to reduce the impact on the normal operation of the business system. The metadata monitoring unit is used to acquire comprehensive metadata of the data source, such as table structure, field attributes, index information, and statistical information, in real time through an automated real-time monitoring mechanism. Changes in metadata information are pushed to the data change perception module, enabling timely adjustment of the configuration of data collection, data transformation, and data fusion processing stages after changes in the business system's metadata.
[0013] Furthermore, in the apparatus for the intelligent construction method of data warehouse based on large models, the change perception module is used to monitor the state change event information of each module, track and distribute the state events, and includes a change monitoring unit, a state event tracking unit, and a state event distribution unit.
[0014] 1) The change monitoring unit is used to capture structural changes, changes in data source operation indicators, and changes in platform service status in real time through an event-driven monitoring mechanism. It implements a rule-based change filtering and aggregation mechanism, and ensures that change events can be processed in a timely manner according to business importance through row-level fine-grained tracking and intelligent event routing.
[0015] 2) Status event tracking unit, used to track the entire process of status events from creation and distribution to completion in real time, and to persistently store status event information, supporting multi-dimensional status event query and statistical analysis.
[0016] 3) Status event distribution unit, used to distribute status events to the corresponding modules according to their type.
[0017] Furthermore, in the apparatus for the intelligent construction method of data warehouse based on large model, the intelligent analysis and decision-making module of large model is used to receive prompt words and large model call requests from various modules and return the analysis results of the large model, including a prompt word construction unit, a large model call unit, a call result parsing unit, and an anomaly reflection unit;
[0018] 1) Prompt word construction unit, used for intelligent prompt word construction for data processing scenarios. Through a multi-dimensional information collection mechanism, it automatically obtains contextual information, including data table structure, field attributes, business rules, domain knowledge and experience knowledge, and dynamically fills and optimizes parameters based on a scenario-based prompt word template library and specific task requirements.
[0019] 2) The large model calling unit is used to send prompts for various data processing task types to the large model, which then returns results in a specified format, such as JSON, based on the context requirements of the prompts.
[0020] 3) Call the result parsing unit, which is used to convert the unstructured text returned by the model into standard JSON or other structured formats through the formatting engine, and ensure the legality and integrity of the results through multi-dimensional validation rules.
[0021] 4) The anomaly reflection unit is used to capture various anomalies of each task during the automatic collection and processing process, perform in-depth cause analysis in combination with task context information, and automatically optimize the prompt word construction strategy and call parameter configuration based on the self-optimization process of the reflection mechanism by analyzing anomaly patterns and influencing factors.
[0022] Furthermore, the apparatus for the intelligent construction method of data warehouse based on large model includes a data acquisition module used to analyze the data table acquisition parameters of the business system through the large model, complete the data acquisition through acquisition jobs, and form an original data table. The module includes an acquisition prompt word generation unit, a table feature analysis unit, an acquisition parameter configuration unit, an acquisition job generation unit, and an acquisition status monitoring unit.
[0023] 1) The data collection prompt word generation unit is used to integrate the data table name, description, field information, and sampled data content with the data collection parameter recommendation template and submit them to the intelligent analysis and decision module. The large model then returns the table feature analysis results.
[0024] 2) Table Feature Parsing Unit: This unit is used to parse the table feature analysis results returned by the large model, including whether it is a temporary table, the type of data table, the data collection method, the incremental collection time field, the code identifier field, and the name field. The data table types include entity data, transaction data, and dimension data. The data collection methods include full collection and incremental collection.
[0025] 3) The data acquisition job generation unit is used to automatically generate a complete data acquisition job configuration based on the table feature parsing results, including data source connection parameters, reading strategy, and concurrency, and realize intelligent allocation and dynamic adjustment of computing resources through the resource scheduling engine.
[0026] 4) Data acquisition status monitoring unit, used to detect the running status of data acquisition operations, as well as the integrity, accuracy and consistency of data, and push status information to change perception module.
[0027] (5) Data Exploration Module
[0028] This tool is used to generate exploration result reports from the collected raw data through statistical analysis algorithms. It includes a data distribution analysis unit, an exploration report generation unit, and a cleaning and transformation suggestion generation unit.
[0029] 1) Data distribution analysis unit, used to extract multi-dimensional features from data through statistical analysis algorithms, including analyzing statistical indicators of numerical fields, such as maximum value, minimum value, mean, median, and standard deviation; identifying periodic features of time fields, such as time span, update frequency, and change pattern; and conducting in-depth analysis of the null value distribution of text fields to obtain the proportion of null values.
[0030] 2) Exploration report generation unit, used to automatically generate exploration reports of the collected data tables. The exploration report includes data overview, quality score and problem distribution.
[0031] 3) The data cleaning and transformation suggestion generation unit is used to generate data transformation suggestions based on the characteristics of the field types of the exploration results through a large model. The field types of the exploration results include text fields, time fields, and category fields.
[0032] (6) Data cleaning and conversion module
[0033] Based on the data exploration results, the system cleanses and transforms the collected raw data by recommending business rules and transformation rules through a large model, forming a standardized result table. It includes a business rule binding unit, a data standard association unit, a transformation rule generation unit, and a transformation job execution unit.
[0034] 1) The business rule binding unit is used to construct prompt text by combining the meaning of fields with the rule knowledge base, submitting it to the large model for intelligent matching, and returning whether the field matches the business rule. If so, it automatically establishes the association between the data field and the business rule, and evaluates the applicability and scope of impact of the rule through rule conflict detection mechanisms and verification methods to ensure the accuracy and effectiveness of rule execution. The rule knowledge base consists of rules written by business experts and requirements analysts based on domain business knowledge, including rule name, rule description, rule type, and regular expression.
[0035] 2) The data standard association unit is used to construct prompt text by matching the meaning of category fields with the data standard library, submit it to the large model for intelligent matching, and return whether the field matches the data standard. If it does, the association relationship between the data field and the data standard is automatically established and converted according to the naming of the data standard and the standard dictionary code. The data standard library includes international standards, ministerial standards, and industry standards, which are periodically obtained from public websites or documents and solidified into a data standard library through manual input or automated program parsing.
[0036] 3) The conversion rule generation unit is used to generate appropriate conversion rules based on the field type of the probe results and the mapping relationship with the built-in conversion rule template, such as removing special characters from the name field, uniformly converting the certificate code field to uppercase, and uniformly standardizing the time field. The built-in conversion rule template is solidified into the program during the development stage and can quickly generate conversion rules based on the fields during the runtime stage.
[0037] 4) The conversion job execution unit is used to create and run conversion jobs according to business rules and conversion rules, so as to realize the processing of data from the collected data to the converted data and form standardized result data.
[0038] 5) The conversion status monitoring unit is used to detect the running status of the conversion operation and push the status information to the change sensing module.
[0039] (7) Data Multidimensional Fusion Module
[0040] The standardized result data output by the data cleaning and transformation module is used to identify the target table of the subject layer to be merged through the large model, and complete the data fusion. This includes the target table matching unit, field mapping unit, fusion rule configuration unit, and fusion execution unit.
[0041] 1) The target table matching unit is used to construct prompt text from the standardized table and the target-related information of the topic layer, and submit it to the intelligent analysis and decision module. The large model returns the table feature analysis results. The prompt text is matched and identified with each data table in the topic layer in terms of semantic and structural similarity, and the identification results are returned. The returned identification results include whether the target table matches, the field matching relationship, and the confidence level.
[0042] 2) The fusion rule configuration unit is used to generate fusion rule configurations on the fields of the theme layer data table based on the results of the target table matching unit.
[0043] 3) The fusion execution unit is used to generate data fusion jobs according to fusion rules and execute fusion jobs to integrate the standardized result data output by the data cleaning and transformation module into the target table of the theme layer.
[0044] 4) The fusion status monitoring unit is used to detect the running status of the fusion operation and push the status information to the change perception module.
[0045] Based on the above technical solution, the apparatus and method for intelligent construction of data warehouses based on large models of the present invention have at least the following advantages:
[0046] 1) This invention achieves intelligent processing of the entire data processing process by introducing large model technology, which significantly improves data processing efficiency and accuracy and reduces the cost of manual intervention;
[0047] 2) Practical application verification shows that this invention increases the automation rate of the entire data processing process to over 90%, improves processing efficiency by 5-10 times, increases the data quality compliance rate to 90%, reduces manual intervention by 80%, and saves a lot of data processing manpower costs every year.
[0048] 3) At the same time, the system has the adaptability to mainstream data sources and the ability to handle more than 90% of abnormal scenarios, providing an efficient, accurate and economical solution for enterprise digital transformation. Attached Figure Description
[0049] Figure 1 The diagram shown is an overall structural diagram of the intelligent data warehouse construction device based on a large model according to the present invention.
[0050] Figure 2 The diagram shown is the overall flowchart of the intelligent data warehouse construction method based on a large model according to the present invention;
[0051] Figure 3 This is a schematic diagram of an example of a data table to be collected from a company's human resources business system. Detailed Implementation
[0052] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions provided by this invention will be further described in detail below with reference to the accompanying drawings and embodiments. Obviously, the described embodiments are only some embodiments of this invention, and not all embodiments.
[0053] This invention provides a method and apparatus for intelligent construction of a data warehouse based on a large model, belonging to the field of big data processing technology. The method includes: a user registering data source information for a business system to be processed through a data source registration portal. This data source information includes IP address, port, username, password, database, etc., and is submitted as a data processing task; a task executor receiving the data processing task, obtaining data table information from the registered data source, calling a large model to perform feature analysis on the data table, identifying table type, business attributes, and collection parameters; performing data collection based on the analysis results, where the analysis results include identifying whether the data table is a temporary table, determining the data collection period, and selecting incremental fields; and intelligently probing the collected data to identify field characteristics and data quality, generating probing results, where field characteristics are field types, including numeric, time-based, and text-based; and performing data transformation based on the probing results, automatically binding business rules, associating data standards, and generating transformation rules. Submit and run the data cleaning and transformation job; based on the standardized table and the relevant information of the topic layer target table, construct prompt text, input the prompt text into the large model, intelligently match the target table and field mapping relationship, generate fusion rules and execute fusion. The prompt text includes the standardized table name, table description, field information, domain knowledge and experience knowledge, and relevant information of the topic layer target table. The prompt text is matched and identified with each data table in the topic layer for semantic and structural similarity, and the identification results are returned. The returned identification results include whether the target table matches, field matching relationship, and confidence level. If a match is found, fusion rules are generated and configured on the fields of the topic layer data table. A data fusion job is generated according to the fusion rules and executed to fuse the standardized result data output by the data cleaning and transformation module into the topic layer target table.
[0054] Based on the above method, the present invention provides an intelligent data warehouse construction device based on a large model, such as... Figure 1 It mainly includes a data source registration module, a change perception module, a large-scale intelligent analysis and decision-making module, a data acquisition module, a data exploration module, a data cleaning and transformation module, and a multi-dimensional data fusion module.
[0055] The data source registration module 101 is used to manage data source connection information and monitor changes in data source status and metadata, including a connection configuration unit, a status monitoring unit, and a metadata change monitoring unit.
[0056] The connection configuration unit is primarily responsible for the entire lifecycle management of data source connections. This unit supports connection configuration for various data source types, including MySQL, Oracle, and PostgreSQL. Through a text input box, users enter necessary connection information, such as IP address, port, username, password, and database details. This text is then fed into the large-scale intelligent analysis and decision-making module, which automatically analyzes the required parameters, completes the data source configuration and management, and employs high-strength encryption algorithms to ensure the secure storage of connection information. Simultaneously, this unit provides comprehensive connection testing and verification functions, including network connectivity testing, database response time testing, and permission verification testing. Through performance indicator collection and analysis, it promptly identifies and warns of potential connection problems, ensuring the stability and reliability of data source connections.
[0057] The status monitoring unit is used to collect and track various operational metrics of the data source in real time, including core metrics such as connection response time, data throughput, and error rate, as well as system-level performance metrics such as CPU utilization, memory usage, and I / O performance. These metrics are pushed to the change awareness module, which intelligently adjusts the parallelism of data collection based on the operational metrics, reducing the impact on the normal operation of the business system.
[0058] The metadata monitoring unit is used to acquire comprehensive metadata from the data source, such as table structure, field attributes, index information, and statistical information, through an automated real-time monitoring mechanism. It pushes the changed metadata information to the data change perception module, enabling timely adjustments to the configuration of data collection, data transformation, and data fusion processes after changes in the business system's metadata.
[0059] The change sensing module 102 is used to monitor the state change event information of each module, track and distribute the state events, including a change monitoring unit, a state event tracking unit, and a state event distribution unit.
[0060] 1) The change monitoring unit is used to capture structural changes (such as table creation, field changes, index modifications), changes in data source operation metrics, and changes in platform service status in real time through an event-driven monitoring mechanism. It implements a rule-based change filtering and aggregation mechanism, and ensures that change events can be processed in a timely manner according to business importance through row-level fine-grained tracking and intelligent event routing.
[0061] 2) The status event tracking unit mainly tracks the entire process of status events from creation and distribution to completion in real time, and persists the status event information for storage, supporting multi-dimensional status event query and statistical analysis.
[0062] 3) The status event distribution unit is used to distribute status events to the corresponding modules according to their type.
[0063] The large model intelligent analysis and decision-making module 103 is used to receive prompt words and large model call requests from various modules, and return the analysis results of the large model, including a prompt word construction unit, a large model call unit, a call result parsing unit, and an anomaly reflection unit.
[0064] 1) The prompt word construction unit is used for intelligent prompt word construction for data processing scenarios. Through a multi-dimensional information collection mechanism, it automatically obtains contextual information such as data table structure, field attributes, and business rules, as well as domain knowledge and experience knowledge. Based on the scenario-based prompt word template library (including various templates such as data collection parameter recommendation and field mapping relationship recognition), it dynamically fills in and optimizes parameters according to specific task requirements.
[0065] 2) The large model calling unit is used to send the prompts for each data processing task type to the large model, which then returns the results in a specified format, such as JSON, according to the context requirements of the prompts.
[0066] 3) The result parsing unit is used to convert the unstructured text returned by the model into standard JSON or other structured formats through the formatting engine, and to ensure the legality and integrity of the results through multi-dimensional validation rules.
[0067] 4) The anomaly reflection unit is used to capture various anomalies (such as call timeouts, result anomalies, parsing failures, and collection task anomalies) in each task during automatic data collection and processing. It performs in-depth root cause analysis based on task context information and, based on the self-optimization process of the reflection mechanism, automatically optimizes the prompt word construction strategy and call parameter configuration by analyzing anomaly patterns and influencing factors. The anomaly reflection unit maintains a knowledge base of manual handling opinions, continuously accumulates and summarizes anomaly handling experience, and achieves automatic adjustment and optimization of anomaly handling strategies through periodic fine-tuning of prompt words, thereby continuously improving the system's anomaly handling capabilities and service stability.
[0068] The data acquisition module 104 is used to collect parameters of the data table of the business system through large model analysis, and to complete the data collection through acquisition jobs to form the original data table. It includes a collection prompt word generation unit, a table feature analysis unit, a collection parameter configuration unit, a collection job generation unit, and a collection status monitoring unit.
[0069] 1) The data collection prompt word generation unit is used to integrate the name, description, field information, and sampled data content of the data table to be collected with the data collection parameter recommendation template, and submit them to the intelligent analysis and decision module, which returns the table feature analysis results through the large model.
[0070] 2) The table feature parsing unit is used to parse the table feature analysis results returned by the large model, including whether it is a temporary table, the type of data table (entity data, transaction data, dimension data), the data collection method (full collection, incremental collection), the incremental collection time field, the code identifier field, the name field, etc.
[0071] 3) The data acquisition job generation unit is used to automatically generate a complete data acquisition job configuration based on the table feature parsing results, including data source connection parameters, reading strategies, concurrency, etc., and realizes intelligent allocation and dynamic adjustment of computing resources through the resource scheduling engine.
[0072] 4) The data acquisition status monitoring unit is used to detect the running status of the data acquisition operation, as well as indicators such as data integrity, accuracy, and consistency, and push the status information to the change perception module.
[0073] The data exploration module 105 is used to generate exploration result reports from the collected raw data through statistical analysis algorithms, including a data distribution analysis unit, an exploration report generation unit, and a cleaning and transformation suggestion generation unit.
[0074] 1) The data distribution analysis unit is used to extract multi-dimensional features from data through statistical analysis algorithms, including statistical indicator analysis of numerical data (maximum, minimum, mean, median, standard deviation, etc.), periodic feature identification of time data (time span, update frequency, change pattern, etc.), and in-depth analysis of the distribution of missing values (proportion of missing values, pattern recognition, cause analysis, etc.).
[0075] 2) The investigation report generation unit is used to automatically generate investigation reports for the collected data tables. The investigation reports cover statistical analysis from multiple dimensions, including data overview, quality score, and problem distribution.
[0076] 3) The cleaning and transformation suggestion generation unit is used to generate data transformation suggestions based on the characteristics of the field types (text fields, time fields, category fields) of the exploration results through a large model.
[0077] The data cleaning and transformation module 106 is used to clean and transform the collected raw data based on the data exploration results, through business rules and transformation rules recommended by the large model, and form a standardized result table, including a business rule binding unit, a data standard association unit, a transformation rule generation unit, and a transformation job execution unit.
[0078] 1) The business rule binding unit is used to construct prompt text by combining field meanings with the rule knowledge base, submitting it to the large model for intelligent matching, and returning whether the field matches the business rule. If so, it automatically establishes the association between the data field and the business rule, and evaluates the applicability and scope of impact of the rule through rule conflict detection mechanisms and verification methods to ensure the accuracy and effectiveness of rule execution. The rule knowledge base consists of rules written by business experts and requirements analysts based on domain business knowledge, including rule name, rule description, rule type, and regular expression. The rule knowledge base also evaluates the applicability and scope of impact of the rules through rule conflict detection mechanisms and verification methods to ensure the accuracy and effectiveness of rule execution.
[0079] 2) The data standard association unit is used to construct prompt text by matching the meaning of category fields with the data standard library, submit it to the large model for intelligent matching, and return whether the field matches the data standard. If it does, the association relationship between the data field and the data standard is automatically established, and the data is converted according to the naming of the data standard and the standard dictionary code. The data standard library includes international standards, ministerial standards, and industry standards, which are periodically obtained from public sites or documents and solidified into a data standard library through manual input or automated program parsing.
[0080] 3) The conversion rule generation unit is used to generate appropriate conversion rules based on the field type of the probe results and the mapping relationship with the built-in conversion rule template, such as removing special characters from the name field, uniformly converting the certificate code field to uppercase, and unifying the time field to a standard format.
[0081] 4) The conversion job execution unit is used to create and run conversion jobs according to business rules and conversion rules, so as to realize the processing of data from collected data to converted data and form standardized result data.
[0082] 5) The conversion status monitoring unit is used to detect the running status of the conversion operation and push the status information to the change sensing module.
[0083] The multidimensional data fusion module 107 is used to identify the target table of the subject layer to be fused through the large model, and complete the data fusion by using the standardized result data output by the data cleaning and transformation module. It includes a target table matching unit, a field mapping unit, a fusion rule configuration unit, and a fusion execution unit.
[0084] 1) The target table matching unit is used to construct prompt word text from the standardized table and the target-related information of the topic layer, and submit it to the intelligent analysis and decision module. The large model returns the table feature analysis results. The prompt word text is matched and identified with each data table in the topic layer in terms of semantic and structural similarity, and the identification results are returned. The returned identification results include whether the target table matches, the field matching relationship, and the confidence level.
[0085] 2) The fusion rule configuration unit is used to generate fusion rule configurations on the fields of the theme layer data table based on the results of the target table matching unit.
[0086] 3) The fusion execution unit is used to generate data fusion jobs according to fusion rules and execute fusion jobs to realize the fusion of standardized result data output by the data cleaning and transformation module into the target table of the theme layer.
[0087] 4) The fusion status monitoring unit is used to detect the running status of the fusion operation and push the status information to the change perception module.
[0088] The following is in conjunction with the appendix Figure 2-3 Tables 1-3 provide a detailed explanation of the main implementation principles, specific implementation methods, and corresponding beneficial effects of the technical solutions in the embodiments of this application.
[0089] Example 1:
[0090] This embodiment takes the construction of a data warehouse for data collection, processing, and fusion in a company's human resources business system as an example to illustrate the specific implementation process of the present invention.
[0091] The main methods and processes are as follows: Figure 2 As shown.
[0092] Step 201, Intelligent Data Source Registration. Users register the data source information of the business system to be collected through the data source registration portal, including IP address, port, username, password, database, etc., and submit it as a data processing task.
[0093] The registration information provided in this embodiment is shown in Table 1 below:
[0094] Table 1 Data source information of the business systems to be collected
[0095]
[0096]
[0097] Based on the provided registration information, combined with domain knowledge and experience knowledge, a prompt word is generated and submitted to the large model intelligent analysis and decision-making module 103. The module returns the recognized JSON result data structure. The data source registration module 101 automatically calls the data source registration API to complete the data source registration, avoiding the manual input of each item in the traditional data source registration form.
[0098] Step 202, Intelligent Data Acquisition. Obtain data table information from the registered data source, perform feature analysis on the data tables using a large model, and identify table type, business attributes, and acquisition parameters. Based on the analysis results, execute data acquisition, including identifying whether it is a temporary table, determining the acquisition period, selecting incremental fields, etc., and submit and run the data acquisition job. Use open-source large models for private deployment, such as QWen 2.5 and Llama 3.1, or commercial large models like ChatGPT, Deepseek, and Gemini. These all provide standard, unified APIs that can receive prompt text and return table feature analysis results, including whether it is a temporary table, the type of data table (entity data, transaction data, dimensional data), the data acquisition method (full acquisition, incremental acquisition), the incremental acquisition time field, the encoding identifier field, and the name field.
[0099] In this embodiment, metadata information from employee registration forms, employee files, and gender tables will be obtained, such as... Figure 3 As shown.
[0100] Then, the metadata information of each table, combined with domain knowledge and experience knowledge, is merged into prompt words and submitted to the large model intelligent analysis and decision-making module 103. The module returns the recognized JSON result data structure. Taking the employee registration form as an example,
[0101] The prompt words are shown in Table 2 below:
[0102] Table 2 analyzes the prompts in the employee onboarding registration forms to be collected.
[0103]
[0104]
[0105] The returned results are shown in Table 3 below:
[0106] Table 3. Identification Results of Employee Onboarding Registration Forms to be Collected
[0107]
[0108] The data acquisition module 104 automatically calls the acquisition job creation API to create and execute the acquisition job, avoiding the manual analysis and input operations required for traditional acquisition job configuration. The change sensing module 102 senses the job's running status in real time, and proceeds to step 203 after the job runs successfully.
[0109] Step 203, Intelligent Data Exploration. The data exploration module 105 performs intelligent exploration on the collected data, identifying field characteristics and data quality. Based on the characteristics of text, time, and category fields, it generates data transformation suggestions through a large-scale model.
[0110] In this embodiment, an employee registration form is used as an example. Data transformation suggestions are generated through statistical analysis algorithms and large-scale models. The results are shown in Table 4 below:
[0111] Table 4: Exploration Results and Conversion Suggestions for Employee Registration Forms
[0112]
[0113]
[0114] The change sensing module 102 senses and detects the operation status of the operation in real time. When the operation is successfully completed, it proceeds to step 204.
[0115] Step 204, Intelligent Data Cleaning and Transformation. Based on the exploration results, data transformation is performed, including automatically binding business rules, associating data standards, and generating transformation rules. The data cleaning and transformation job is then submitted and executed. By constructing prompt text using field meanings and a rule knowledge base, this text is submitted to a large model for intelligent matching. The model returns whether the field matches the business rule. If so, it automatically establishes the association between the data field and the business rule (including format rules, value rules, and logical rules). Through rule conflict detection mechanisms and verification methods, the applicability and scope of impact of the rules are evaluated to ensure the accuracy and effectiveness of rule execution. The rule knowledge base consists of rules written by business experts and requirements analysts based on domain business knowledge, including rule names, rule descriptions, rule types, and regular expressions.
[0116] Next, by constructing prompt text from the category field meanings and the data standard library, this text is submitted to the large model for intelligent matching. The system returns whether the field matches the data standard; if so, it automatically establishes an association between the data field and the data standard, and performs conversion according to the data standard's naming conventions and standard dictionary codes. The data standard library includes international standards, ministerial standards, and industry standards, which are periodically obtained from public websites or documents and solidified into a data standard library through manual input or automated program parsing. Then, by probing the field types of the results and mapping them to the built-in conversion rule templates, appropriate conversion rules are generated, such as removing special characters from name fields, uniformly converting document code fields to uppercase, and standardizing the format of time fields. The built-in conversion rule templates are solidified into the program during the development phase and can quickly generate conversion rules based on the fields during runtime.
[0117] As shown in Table 5, this embodiment uses the employee onboarding registration form as an example. Based on the exploration results and transformation suggestions, the data cleaning and transformation module 106 automatically binds suggested rules to the employee name, mobile phone number, gender, department name, and onboarding time fields through the rule binding API, and creates and runs the data cleaning and transformation job. The change perception module 102 perceives the running status of the transformation job in real time. When the job runs successfully, it proceeds to step 205.
[0118] Step 205: Intelligent Multidimensional Data Fusion Processing. Finally, data fusion is performed. The large-scale model intelligently matches the target table and field mapping relationships at the topic layer, generates fusion rules, and executes the data fusion operation. In this embodiment, taking the employee registration form as an example, the standardized result data output by the data cleaning and transformation module 106 is combined with domain knowledge and experience knowledge to form prompt words. These prompt words include table name, table description, and field information. Semantic analysis and structural similarity matching are then performed on each data table at the topic layer. The results are submitted to the large-scale model intelligent analysis and decision module 103, which returns information such as whether the target table matches, field matching relationships, and confidence levels. The results are shown in Table 5 below.
[0119] Table 5. Mapping Relationship between the Human Resources Pool Testing System (Employee Onboarding Registration Form) and the Thematic Level Target Table (Employee Master Table)
[0120]
[0121] The data multidimensional fusion module 107 automatically calls the fusion job creation API based on the recognition results to create and execute the fusion job, avoiding the manual analysis and input operations required for traditional fusion job configuration. The change perception module 102 senses the job running status in real time. When the job runs successfully, it completes the intelligent and automatic collection, processing, and fusion of data from the business system into the target table at the theme layer, realizing full automation of the data processing process.
[0122] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Therefore, any simple modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A method for intelligently constructing a data warehouse based on a large model, characterized in that, The method includes the following steps: Step S1: Users register the data source information of the business system to be collected through the data source registration portal and submit it as a data processing task; Step S2: Receive data processing tasks through the task executor, obtain data table information from the registered data source, call the large model to perform feature analysis on the data table, and identify the table type, business attributes and collection parameters; Step S3: Perform data collection based on the analysis results. The analysis results include identifying whether the data table is a temporary table, determining the data collection period, and selecting incremental fields. The collected data is then intelligently explored to identify field characteristics and data quality, and exploration results are generated. Field characteristics are field types, including numeric, time, and text types. Step S4: Based on the exploration results, perform data transformation, automatically bind business rules, associate data standards and generate transformation rules, submit and run the data cleaning and transformation job; based on the standardized table and the relevant information of the target table in the theme layer, construct prompt text, input the prompt text into the large model, intelligently match the target table and field mapping relationship, generate fusion rules and execute fusion; The prompt text includes standardized table names, table descriptions, field information, domain knowledge and experience knowledge, as well as information related to the target table in the topic layer. The prompt text is matched and identified semantically and structurally with each data table in the topic layer, and the identification results are returned. The returned identification results include whether the target table matches, the field matching relationship, and the confidence level. If a match is found, fusion rules are generated and configured on the fields of the topic layer data table. A data fusion job is generated according to the fusion rules and executed to fuse the standardized result data output by the data cleaning and transformation module into the target table in the topic layer.
2. The intelligent data warehouse construction method based on a large model as described in claim 1, characterized in that: The data source information includes IP address, port number, username, password, and database.
3. The intelligent data warehouse construction method based on a large model as described in claim 1, characterized in that: The steps described involve intelligently probing the collected data to identify field characteristics and data quality, including the following steps: Analyze the statistical indicators of numerical fields, including maximum value, minimum value, mean, median, and standard deviation. Identify the periodic characteristics of time-based fields, including time span, update frequency, and change patterns. Perform in-depth analysis on the distribution of null values in text fields to obtain the proportion of null values.
4. The intelligent data warehouse construction method based on a large model as described in claim 1, characterized in that: The steps described above involve data transformation based on the exploration results, automatically binding business rules, associating data standards, generating transformation rules, submitting and running the data cleaning and transformation job, including the following steps: The meanings of the fields and the rule knowledge base are used to build prompt word text, which is then input into a large model for intelligent matching, and the matching results are returned. If a field matches a business rule, the association between the data field and the business rule is automatically established, and the data is converted according to the naming of the data standard and the standard dictionary code. The data field and the business rule include format rules, value rules, and logical rules. By using rule conflict detection mechanisms and verification methods, the applicability and scope of impact of data fields and business rules are assessed. By examining the field types of the results and mapping them to the built-in conversion rule templates, conversion rules are generated. These rules include removing special characters from name fields, converting document code fields to uppercase, and standardizing time fields.
5. An apparatus for implementing the intelligent data warehouse construction method based on a large model as described in any one of claims 1-4, characterized in that: It includes a data source registration module, a change perception module, a large-scale intelligent analysis and decision-making module, a data acquisition module, a data exploration module, a data cleaning and transformation module, and a multi-dimensional data fusion module; The data source registration module is used to manage data source connection information and monitor changes in data source status and metadata, including a connection configuration unit, a status monitoring unit, and a metadata change monitoring unit.
6. The intelligent data warehouse construction device based on a large model as described in claim 5, characterized in that: The connection configuration unit is used for the full lifecycle management of the data source connection. It obtains the text content input by the user and transmits the text content to the large model intelligent analysis and decision module. It automatically analyzes the parameters required for the connection, completes the configuration and management of the data source, and uses a high-strength encryption algorithm to ensure the secure storage of connection information. The text content includes IP, port, username, and password. The connection configuration unit is also used to perform network connectivity testing, database response time testing, and permission verification testing, and to promptly detect and warn of potential connection problems through performance indicator collection and analysis.
7. The intelligent data warehouse construction device based on a large model as described in claim 5, characterized in that: The status monitoring unit is used to collect and track various operational indicators of the data source in real time, including core indicators and system-level performance indicators. Core indicators include connection response time, data throughput, and error rate, while system-level performance indicators include CPU utilization, memory usage, and IO performance. These indicators are then pushed to the change perception module, which intelligently adjusts the parallelism of data collection based on the various operational indicators. The metadata change monitoring unit is used to acquire comprehensive metadata in real time through an automated real-time monitoring mechanism, including the table structure, field attributes, index information, and statistical information of the data source. It pushes the changed metadata information to the data change perception module, so that after the metadata of the business system changes, the configuration of the processing links, including data collection, data transformation, and data fusion, can be adjusted in a timely manner.
8. The intelligent data warehouse construction device based on a large model as described in claim 5, characterized in that: The change sensing module is used to monitor the state change event information of each module, track and distribute state events, including a change monitoring unit, a state event tracking unit, and a state event distribution unit. 1) The change monitoring unit is used to capture structural changes, changes in data source operation indicators, and changes in platform service status in real time through an event-driven monitoring mechanism. It implements a rule-based change filtering and aggregation mechanism, and ensures that change events can be processed in a timely manner according to business importance through row-level fine-grained tracking and intelligent event routing. 2) Status event tracking unit, used to track the entire process of status events from creation and distribution to completion in real time, and to persistently store status event information, supporting multi-dimensional status event query and statistical analysis; 3) Status event distribution unit, used to distribute status events to the corresponding modules according to their type.
9. The intelligent data warehouse construction device based on a large model as described in claim 5, characterized in that: The large model intelligent analysis and decision-making module is used to receive prompts and large model call requests from various modules, and return the analysis results of the large model. It includes a prompt construction unit, a large model call unit, a call result parsing unit, and an anomaly reflection unit. 1) Prompt word construction unit, used for intelligent prompt word construction for data processing scenarios. Through a multi-dimensional information collection mechanism, it automatically obtains contextual information, including data table structure, field attributes, business rules, domain knowledge and experience knowledge, and dynamically fills and optimizes parameters based on a scenario-based prompt word template library and specific task requirements. 2) The large model calling unit is used to send prompts for various data processing task types to the large model, which then returns results in a specified format based on the context of the prompts. 3) Call the result parsing unit, which is used to convert the unstructured text returned by the model into standard JSON or other structured formats through the formatting engine, and to ensure the legality and integrity of the results through multi-dimensional validation rules; 4) The anomaly reflection unit is used to capture various anomalies of each task during the automatic collection and processing process, perform in-depth cause analysis in combination with task context information, and automatically optimize the prompt word construction strategy and call parameter configuration based on the self-optimization process of the reflection mechanism by analyzing anomaly patterns and influencing factors.
10. The intelligent data warehouse construction device based on a large model as described in claim 5, characterized in that: The data acquisition module is used to analyze the data table acquisition parameters of the business system through a large model, complete the data acquisition through acquisition jobs, and form the original data table. It includes an acquisition prompt word generation unit, a table feature analysis unit, an acquisition parameter configuration unit, an acquisition job generation unit, and an acquisition status monitoring unit. 1) The data collection prompt word generation unit is used to integrate the data table name, description, field information, and sampled data content with the data collection parameter recommendation template and submit them to the intelligent analysis and decision module, which returns the table feature analysis results through the large model; 2) Table feature parsing unit, used to parse the table feature analysis results returned by the large model, including whether it is a temporary table, the type of data table, the data collection method, the incremental collection time field, the code identifier field, and the name field. The data table types include entity data, transaction data, and dimension data. The data collection methods include full collection and incremental collection. 3) Data acquisition job generation unit, which is used to automatically generate a complete data acquisition job configuration based on the table feature parsing results, including data source connection parameters, reading strategy, concurrency, and realize intelligent allocation and dynamic adjustment of computing resources through the resource scheduling engine; 4) Data acquisition status monitoring unit, used to detect the running status of data acquisition operations, as well as the integrity, accuracy and consistency of data, and push status information to the change perception module; (5) Data Exploration Module This is used to generate exploration result reports from the collected raw data through statistical analysis algorithms, including a data distribution analysis unit, an exploration report generation unit, and a cleaning and transformation suggestion generation unit; 1) The data distribution analysis unit is used to extract multi-dimensional features from data through statistical analysis algorithms. This includes analyzing statistical indicators of numerical fields, such as maximum, minimum, mean, median, and standard deviation; identifying periodic features of time fields, such as time span, update frequency, and change patterns; and conducting in-depth analysis of the distribution of null values in text fields to obtain the proportion of null values. 2) Exploration report generation unit, used to automatically generate exploration reports of the collected data tables. The exploration report includes data overview, quality score and problem distribution; 3) The data cleaning and transformation suggestion generation unit is used to generate data transformation suggestions based on the characteristics of the field types of the exploration results through a large model. The field types of the exploration results include text fields, time fields, and category fields. (6) Data cleaning and conversion module Based on the data exploration results, the data cleaning and transformation module is used to clean and transform the collected raw data by recommending business rules and transformation rules through a large model, and form a standardized result table. The data cleaning and transformation module includes a business rule binding unit, a data standard association unit, a transformation rule generation unit, and a transformation job execution unit. 1) The business rule binding unit is used to construct prompt text by combining the meaning of fields with the rule knowledge base, submit it to the large model for intelligent matching, and return whether the field matches the business rule. If it does, the association between the data field and the business rule is automatically established. Through rule conflict detection mechanism and verification methods, the applicability and scope of impact of the rule are evaluated to ensure the accuracy and effectiveness of rule execution. The rule knowledge base consists of rules written by business experts and requirements analysts based on domain business knowledge, including rule name, rule description, rule type, and regular expression. 2) The data standard association unit is used to construct prompt text by matching the meaning of category fields with the data standard library, submit it to the large model for intelligent matching, and return whether the field matches the data standard. If it does, the association relationship between the data field and the data standard is automatically established, and the data is converted according to the naming of the data standard and the standard dictionary code. The data standard library includes international standards, ministerial standards, and industry standards, which are periodically obtained from public websites or documents and solidified into the data standard library through manual input or automated program parsing. 3) The conversion rule generation unit is used to generate appropriate conversion rules based on the field type of the exploration results and the mapping relationship with the built-in conversion rule template, such as removing special characters from the name field, uniformly converting the certificate code field to uppercase, and uniformly standardizing the time field. 4) The conversion job execution unit is used to create and run conversion jobs according to business rules and conversion rules, so as to realize the processing of data from the collected data to the converted data and form standardized result data; 5) The conversion status monitoring unit is used to detect the running status of the conversion operation and push the status information to the change sensing module; (7) Data Multidimensional Fusion Module The standardized result data output by the data cleaning and transformation module is used to identify the target table of the theme layer to be merged through the large model, and complete the data fusion. This includes the target table matching unit, field mapping unit, fusion rule configuration unit, and fusion execution unit. 1) The target table matching unit is used to construct prompt text from the standardized table and the target-related information of the topic layer, and submit it to the intelligent analysis and decision module. The large model returns the table feature analysis results; the prompt text is matched and identified with each data table of the topic layer in terms of semantic and structural similarity, and the identification results are returned. The returned identification results include whether the target table matches, the field matching relationship, and the confidence level. 2) The fusion rule configuration unit is used to generate fusion rule configurations on the fields of the theme layer data table based on the results of the target table matching unit; 3) The fusion execution unit is used to generate data fusion jobs according to fusion rules and execute fusion jobs to integrate the standardized result data output by the data cleaning and transformation module into the target table of the theme layer; 4) The fusion status monitoring unit is used to detect the running status of the fusion operation and push the status information to the change perception module.
Citation Information
Patent Citations
Multi-source heterogeneous data mapping method of dynamic ontology semantic fusion model
CN115630066A
Commercial intelligent decision-making question-answering system and method of knowledge graph-driven large model
CN118227767A