A method, system, electronic device, and storage medium for constructing a data warehouse
By matching the correlation between internal and external business data warehouses for new business needs, and using data sets with high correlation or temporary data models to build data warehouses, the problem of low construction rate of traditional data warehouses is solved, and efficient and accurate data warehouse construction is achieved.
Patent Information
- Application Number
- CN202410656623.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-24
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-05-24
AI Technical Summary
Traditional data warehouses are shared by multiple subsidiaries in the face of new business needs, and cannot meet the needs of large-scale data processing.
By obtaining new business demand information, the correlation matching between internal and external business data warehouses is performed. If there is a business data set whose correlation exceeds the preset threshold, the target business data set is built using the data model of the data set. If it does not exist, a temporary data model is selected for construction.
It improves the efficiency and accuracy of data warehouse construction, can promptly respond to new business needs, reduce the difficulty of the ETL process, has high compatibility, and meets the needs of large-scale data processing.
Smart Images

Figure CN118568183B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data warehouses, and in particular to a method, system, electronic device, and storage medium for constructing a data warehouse. Background Art
[0002] A data warehouse is an integrated, subject-oriented, non-volatile data storage area for supporting management decision-making and enterprise analysis. It is a collection for storing a large amount of historical data, which has been cleaned, transformed, and integrated to meet the analysis needs of internal and external users of an enterprise.
[0003] For a traditional data warehouse, when a new business requirement is put forward by an enterprise, generally, data is retrieved from different systems again, and all data is subjected to an ETL process and stored in the data warehouse according to a unified format and standard. Then, a data model is constructed according to the business requirements and data characteristics. After that, through data governance and security management mechanisms, unified control of the data warehouse is carried out, and then visual display is performed according to the requirements.
[0004] However, when multiple subsidiaries of an enterprise share a data warehouse and a new business requirement volume is put forward, it is necessary to re-pull tables to construct a data model. Among them, there may be a large number of models with the same business requirements. The repeated construction process results in a low construction rate of the data warehouse and cannot meet the requirements of large-scale data processing. Therefore, further improvement is needed. Summary of the Invention
[0005] The purpose of this application is to provide a method, system, electronic device, and storage medium for constructing a data warehouse, which has the characteristics of being able to respond to new business requirements in a timely manner to construct a data warehouse and improving the construction efficiency.
[0006] In a first aspect, this application provides a method for constructing a data warehouse, adopting the following technical solutions:
[0007] A method for constructing a data warehouse includes:
[0008] Obtain new business requirement information, where the new business requirement information includes a subsidiary code, and retrieve the associated internal business data warehouse according to the subsidiary code;
[0009] Perform a relevance match between the new business requirement information and each business data set in the internal business data warehouse to determine whether there is a business data set with a correlation degree exceeding a preset threshold;
[0010] If there is a business data set with a correlation degree exceeding the preset threshold, filter out the basic business data set according to the preset screening rule, and construct the target business data set by using the data model included in the basic business data set;
[0011] If there is no business data set with a correlation degree exceeding the preset threshold, the new business requirement information is correlated and matched with each business data set in the external business data warehouse related to other subsidiary codes to determine whether there is a business data set with a correlation degree exceeding the preset threshold;
[0012] If there is a business data set with a correlation degree exceeding the preset threshold, a temporary data model is selected according to the preset imitation rule, and the target business data set is constructed by using the temporary data model.
[0013] By adopting the above technical solution, first match the new business requirements with the internal business data warehouse, and then match with the external business data warehouse. This method can effectively utilize the construction logic of the existing business data sets to construct the target business data set, thereby improving the efficiency of constructing the data warehouse and promptly responding to new business requirements. Among them, the internal and external business data warehouses are matched successively. Distributed matching is conducive to improving the retrieval and matching efficiency. Since the data involved in the internal business data warehouse has a high consistency with the data involved in the new business requirements, therefore, retrieve in the internal business data warehouse to find a business data set with a high matching degree. The consistency of the data is more conducive to reducing the difficulty of the ETL process. Then select the data model. This process also realizes automatic screening, directly using the data models in the existing business data sets, so the compatibility is relatively high, greatly improving the efficiency of constructing the data warehouse, and does not require developers to re-pull the table to select the model and then construct the business data set, which is conducive to meeting the large-scale data processing requirements.
[0014] Optionally, in the step of, if there is a business data set with a correlation degree exceeding the preset threshold, screening out the basic business data set according to the preset screening rule and constructing the target business data set by using the data model included in the basic business data set, includes:
[0015] If there is a business data set with a correlation degree exceeding the preset threshold, test the complexity of data integration and transformation of the data involved in the new business requirement information to obtain a test score;
[0016] Add the test score to the correlation degree, screen out the business data set with the highest correlation degree to obtain the basic business data set;
[0017] Construct the target business data set by using the data model included in the basic business data set.
[0018] By adopting the above technical solution, there are multiple business data sets exceeding the preset threshold. The data involved in the new business requirement information is successively integrated and transformed according to the ETL methods of these business data sets, and then the complexity of the new business requirement information is tested. Then, the business data set with the highest correlation degree is selected as the basic business data set. At this time, the selected basic business data set has the lowest complexity, which is beneficial to ensuring the smooth operation after the subsequent construction of the target business data set.
[0019] Optionally, after the step of increasing the test score to the correlation degree and screening out the business data set with the highest correlation degree, the following steps are further included:
[0020] Copy the business data set with the highest correlation degree to obtain a temporary business data set;
[0021] Compare the temporary business data set with the data involved in the new business requirement information, and delete the irrelevant data in the temporary business data set to obtain the basic business data set.
[0022] By adopting the above technical solution, copying the business data set with the highest correlation degree as the temporary business data set, and performing data comparison and screening for the new business requirement information can achieve precise data processing. Furthermore, it can ensure that the basic business data set contains data related to the new requirements, thus better meeting the needs of business analysis and applications. It can also reduce the cost and complexity of data processing, save processing time and resources, improve the efficiency of data processing, reduce data noise and redundancy, and improve the credibility and effectiveness of data analysis and applications.
[0023] Optionally, the step of, if there is a business data set with a correlation degree exceeding the preset threshold, selecting a temporary data model according to the preset imitation rule and using the temporary data model to construct the target business data set includes:
[0024] If there is a business data set with a correlation degree exceeding the preset threshold, select the business data set with a correlation degree exceeding the preset threshold to form an external business data cluster;
[0025] Count the data model types of the external business data cluster, and arrange the data model types in priority according to the preset arrangement rule;
[0026] Collect sample data from the corresponding system according to the new business requirement information, preprocess the sample data, and then successively select the model types in the order of priority to construct a temporary data model, and perform performance testing on the temporary data model. When the temporary data model meets the standard, select the currently qualified temporary data model to construct the target business data set.
[0027] By adopting the above technical solution, setting priorities for the types of data models involved in the external business data cluster and arranging them according to the priorities can improve the accuracy of model selection. Then, build temporary data models in sequence according to the arrangement order, and then conduct performance tests on the temporary models. Once a temporary data model meets the standards, select the currently qualified temporary data model to build the target business data set. This method can accelerate the selection of a data model set suitable for new business requirements.
[0028] Optionally, the step of statistically analyzing the types of data models in the external business data cluster and arranging the priorities of the data model types according to the preset arrangement rules includes:
[0029] Statistically analyze the number of data models of the same type in the external business data cluster, and arrange the priorities of the data model types in descending order according to the number.
[0030] By adopting the above technical solution, arranging the priorities of the data model types in descending order of the number helps to improve the utilization rate of the data models. Priority is given to processing the data model types with a larger number, and the utilization rate is higher, so that the data information provided by these models can be more fully utilized to meet business requirements. For the data model types with a smaller number, their processing priorities can be appropriately reduced to avoid affecting the processing efficiency of the overall business data.
[0031] Optionally, the step of, if there is no business data set with a correlation degree exceeding the preset threshold, performing a correlation match between the new business requirement information and each business data set in the external business data warehouse related to the codes of other subsidiaries, and determining whether there is a business data set with a correlation degree exceeding the preset threshold further includes:
[0032] If there is a business data set with a correlation degree exceeding the preset threshold, compare the company type labels of the subsidiaries to which the codes in the new business requirement information belong with the company type labels of the subsidiaries to which the codes of other subsidiaries belong, screen out the codes of other subsidiaries with the same company type labels, and then perform a correlation match between the new business requirement information and each business data set in the external business data warehouse related to the codes of other subsidiaries with the same company type labels, and determine whether there is a business data set with a correlation degree exceeding the preset threshold.
[0033] By adopting the above technical solution, the accuracy of building the target business data set can be improved, and the efficiency of building the target business data set can also be effectively improved.
[0034] In a second aspect, the present application provides a data warehouse construction system, adopting the following technical solution:
[0035] A data warehouse construction system includes:
[0036] Information acquisition module: used to acquire new business requirement information, where the new business requirement information includes a subsidiary code, and retrieve the associated internal business data warehouse according to the subsidiary code;
[0037] Internal matching module: used to perform relevance matching between the new business requirement information and each business data set in the internal business data warehouse, and determine whether there is a business data set with an association degree exceeding a preset threshold;
[0038] Internal construction module: If there is a business data set with an association degree exceeding the preset threshold, filter out the basic business data set according to the preset screening rules, and construct the target business data set using the data model included in the basic business data set;
[0039] External matching module: If there is no business data set with an association degree exceeding the preset threshold, perform relevance matching between the new business requirement information and each business data set in the external business data warehouse related to other subsidiary codes, and determine whether there is a business data set with an association degree exceeding the preset threshold;
[0040] External construction module: If there is a business data set with an association degree exceeding the preset threshold, select a temporary data model according to the preset imitation rules, and construct the target business data set using the temporary data model.
[0041] Thirdly, the present application provides an electronic device, adopting the following technical solution:
[0042] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned data warehouse construction method are implemented.
[0043] Fourthly, the present application provides a computer storage medium, adopting the following technical solution:
[0044] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned data warehouse construction method are implemented.
[0045] In summary, the beneficial technical effects included in the present application are:
[0046] First, match the new business requirements with the internal business data warehouse, and then with the external business data warehouse. This method can effectively utilize the construction logic of the existing business data set to construct the target business data set, thereby improving the efficiency of constructing the data warehouse and promptly responding to new business requirements. Among them, by performing the matching of the internal and external business data warehouses successively, distributed matching is conducive to improving the retrieval and matching efficiency. Since the data involved in the internal business data warehouse has a high consistency with the data involved in the new business requirements, therefore, retrieve in the internal business data warehouse to find business data sets with high matching degrees. The consistency of the data is more conducive to reducing the difficulty of the ETL process. Then, select the data model. This process also realizes automatic screening and directly uses the data models in the existing business data sets, so the compatibility is relatively high, greatly improving the efficiency of constructing the data warehouse. It does not require developers to pull tables again to select models and then construct business data sets, which is conducive to meeting the large-scale data processing requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is a schematic flowchart of a method for constructing a data warehouse in one embodiment of the present application.
[0048] Figure 2 It is a schematic flowchart of the specific steps of step S20 in another embodiment of the present application.
[0049] Figure 3 It is a schematic flowchart of the steps after step S201 in another embodiment of the present application.
[0050] Figure 4 It is a schematic flowchart of the specific steps of step S210 in another embodiment of the present application.
[0051] Figure 5 It is a schematic structural diagram of a data warehouse construction system in one embodiment of the present application.
[0052] Figure 6 It is a schematic block diagram of the principle of an electronic device in one embodiment of the present application.
[0053] In the figure, 1. Information acquisition module; 2. Internal matching module; 3. Internal construction module; 4. External matching module; 5. External construction module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments.
[0055] In the description of the embodiments of the present application, words such as "for example" or "for illustration" are used to give examples, illustrations or explanations. Any embodiment or design solution described as "for example" or "for illustration" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "for example" or "for illustration" is intended to present the relevant concepts in a specific manner.
[0056] In the description of the embodiments of the present application, the term "plurality" means two or more. For example, a plurality of systems means two or more systems, and a plurality of screen terminals means two or more screen terminals. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly indicating the technical features indicated. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0057] The embodiments of the present application provide a method for constructing a data warehouse. In one embodiment, please refer to Figure 1 , Figure 1 is a schematic flowchart of a method for constructing a data warehouse provided by the embodiments of the present application. This method can be implemented depending on a computer program, which can be integrated into an application or run as an independent tool-class application. This method can also be implemented depending on a single-chip microcomputer and can also run in a data warehouse construction system. Specifically, this method may include the following steps:
[0058] S1. Obtain new business requirement information, where the new business requirement information includes a subsidiary code, and retrieve the associated internal business data warehouse according to the subsidiary code.
[0059] Among them, there are multiple subsidiaries in the business model of the enterprise that share a database system, but the subsidiaries are responsible for their own profits and losses. Therefore, the data storage of each subsidiary needs to be independent of each other in order to be managed independently. For example, each subsidiary has its own financial analysis, sales data analysis, customer behavior analysis system, etc. For the head office, sharing a database system is conducive to the head office's management of subsidiaries and facilitates timely adjustments to business strategies. When a subsidiary proposes a new business analysis requirement, the developer can analyze the requirements and integrate the data involved in the demand analysis. For example, a subsidiary has a new business requirement for customer behavior analysis, and the data involved may include basic customer information, sales data, geographic location information, feedback comment data, etc. The new business analysis requirement also includes the subsidiary code, which is a unique identification code corresponding to each subsidiary, and is used to retrieve the internal business data warehouse related to the subsidiary in the data warehouse. The internal business data warehouse includes multiple business data sets, and one business requirement corresponds to one business data set.
[0060] S2. Perform correlation matching between the new business demand information and each business data set in the internal business data warehouse to determine whether there is a business data set whose correlation exceeds a preset threshold.
[0061] Among them, when the data involved in the new business demand information contains an existing business data set, the correlation is strongest. For example, the business demand for customer behavior analysis is a new business demand, and the business data set for sales data analysis is an existing business data set. The data involved in the business demand for customer behavior analysis contains the data involved in the sales data analysis business data set, and the two are most strongly correlated.
[0062] S20: If there is a business data set whose correlation exceeds a preset threshold, filter out the basic business data set according to a preset filtering rule, and construct a target business data set using the data model included in the basic business data set.
[0063] Specifically, by comparing and matching the data consistency between the new business requirement information and each business data set in the internal business data warehouse, the correlation degree is obtained, and it is judged whether there is a business data set with a correlation degree exceeding the preset threshold. If there is a business data set with a correlation degree exceeding the preset threshold, the basic business data set is screened out according to the preset screening rules, and the data model contained in the basic business data set is used as the basis for generating the target business data set for the new business requirement to construct the target business data set. Above, on the basis of automatically matching the appropriate business data set, the existing data model is used to construct the new business data set, effectively utilizing the development basis of the previous business data set. Since the preset threshold for the correlation degree is set relatively high, most of the data ETL processes, that is, data integration, cleaning, and transformation, in the new business requirement information can be directly based on the existing business data set. The process of selecting the model is also automated, directly using the data model in the existing business data set, so the compatibility is relatively high, greatly improving the efficiency of constructing the data warehouse. It is not necessary for developers to re-pull the table to select the model and then construct the business data set, which is conducive to meeting the large-scale data processing requirements.
[0064] S21. If there is no business data set with a correlation degree exceeding the preset threshold, the new business requirement information is correlated and matched with each business data set in the external business data warehouse related to the codes of other subsidiaries, and it is judged whether there is a business data set with a correlation degree exceeding the preset threshold.
[0065] Specifically, the business data warehouse related to the codes of other subsidiaries becomes the external business data warehouse. If there is no business data set with a correlation degree exceeding the preset threshold, the new business requirement information is correlated and matched with each business data set in the external business data warehouse, and it is judged again whether there is a business data set with a correlation degree exceeding the preset threshold. This application first performs matching in the internal business data warehouse and then in the external business data warehouse. Distributed matching is conducive to improving the retrieval and matching efficiency. And because the data involved in the internal business data warehouse has a relatively high consistency with the data involved in the new business requirement, therefore, by retrieving in the internal business data warehouse to find the business data set with a high matching degree, the data consistency is more conducive to reducing the difficulty of the ETL process and improving the efficiency of constructing the target business data set.
[0066] S210. If there is a business data set with a correlation degree exceeding the preset threshold, a temporary data model is selected according to the preset imitation rules, and the temporary data model is used to construct the target business data set.
[0067] Specifically, when there is a business data set in the external business warehouse with a connection degree exceeding the preset threshold, a temporary data model is selected from the business data set, and then the temporary model is used to construct the target business data set. Since there are certain deviations between the data involved in the business data set of the external business warehouse and the data involved in the new business requirements, the data model of the external business warehouse can be borrowed, and on the basis of the basic logic of this temporary data model, the target business data set is constructed by imitation, which is also conducive to improving the efficiency of constructing the data warehouse.
[0068] S211. If there is no business data set with a connection degree exceeding the preset threshold, a signal requiring manual creation is sent to the development end.
[0069] Specifically, when there is no business data set with a connection degree exceeding the preset threshold, it indicates that the new business requirements have not been constructed yet. Therefore, it is impossible to utilize the existing business data set basis, and a corresponding signal can be sent to the development end.
[0070] Refer to Figure 2 , on the basis of the above embodiments, as an optional embodiment, step S20: If there is a business data set with a connection degree exceeding the preset threshold, the basic business data set is screened according to the preset screening rules, and the target business data set is constructed by using the data model included in the basic business data set, including:
[0071] S200. If there is a business data set with a connection degree exceeding the preset threshold, the complexity of data integration and conversion of the data involved in the new business requirement information is tested to obtain a test score.
[0072] S201. The test score is added to the connection degree, and the business data set with the highest connection degree is screened out to obtain the basic business data set.
[0073] S202. The target business data set is constructed by using the data model included in the basic business data set.
[0074] As described above, if there are business data sets with a correlation degree exceeding the preset threshold and there are multiple business data sets exceeding the preset threshold, the ETL processes involved in the construction of these business data sets are successively used in the new business requirement information data, that is, the data related to the new business requirement information is successively integrated and transformed according to the ETL methods of these business data sets, and then the complexity of the new business requirement information is tested. The complexity can be evaluated according to the conversion time and memory consumption, etc. After obtaining the test score, the test score is presented in the form of a negative number. The higher the complexity, the smaller the value of the test score. Then, the test score is increased to the correlation degree, the business data set with the highest correlation degree is selected to obtain the basic business data set, and then the target business data set is constructed by using the data model included in the basic business data set. At this time, the selected basic business data set has the lowest complexity, which is beneficial to the smooth construction and operation of the target business data set.
[0075] Refer to Figure 3 , on the basis of the above embodiments, as an optional embodiment, after step S201: increasing the test score to the correlation degree and selecting the business data set with the highest correlation degree, it further includes:
[0076] S2010. Copy the business data set with the highest correlation degree to obtain a temporary business data set.
[0077] S2011. Compare the temporary business data set with the data related to the new business requirement information, and delete the irrelevant data in the temporary business data set to obtain the basic business data set.
[0078] Among them, by copying the business data set with the highest correlation degree as the temporary business data set and performing data comparison and screening for the new business requirement information, precise processing of the data can be realized, thereby ensuring that the basic business data set contains data related to the new requirements, better meeting the needs of business analysis and applications, reducing the cost and complexity of data processing, saving processing time and resources, improving the efficiency of data processing, reducing data noise and redundancy, and improving the credibility and effectiveness of data analysis and applications.
[0079] Refer to Figure 4 , on the basis of the above embodiments, as an optional embodiment, step S210: if there is a business data set with a correlation degree exceeding the preset threshold, select a temporary data model according to the preset imitation rule, and use the temporary data model to construct the target business data set, including:
[0080] S2101. If there is a business data set with a correlation degree exceeding the preset threshold, select the business data set with a correlation degree exceeding the preset threshold to form an external business data cluster.
[0081] S2102. Statistically analyze the data model types in the external business data cluster, and prioritize the data model types according to the preset arrangement rules.
[0082] S2103. Collect sample data from the corresponding system according to the new business requirement information, preprocess the sample data, then sequentially select model types according to the priority arrangement to construct a temporary data model, and perform performance testing on the temporary data model. When a temporary data model meets the standard, select the currently qualified temporary data model to construct the target business data set.
[0083] As described above, by setting priorities for the data model types involved in the external business data cluster and arranging them according to the priorities, the accuracy of model selection can be improved. Then, construct a temporary data model sequentially according to the arrangement order. The sample data designed for the temporary data model is limited and the data volume is not large. Then, perform performance testing on the temporary model. Once a temporary data model meets the standard, select the currently qualified temporary data model to construct the target business data set. This method can accelerate the selection of a data model suitable for the new business requirements. Among them, performance testing involves multiple aspects, such as data integrity, data accuracy, response time, query efficiency, and resource consumption testing. Conduct a score evaluation comprehensively considering multiple aspects. Once the score meets the standard, select the currently qualified temporary data model to construct the target business data set. If all the temporary data models do not meet the standard, send a signal indicating the need for manual creation to the development end.
[0084] Based on the above embodiments, as an alternative embodiment, step S2102: The step of statistically analyzing the data model types in the external business data cluster and prioritizing the data model types according to the preset arrangement rules includes:
[0085] Statistically count the number of data models of the same type in the external business data cluster, and prioritize the data model types in descending order according to the quantity.
[0086] As described above, prioritizing the data model types in descending order of quantity helps to improve the utilization rate of the data models. Prioritize the processing of data model types with a larger quantity, which have a higher usage rate, and can make more full use of the data information provided by these models to meet business requirements. For data model types with a smaller quantity, their processing priority can be appropriately reduced to avoid affecting the processing efficiency of the overall business data.
[0087] Based on the above embodiments, as an alternative embodiment, step S21: If there is no business data set with a correlation degree exceeding a preset threshold, then the step of performing a correlation match between the new business requirement information and each business data set in the external business data warehouse related to other subsidiary codes to determine whether there is a business data set with a correlation degree exceeding the preset threshold further includes:
[0088] If there is a business data set with a correlation degree exceeding the preset threshold, then compare the company type labels of the subsidiary codes in the new business requirement information with the company type labels of other subsidiary codes, screen out the other subsidiary codes with the same company type labels, and then perform a correlation match between the new business requirement information and each business data set in the external business data warehouse related to the other subsidiary codes with the same company type labels to determine whether there is a business data set with a correlation degree exceeding the preset threshold.
[0089] As described above, corresponding attribute labels are assigned to the company types of the subsidiaries. Subsidiaries with the same company type label have highly similar businesses, and thus, highly similar data types. Therefore, by first performing a correlation match on the external business data warehouses corresponding to the subsidiaries with the same company type label, some subsidiaries with overly large differences in business types can be screened out, further improving the accuracy of constructing the target business data set and also effectively improving the efficiency of constructing the target business data set.
[0090] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not indicate the order of execution. The order of execution of each process should be determined based on its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0091] The embodiments of the present application further provide a data warehouse construction system, which corresponds one-to-one with a data warehouse construction method in the embodiments. Refer to Figure 5 and this data warehouse construction system includes: an information acquisition module 1, an internal matching module 2, an internal construction module 3, an external matching module 4, and an external construction module 5. The detailed descriptions of each functional module are as follows:
[0092] The information acquisition module 1: is used to acquire new business requirement information, which includes subsidiary codes, and retrieve the related internal business data warehouse according to the subsidiary codes.
[0093] The internal matching module 2: is used to perform a correlation match between the new business requirement information and each business data set in the internal business data warehouse to determine whether there is a business data set with a correlation degree exceeding a preset threshold.
[0094] Internal construction module 3: If there is a business data set with an association degree exceeding the preset threshold, the basic business data set is screened out according to the preset screening rules, and the target business data set is constructed by using the data model included in the basic business data set.
[0095] External matching module 4: If there is no business data set with an association degree exceeding the preset threshold, the new business requirement information is matched with each business data set in the external business data warehouse related to the codes of other subsidiaries to determine whether there is a business data set with an association degree exceeding the preset threshold.
[0096] External construction module 5: If there is a business data set with an association degree exceeding the preset threshold, a temporary data model is selected according to the preset imitation rules, and the target business data set is constructed by using the temporary data model.
[0097] Among them, the system can obtain the new business requirement information through the information acquisition module 1. The new business requirement information includes the subsidiary code. The relevant internal business data warehouse is retrieved according to the subsidiary code, and then the internal matching module 2 is used to match the new business requirement information with each business data set in the internal business data warehouse to determine whether there is a business data set with an association degree exceeding the preset threshold. If there is a business data set with an association degree exceeding the preset threshold, the internal construction module 3 is used to screen out the basic business data set according to the preset screening rules, and the target business data set is constructed by using the data model included in the basic business data set. If there is no business data set with an association degree exceeding the preset threshold, the external matching module 4 is used to match the new business requirement information with each business data set in the external business data warehouse related to the codes of other subsidiaries to determine whether there is a business data set with an association degree exceeding the preset threshold. If there is a business data set with an association degree exceeding the preset threshold, the external construction module 5 is used to select a temporary data model according to the preset imitation rules, and the target business data set is constructed by using the temporary data model.
[0098] For the specific limitations of the data warehouse construction system, reference can be made to the limitations on the data warehouse construction method in the context, which will not be elaborated here. Each module in the above data warehouse construction system can be implemented in whole or in part by software, hardware and their combination. The above modules can be embedded in the processor of the electronic device in the form of hardware or independent of it, or stored in the memory of the electronic device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0099] In this embodiment, an electronic device is provided, and the electronic device is a computer. Refer to Figure 6, the electronic device includes a processor, a memory, a network interface, and a database connected via a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the electronic device is used to store detection data tables. The network interface of the electronic device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it implements a data warehouse construction method.
[0100] In one embodiment, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:
[0101] S1. Obtain new business requirement information, which includes a subsidiary code, and retrieve the associated internal business data warehouse according to the subsidiary code.
[0102] S2. Perform a relevance match between the new business requirement information and each business data set in the internal business data warehouse, and determine whether there is a business data set with a relevance degree exceeding a preset threshold.
[0103] S20. If there is a business data set with a relevance degree exceeding the preset threshold, filter out the basic business data set according to the preset filtering rules, and construct the target business data set by using the data model included in the basic business data set.
[0104] S21. If there is no business data set with a relevance degree exceeding the preset threshold, perform a relevance match between the new business requirement information and each business data set in the external business data warehouse related to other subsidiary codes, and determine whether there is a business data set with a relevance degree exceeding the preset threshold.
[0105] S210. If there is a business data set with a relevance degree exceeding the preset threshold, select a temporary data model according to the preset imitation rules, and construct the target business data set by using the temporary data model.
[0106] S211. If there is no business data set with a relevance degree exceeding the preset threshold, send a signal requiring manual creation to the development end.
[0107] Step S20 specifically includes the following steps:
[0108] S200. If there is a business data set with a relevance degree exceeding the preset threshold, test the complexity of data integration and conversion for the data involved in the new business requirement information to obtain a test score.
[0109] S201. Increase the test score to the correlation level, screen out the business data set with the highest correlation, and obtain the basic business data set.
[0110] S202. Use the data model included in the basic business data set to construct the target business data set.
[0111] After step S201, the following steps are further included:
[0112] S2010. Copy the business data set with the highest correlation to obtain a temporary business data set.
[0113] S2011. Compare the temporary business data set with the data involved in the new business requirement information, and delete the irrelevant data in the temporary business data set to obtain the basic business data set.
[0114] Step S210 specifically includes the following steps:
[0115] S2101. If there is a business data set whose correlation exceeds the preset threshold, select the business data set whose correlation exceeds the preset threshold to form an external business data cluster.
[0116] S2102. Count the data model types of the external business data cluster, and arrange the data model types in priority according to the preset arrangement rules.
[0117] S2103. Collect sample data from the corresponding system according to the new business requirement information, preprocess the sample data, then sequentially select the model types in the order of priority to construct a temporary data model, and perform a performance test on the temporary data model. When the temporary data model meets the standard, select the currently qualified temporary data model to construct the target business data set.
[0118] The embodiment of the present application also discloses a computer-readable storage medium storing a computer program that can be loaded and executed by a processor. When the computer program is executed by the processor, the steps of any of the above data warehouse construction methods are implemented, and the same effect can be achieved.
[0119] Among them, the computer-readable storage medium includes, for example: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.
[0120] The above are all preferred embodiments of the present application, and do not limit the protection scope of the present application accordingly. Any feature disclosed in this specification (including the abstract and drawings), unless specifically described, can be replaced by other equivalent or similar-purpose alternative features. That is, unless specifically described, each feature is only an example of a series of equivalent or similar features.
Claims
1. A method for constructing a data warehouse, characterized in that: include: Acquire new business demand information, wherein the new business demand information includes a subsidiary code, and retrieve the associated internal business data warehouse according to the subsidiary code; Perform correlation matching between new business demand information and various business data sets in the internal business data warehouse to determine whether there is a business data set with a correlation exceeding a preset threshold; If there is a business data set whose correlation exceeds the preset threshold, the basic business data set is filtered out according to the preset filtering rules, and the target business data set is constructed using the data model contained in the basic business data set, including: If there is a business data set whose correlation exceeds the preset threshold, the complexity of data integration and conversion of the data involved in the new business demand information is tested to obtain a test score; The test score is increased to the correlation degree, and the business data set with the highest correlation degree is selected to obtain the basic business data set; Utilize the data model contained in the basic business data set to build the target business data set; If there is no business data set whose correlation exceeds the preset threshold, the new business demand information is matched with each business data set in the external business data warehouse related to other subsidiary codes to determine whether there is a business data set whose correlation exceeds the preset threshold, further comprising: if there is a business data set whose correlation exceeds the preset threshold, the company type label to which the subsidiary code in the new business demand information belongs is compared with the company type label to which other subsidiary codes belong, and other subsidiary codes with the same company type label are screened out, and then the new business demand information is matched with each business data set in the external business data warehouse related to other subsidiary codes with the same company type label to determine whether there is a business data set whose correlation exceeds the preset threshold; If there is a business data set whose correlation exceeds a preset threshold, a temporary data model is selected according to a preset imitation rule, and the temporary data model is used to construct a target business data set.
2. The method according to claim 1, characterized in that: After the step of increasing the test score to the correlation degree and selecting the business data set with the highest correlation degree, the method further includes: Copy the business data set with the highest correlation to obtain a temporary business data set; The temporary business data set is compared with the data involved in the new business demand information, and irrelevant data in the temporary business data set is deleted to obtain the basic business data set.
3. The method according to claim 1, characterized in that If there is a business data set with a correlation exceeding a preset threshold, the step of selecting a temporary data model according to a preset imitation rule and using the temporary data model to construct a target business data set includes: If there are business data sets whose correlation exceeds the preset threshold, then the business data sets whose correlation exceeds the preset threshold are selected to form an external business data cluster; Collect statistics on the data model types of external business data clusters and prioritize the data model types according to preset arrangement rules; According to the new business demand information, sample data is collected from the corresponding system, pre-processed, and then the model types are selected in order of priority to build a temporary data model. The temporary data model is performance tested. When a temporary data model meets the standards, the current temporary data model that meets the standards is selected to build the target business data set.
4. The method according to claim 3, characterized in that: The step of counting the data model types of the external business data cluster and prioritizing the data model types according to a preset arrangement rule includes: Count the number of data models of the same type in the external business data cluster, and prioritize the data model types in descending order based on the number.
5. A data warehouse construction system, characterized in that: include: Information acquisition module (1): used to acquire new business demand information, wherein the new business demand information includes a subsidiary code, and retrieve the associated internal business data warehouse according to the subsidiary code; Internal matching module (2): used to match the new business demand information with each business data set in the internal business data warehouse, and determine whether there is a business data set whose correlation exceeds a preset threshold; Internal construction module (3): If there is a business data set whose correlation exceeds a preset threshold, the basic business data set is filtered out according to the preset filtering rules, and the target business data set is constructed using the data model contained in the basic business data set, including: If there is a business data set whose correlation exceeds the preset threshold, the complexity of data integration and conversion of the data involved in the new business demand information is tested to obtain a test score; The test score is increased to the correlation degree, and the business data set with the highest correlation degree is selected to obtain the basic business data set; Utilize the data model contained in the basic business data set to build the target business data set; External matching module (4): if there is no business data set with a correlation exceeding a preset threshold, then matching the new business demand information with each business data set in the external business data warehouse related to other subsidiary codes, and determining whether there is a business data set with a correlation exceeding the preset threshold, further comprising: if there is a business data set with a correlation exceeding the preset threshold, then comparing the company type label to which the subsidiary code in the new business demand information belongs with the company type label to which the other subsidiary codes belong, screening out other subsidiary codes with the same company type label, and then matching the new business demand information with each business data set in the external business data warehouse related to the other subsidiary codes with the same company type label, and determining whether there is a business data set with a correlation exceeding the preset threshold; External construction module (5): If there is a business data set whose correlation exceeds a preset threshold, a temporary data model is selected according to a preset imitation rule, and the temporary data model is used to construct the target business data set.
6. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and execute the data warehouse construction method according to any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that: The computer program that can be loaded by a processor and execute the data warehouse construction method according to any one of claims 1 to 4 is stored.
Citation Information
Patent Citations
Collaborative generation method, device and equipment of data model
CN116483797A
Index generation method and device, computer equipment and storage medium
CN116841505A