Method and device for data warehouse, equipment, medium and product
By obtaining metadata and business knowledge of the data table and combining upstream indicator caliber feeding into the language model, the problem of inaccurate inference of data warehouse indicator caliber in the existing technology is solved, and the automated generation and accuracy of multi-layer indicator caliber is achieved.
Patent Information
- Application Number
- CN202510593076.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-08-15
AI Technical Summary
When the prior art infers the indicator caliber of a data warehouse, it only focuses on the indicators in a single data table, making it difficult to fully reflect the core information of the indicators on the upstream processing link, and the language model cannot understand the business characteristics, resulting in the inferred indicator caliber information inaccurate.
By obtaining metadata information and business knowledge of the data table, and feeding it into the language model with upstream indicator caliber, it realizes the automated generation of multi-layer indicator caliber, and promotes the language model to understand the business characteristics of the data warehouse.
The information accumulation of the index caliber along the index processing link is realized, the inference accuracy of the index caliber is enhanced, and the key information on the link is fully reflected.
Smart Images

Figure CN120492548A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate generally to the field of computers, and in particular to methods and apparatuses, devices, media, and products for data warehousing. Background Art
[0002] In recent years, language models have rapidly developed and become an influential and highly sought-after technology. Large language models (LLMs), a representative example of language models, are widely used for their powerful natural language understanding and generalization capabilities. One of the core functions of language models is to significantly simplify the previously complex and cumbersome knowledge acquisition process. Users can overcome previously high knowledge barriers and achieve functionality comparable to that of difficult-to-use specialized tools, thereby lowering the entry threshold and learning costs. Language models are increasingly being used in the big data field. Using natural language as a bridge, users can interact with data more conveniently and intuitively through language models, enabling data analysis and processing. Summary of the Invention
[0003] An embodiment of the present disclosure provides a solution for a data warehouse.
[0004] In a first aspect of the present disclosure, a method for a data warehouse is provided, the method comprising obtaining metadata information and business knowledge of a data table, the metadata information comprising table association data indicating association relationships between a plurality of data tables corresponding to the data warehouse. The method further comprises obtaining upstream indicator calibers of upstream data tables associated with the data table based on the table association data. The method further comprises inferring the indicator caliber of the data table by feeding the obtained metadata information and business knowledge, as well as the upstream indicator caliber, into a language model to describe indicators of the data warehouse associated with the data table.
[0005] In a second aspect of the present disclosure, a device for a data warehouse is provided, the device including a data and knowledge acquisition module configured to acquire metadata information and business knowledge of a data table, the metadata information including table association data indicating an association relationship between multiple data tables corresponding to the data warehouse. The device also includes an upstream caliber acquisition module configured to acquire the upstream indicator caliber of an upstream data table associated with the data table based on the table association data. The device also includes a caliber inference module configured to infer the indicator caliber of the data table by feeding the acquired metadata information and business knowledge, as well as the upstream indicator caliber, into a language model to describe the indicators of the data warehouse associated with the data table.
[0006] According to a third aspect of the present disclosure, an electronic device is provided. The computing device includes a processor and a memory, wherein the memory stores instructions that, when executed by the processor, cause the processor to perform a method or process according to an embodiment of the present disclosure.
[0007] According to a fourth aspect of the present disclosure, a machine-readable storage medium is provided, wherein machine-executable instructions are stored on the machine-readable storage medium, and when executed by a processor, the machine-executable instructions cause the processor to perform a method or process according to an embodiment of the present disclosure.
[0008] In a fifth aspect of the present disclosure, a computer program product is provided, which is tangibly stored on a non-transitory computer-readable storage medium and includes a computer program that, when executed by a processor of a computer, causes the processor to perform a method or process according to an embodiment of the present disclosure.
[0009] Please note that the invention summary is provided to introduce a series of concepts in a simplified form, which will be further described in the detailed description below. The invention summary is not intended to identify key features or essential features of the disclosure, nor is it intended to limit the scope of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other objects, features and advantages of the present disclosure will become more apparent through a more detailed description of the embodiments of the present disclosure in conjunction with the accompanying drawings, in which:
[0011] Figure 1 is a diagram schematically illustrating an example environment in which methods and / or processes according to embodiments of the present disclosure may be implemented;
[0012] Figure 2 is a flowchart schematically illustrating a method for a data warehouse according to an embodiment of the present disclosure;
[0013] Figure 3A A diagram schematically illustrating an exemplary association relationship between data tables according to an embodiment of the present disclosure;
[0014] Figure 3B A diagram schematically illustrates an exemplary example of an indicator processing link according to an embodiment of the present disclosure;
[0015] Figure 4 is a diagram schematically illustrating a schematic process of indicator caliber reasoning based on a language model according to an embodiment of the present disclosure;
[0016] Figure 5 is a diagram schematically illustrating an exemplary link execution order of metric-caliber reasoning according to an embodiment of the present disclosure;
[0017] Figure 6 is a diagram schematically illustrating an apparatus for a data warehouse according to an embodiment of the present disclosure;
[0018] Figure 7 is a schematic block diagram of an example device that can be used to implement embodiments according to the present disclosure.
[0019] Throughout the drawings, same or similar reference numbers generally refer to same or similar elements. DETAILED DESCRIPTION
[0020] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0021] In the description of the embodiments of the present disclosure, the term "including" and its variations should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. can refer to different or the same objects, unless explicitly indicated to be different.
[0022] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0023] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0024] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0025] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0026] As mentioned above, language modeling technology can be applied to facilitate data analysis and processing. Language models such as LLM can be built based on deep learning techniques. By training on large amounts of natural language data, they can understand the structure, semantics, and grammar of language, thereby generating natural and fluent responses. Incorporating language models into data tasks can improve the ease of use and interactivity of the analysis and processing process.
[0027] In some relevant solutions, a reasoning strategy combined with language models can be used to determine the metric scope of the data warehouse. The metric scope of a data warehouse can refer to the standard specifications for defining and calculating metrics within the data warehouse. It clarifies key elements such as the business meaning of the metric, data source, data scope, calculation method, and data processing process, and is an important basis for ensuring the accuracy, consistency, and availability of data within the data warehouse.
[0028] However, these current solutions still have many shortcomings and face challenges in practical application. First, these solutions only focus on the indicators in a single data table when inferring the indicator caliber of the data warehouse. If you only look at the single-layer information of the indicator itself, it is difficult to fully reflect the core information of the indicator in the upstream processing link. In the indicator processing link of the data warehouse, the key calculation logic of the indicator is often implemented in the middle-layer table. If you reason separately for the downstream data table, you cannot see the processing logic of the middle layer. Furthermore, only using the technical data of the data warehouse itself (for example, metadata indicating the table structure or indicator calculation logic) for model training and reasoning makes it impossible for the language model to understand and consider business characteristics, resulting in the inferred indicator caliber information not meeting the accuracy requirements.
[0029] In order to at least solve at least some of the above-mentioned and other potential problems, an embodiment of the present disclosure proposes a solution for a data warehouse. The solution includes obtaining metadata information and business knowledge of a data table, and the metadata information includes table association data indicating the association relationship between multiple data tables corresponding to the data warehouse. The solution also includes obtaining the upstream indicator caliber of the upstream data table associated with the data table based on the table association data. The solution also includes inferring the indicator caliber of the data table by feeding the obtained metadata information and business knowledge, as well as the upstream indicator caliber into a language model to describe the indicators of the data warehouse associated with the data table.
[0030] The data warehouse solution according to the embodiments of the present disclosure provides a multi-layered automated generation strategy for indicator calibers for data warehouses. This strategy enables the accumulation of information along the indicator processing chain, thereby fully reflecting the key information along the chain. Furthermore, multimodal indicator caliber reasoning, combining business data with technical data, can facilitate language models' understanding of the business characteristics of data warehouses, thereby enhancing the accuracy of indicator caliber reasoning.
[0031] Reference below Figures 1 to 7 It should be understood that these exemplary embodiments are provided only to enable those skilled in the art to better understand and implement the embodiments of the present disclosure, and are not intended to limit the scope of the present disclosure in any way.
[0032] Figure 1 FIG. 1 is a diagram schematically illustrating an example environment 100 in which methods and / or processes according to embodiments of the present disclosure may be implemented. Figure 1 As shown in , the example environment 100 includes a data warehouse 105, an input 110, a computing system 120 including computing nodes (e.g., a locally disposed computing node 121 and a cloud-disposed computing node 122), and an output 130. Figure 1 The implementation environment of the embodiments of the present disclosure is described using a hybrid computing system as an example. It should be understood that this is merely illustrative and non-restrictive, and that other types of computing systems are also feasible, such as distributed or centralized computing systems. The appropriate computing environment configuration can be selected based on actual usage requirements.
[0033] exist Figure 1 In the figure, only limited components and exemplary connection relationships are shown. It should be understood that this is for the purpose of ease of explanation and illustration and is not intended to limit the scope of the present disclosure, and other different components may also exist. For example, a display component and an input component, etc. By way of example and not limitation, the indicator caliber of the inferred data table can be displayed on the display component, and the indicator caliber of the inferred data table can be edited through the input component.
[0034] According to an embodiment of the present disclosure, the data warehouse 105 can be a subject-oriented data collection that integrates data related to a certain subject to facilitate in-depth analysis of the subject. The data warehouse 105 can be integrated, and the data therein can come from multiple different data sources. These data from different data sources may have different formats, so it is necessary to perform standardization processing (such as data cleaning, conversion, and integration) on them to ensure data consistency.
[0035] The data in data warehouse 105 is primarily used for analysis and decision support and is typically not updated frequently. Once data enters the data warehouse, it remains stable for a certain period of time, facilitating multi-dimensional comparisons and trend analysis. Data warehouse 105 can record historical data values, thereby reflecting how data changes over time. By analyzing historical data, business trends, patterns, and regularities can be discovered, providing a basis for forecasting and decision-making.
[0036] According to an embodiment of the present disclosure, the indicator caliber of the data warehouse 105 is calculated layer by layer by feeding the input 110 associated with the data warehouse 105 into the computing system 120. That is, the indicator caliber of each data table of the data warehouse 105 is calculated. The input 110 includes technical data of the data table, indicating, for example, the table structure of the data table, the organization method of the data, and the type of data. The input 110 also includes business data. By way of example and not limitation, assuming that the data table is about the field of education, the business data of the data table may include course attendance, subject professional terminology and technical background, and keywords of the course or subject.
[0037] Input 110 associated with the data warehouse 105 may be fed into a system having computing capabilities (e.g., Figure 1 ). Based on the input 110 associated with the data warehouse 105, the computing system 120 can be configured to calculate the indicator caliber of the corresponding data table for each of the multiple data tables corresponding to the data warehouse, thereby describing the indicators of the data warehouse 105 associated with the corresponding data table. Examples of data tables may include but are not limited to Hive tables. A language model is configured in the computing system 120, which infers indicator caliber information based on the input 110. The inferred indicator caliber information can be presented as an output 130 to facilitate the user to understand the associated indicators. In the following, the indicator caliber reasoning based on the language model according to the embodiment of the present disclosure will be further described in detail.
[0038] As described above, at least one computing node included in the computing system 120 can perform processing and operations corresponding to the generation of multi-layer indicator calibers according to an embodiment of the present disclosure. By way of example and not limitation, the computing node 121 can be arranged locally at the user, and the computing node 122 can be arranged in the cloud. The cloud can indicate a service model built based on cloud or distributed technology. In this model, computing resources, storage resources, etc. are coupled together through a network to form a schedulable and scalable resource collection. At least a portion of these resources can be dynamically accessed or allocated to complete tasks without the need to locally own or manage these physical resources.
[0039] A computing node may refer to a computing resource and may be a device with computing capabilities. For example, a computing node may be configured with a processor and memory, etc., or may be equipped with a dedicated accelerator (such as a graphics processing unit (GPU)). In addition, a computing node may store and maintain data. In addition, a language model (e.g., a large language model) involved in the solution for a data warehouse according to an embodiment of the present disclosure may be deployed on one computing node or across multiple computing nodes. Of course, a model outside the system may also be called.
[0040] like Figure 1 As shown in the example, the local computing node 121 and the computing node 122 at the cloud can be interconnected via a network to communicate with each other. Figure 1 Taking the architecture in the example, computing nodes 121 and 122 can communicate via a network to achieve, for example, data synchronization or sharing between nodes. Multiple computing nodes in example environment 100 can process computing tasks in parallel. Furthermore, multiple computing nodes in example environment 100 can have redundancy and fault tolerance mechanisms to ensure reliable execution of computing tasks. When a computing node fails, the system can automatically migrate the job to another functioning node, ensuring task continuity and availability.
[0041] Examples of computing nodes may include supercomputers, personal computers, laptop computers, vehicle-mounted computing devices, mobile devices (such as smartphones, tablet computers, etc.), wearable electronic devices, multimedia devices, personal digital assistants (PDAs), or a combination of any one or more of the above devices. It should be understood that the computing nodes described herein are merely exemplary and non-limiting, and for example, other different types of computing nodes may also be used.
[0042] Combined with the above Figure 1 Describes an example environment in which the methods and / or processes according to embodiments of the present disclosure may be implemented. Figure 2 The method 200 for a data warehouse according to an embodiment of the present disclosure is described. Through this method 200, the key calculation logic on the indicator processing link can be fully reflected, and the indicator caliber information can be accurately inferred.
[0043] Figure 22 is a flow chart schematically illustrating a method 200 for a data warehouse according to an embodiment of the present disclosure. At 210, metadata information and business knowledge for a data table are obtained. This metadata information includes table-related data indicating the relationships between multiple data tables corresponding to the data warehouse. According to an embodiment of the present disclosure, metadata information and business knowledge for each data table in the multiple data tables corresponding to the data warehouse are obtained. This metadata information can describe the production process of the data warehouse at the data table granularity, covering the table structure of the corresponding data table, the calculation logic of fields within and between tables, and the relationships between tables. The table-related data indicating the relationships between tables in the metadata information is also called table lineage data. This data can be used to determine the relationships between a data table and upstream and downstream data tables in the indicator processing chain, such as the indicator relationships between data tables. Furthermore, this business knowledge can summarize and explain the business characteristics of the data table. Because business characteristics are often discrete and diverse between data warehouses and between data tables, considering business knowledge can promote a deeper understanding of the data tables.
[0044] At 220, based on the table associated data, the upstream indicator caliber of the upstream data table associated with the data table is obtained. According to an embodiment of the present disclosure, by indicating the association relationship between the data table and other corresponding data tables on the indicator processing link, the upstream data table of the data table can be determined and the upstream indicator caliber of the upstream data table can be obtained, wherein the upstream indicator caliber is a historical reasoning result that has been inferred by the language model and saved for subsequent reasoning. In this way, the upstream and downstream connection logic of the indicator caliber can be considered in the indicator caliber reasoning based on the language model, without being limited to the single layer being reasoned, thereby realizing multi-layer information transmission without missing key information.
[0045] At 230, the acquired metadata information, business knowledge, and upstream metric metric are fed into a language model to infer the metric metric of the data table, thereby describing the metrics associated with the data table in the data warehouse. According to an embodiment of the present disclosure, before feeding the acquired metadata information, business knowledge, and upstream metric metric into the language model, the metadata information, business knowledge, and upstream metric metric can be organized into model inputs based on natural language descriptions. For example, the acquired metadata information, business knowledge, and upstream metric metric can be assembled into prompt word engineering (PE) input information.
[0046] According to the method 200 for a data warehouse according to an embodiment of the present disclosure, a multi-layered automated metric caliber generation strategy for a data warehouse is provided. This strategy enables the accumulation of metric caliber information along the metric processing chain, thereby fully reflecting the key information along the chain. Furthermore, multimodal metric caliber reasoning, which combines business data with technical data, can facilitate language models' understanding of the business characteristics of the data warehouse, thereby enhancing the accuracy of metric caliber reasoning.
[0047] Figure 3A A diagram schematically illustrates an example 300A of an association relationship between data tables according to an embodiment of the present disclosure. As described above, a data warehouse may correspond to multiple data tables. A data table may be the basic structural unit for organizing data in a data warehouse. Data tables generally do not exist in isolation, but may be associated through various association relationships, such as one-to-one, one-to-many, etc.
[0048] By way of example and not limitation, Figure 3A As shown in FIG, illustrative example 300A includes two data tables (i.e., a first data table 310 and a second data table 320). First data table 310 may include indicators A1, A2, ..., An (i.e., the columns of first data table 310), and may also include records X1, X2, ..., Xn (i.e., the rows of first data table 310). Similarly, second data table 320 may include indicators B1, B2, ..., Bn (i.e., the columns of second data table 320), and may also include records Y1, Y2, ..., Yn (i.e., the rows of second data table 320).
[0049] like Figure 3A As schematically shown by the arrows in FIG, the index B1 of the second data table 320 is associated with the index An of the first data table 310. For example, the index B1 is calculated based on the index An. Furthermore, the index Bn of the second data table 320 is associated with the indexes A1 and A2 of the first data table 310. For example, the index Bn is calculated based on the indexes A1 and A2. An association relationship (i.e., an index association relationship) exists between the indexes of the first data table 310 and the indexes of the second data table 320, thereby associating the first data table 310 with the second data table 320.
[0050] Figure 3BA diagram schematically illustrates an example 300B of an indicator processing chain according to an embodiment of the present disclosure. The indicator processing chain can indicate a production process in a data warehouse, wherein the various data tables thereon are interconnected. As shown in FIG3 , depending on the association relationship between multiple data tables in the indicator processing chain, the upstream data table of data table 2 can be data table 1, and the upstream data tables of data table 3 can be data table 1 and data table 2. In other words, the upstream data table can be one or more data tables located before the data table in the indicator processing chain indicating the association relationship between multiple data tables.
[0051] According to an embodiment of the present disclosure, in the case where the upstream data table is a plurality of data tables, the previous data table on the indicator processing link or the previous data table that meets the correlation requirement can be selected as the target upstream data table for indicator caliber reasoning based on the language model, and its indicator caliber is obtained as the target upstream indicator caliber for reference. In some embodiment sets, the previous data table with the highest indicator correlation with the data table can be selected as the target data table. According to an embodiment of the present disclosure, a predetermined number of previous data tables of the data table or a predetermined number of previous data tables that meet the correlation threshold can also be selected as a target upstream data table set for indicator caliber reasoning based on the language model, and their indicator caliber is obtained as a target upstream indicator caliber set for reference, and the multiple upstream indicator calibers in the target set can be fed into the language model together to facilitate reasoning.
[0052] Figure 4 The figure schematically illustrates a process 400 for inferring indicator caliber based on a language model according to an embodiment of the present disclosure. The process 400 for inferring indicator caliber based on a language model is an iterative inferring process for each data table in a plurality of data tables associated with a data warehouse, layer by layer, until each data table in the indicator processing chain is inferred.
[0053] like Figure 4 As shown in , metadata information of a data table can be obtained, for example, including table structure metadata 410, table processing task metadata 420, and table association data 430. According to an embodiment of the present disclosure, table structure metadata 410 may include at least one of the following: the relationship between the table and the field, table annotations, field annotations, or field data types, and table processing task metadata 420 may include at least one of the following: table production code or code annotations. Table association data (i.e., table lineage data) 430 can be used to determine the association relationship between the data table and upstream and downstream data tables in the indicator processing link (i.e., the upstream and downstream dependency relationship of the table).
[0054] In addition, the business knowledge 440 of the data table can be obtained. According to an embodiment of the present disclosure, the business knowledge 340 may include at least one of the following: business tracking information, business professional terms, business background information, or business keywords. The business tracking information may be based on a data subset corresponding to the business tracking in the log data. Business professional terms may include professional terms in the business field and their corresponding explanations. Business background information may include a description of the business field. In addition, business keywords may indicate key elements in the business field, such as the business meaning of each business type.
[0055] According to an embodiment of the present disclosure, at 470, the upstream indicator caliber of the upstream data table is obtained from the storage device based on the table association data 330. The upstream data table may be one or more data tables located before the data table in the indicator processing chain, and the upstream indicator caliber may be a historical inference result inferred by the language model and saved for subsequent inference. In some embodiments, the upstream indicator caliber information may be cached or persistently stored in the storage device.
[0056] After obtaining the metadata information and business knowledge, as well as the upstream indicator caliber of the upstream data table, at 450, the metadata information, business knowledge, and upstream indicator caliber can be organized into a model input based on a natural language description, and then the model input can be fed into a language model for reasoning to obtain the indicator caliber 460 of the data table. According to an embodiment of the present disclosure, when organizing the metadata information and business knowledge, the metadata information and business knowledge can be converted into key-value pairs based on keyword extraction, keyword relationship pairing, etc. Proceeding to 470, the indicator caliber 460 of the data table of this layer can be stored in a storage device for downstream reasoning.
[0057] After the indicator caliber reasoning for the data table at this layer is completed, the indicator caliber reasoning for the data table at the next layer can be continued, thereby realizing multi-layer information transmission. That is, the indicator caliber reasoning based on the language model according to the embodiment of the present disclosure is performed on the downstream data tables of the data table. For ease of reference, the metadata information of the data table is referred to as the first metadata information, and the business knowledge of the data table is referred to as the first business knowledge.
[0058] According to an embodiment of the present disclosure, after the indicator caliber reasoning for a data table (the data table for which the latest inference result is generated) is completed, the second metadata information and second business knowledge of the downstream data table associated with the data table can be obtained, and the second metadata information includes table association data indicating the upstream and downstream dependencies of the table. The downstream data table is a data table adjacent to the data table on the indicator processing link. The indicator caliber of the data table can be obtained (for example, from a storage device) based on the table association data. The downstream indicator caliber of the downstream data table is inferred by feeding the second metadata information and the second business knowledge, as well as the indicator caliber of the data table into the language model to describe the indicators of the data warehouse associated with the downstream data table. The downstream indicator caliber of the downstream data table can be stored in a storage device for subsequent reasoning. This reasoning process is similar to Figure 4 The processes described in the same or similar manner may also have the same or similar solution details.
[0059] Figure 5 5 is a diagram schematically illustrating an example 500 of a link execution order for indicator-caliber reasoning according to an embodiment of the present disclosure. Figure 5 As shown in the figure, the execution order of the index processing link of the data warehouse is shown in the form of numbers, where multiple nodes ( Figure 5 Each node in the indicator caliber reasoning 510-570) shown in the figure indicates the indicator caliber reasoning of the corresponding layer in the multi-layer indicator caliber reasoning. The indicator caliber deduction starts from the bloodline root node table and deduces downward to the leaf node table, thereby ensuring the order of information transmission.
[0060] According to an embodiment of the present disclosure, after the indicator caliber reasoning for all data tables in the upstream data table is completed, the indicator caliber reasoning for the data table can be enabled. In an exemplary and non-restrictive manner, the indicator caliber reasoning 540 can be started only after the indicator caliber reasoning 510 and the indicator caliber reasoning 520 are completed. The indicator caliber information reasoned upstream can be passed to the indicator caliber reasoning downstream. In the example, the reasoning result of the indicator caliber reasoning 510 can be passed to the indicator caliber reasoning 530 to enable the indicator caliber reasoning 530. In another example, the reasoning results of the indicator caliber reasoning 510 and the indicator caliber reasoning 520 can be passed to the indicator caliber reasoning 540 to enable the indicator caliber reasoning 540.
[0061] According to the embodiments of the present disclosure, each inference is combined with the indicator caliber information inferred for the parent node table of the table, thereby accumulating indicator caliber information along the indicator processing chain of the data warehouse. Ultimately, after the indicator caliber is inferred for all tables, the upstream indicator caliber results are passed on to the downstream indicator caliber inference. In this way, the indicator caliber of the leaf node table can reflect all the core information upstream in the indicator processing chain.
[0062] Figure 6 FIG. 6 is a diagram schematically illustrating an apparatus 600 for a data warehouse according to an embodiment of the present disclosure. The apparatus 600 may include multiple units or modules for performing the steps or actions in the method or process discussed above. Figure 6 As shown in , the device 600 includes a data and knowledge acquisition module 610, which is configured to acquire metadata information and business knowledge of a data table, wherein the metadata information includes table association data indicating the association relationship between multiple data tables corresponding to the data warehouse. The device 600 also includes an upstream caliber acquisition module 620, which is configured to acquire the upstream indicator caliber of the upstream data table associated with the data table based on the table association data. The device 600 also includes a caliber inference module 630, which is configured to infer the indicator caliber of the data table by feeding the metadata information, the business knowledge, and the upstream indicator caliber into a language model to describe the indicators of the data warehouse associated with the data table.
[0063] In some embodiments, the metadata information may include at least one of the following: table structure metadata, including at least one of the following: the relationship between tables and fields, table comments, field comments, or field data types; or table processing task metadata, including at least one of the following: table production code, or code comments.
[0064] In some embodiments, the business knowledge indication may include at least one of the following: business tracking information; business professional terms; business background information; or business keywords.
[0065] In some embodiments, the upstream data table may include at least one data table located before the data table on an indicator processing link indicating an association relationship between the plurality of data tables.
[0066] In some embodiments, the upstream caliber acquisition module 620 can be further configured to: determine a target data table among the at least one data table based on the table association data; and retrieve the target indicator caliber of the target data table from a storage device, wherein the corresponding indicator caliber inferred by the language model for the corresponding data table is stored in the storage device.
[0067] In some embodiments, the apparatus 600 may further include an inference enabling module configured to enable indicator caliber reasoning for the data table after indicator caliber reasoning for all data tables in the upstream data table is completed.
[0068] In some embodiments, wherein the metadata information may be first metadata information, and the business knowledge may be first business knowledge, after the indicator caliber reasoning for the data table is completed, the device 600 may further include a second data and knowledge acquisition module, configured to obtain second metadata information and second business knowledge of a downstream data table associated with the data table, wherein the second metadata information includes the table-associated data. The device 600 may further include a second upstream caliber acquisition module, configured to obtain the indicator caliber of the data table based on the table-associated data. The device 600 may further include a second caliber reasoning module, configured to infer the downstream indicator caliber of the downstream data table by feeding the second metadata information and the second business knowledge, as well as the indicator caliber, into the language model to describe the indicators of the data warehouse associated with the downstream data table.
[0069] In some embodiments, the downstream data table may be a data table that is adjacent to the data table on an indicator processing link indicating an association relationship between the plurality of data tables.
[0070] In some embodiments, the device 600 may also include an input organization module configured to organize the metadata information, the business knowledge, and the upstream indicator caliber into a model input based on a natural language description before feeding the metadata information, the business knowledge, and the upstream indicator caliber into the language model.
[0071] In some embodiments, the input organization module may be further configured to convert the metadata information and the business knowledge into a key-value pair format based on keyword extraction and keyword relationship pairing.
[0072] Figure 7 FIG2 shows a block diagram of an electronic device 700 according to some embodiments of the present disclosure. The device 700 may be a device or apparatus described in the embodiments of the present disclosure. Figure 7As shown, the device 700 includes a central processing unit (CPU) and / or a graphics processing unit (GPU) 701, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 702 or computer program instructions loaded from a storage unit 708 into a random access memory (RAM) 703. Various programs and data required for the operation of the device 700 can also be stored in the RAM 703. The CPU / GPU 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704. Although not shown in FIG. Figure 7 As shown in FIG, device 700 may further include a co-processor.
[0073] Various components in device 700 are connected to I / O interface 705, including an input unit 706, such as a keyboard, mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, optical disk, etc.; and a communication unit 709, such as a network card, modem, wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0074] The various methods or processes described above may be performed by the CPU / GPU 701. For example, in some embodiments, the methods may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the CPU / GPU 701, one or more steps or actions in the methods or processes described above may be performed.
[0075] In some embodiments, the methods and processes described above may be implemented as a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present disclosure.
[0076] Computer-readable storage medium can be a tangible device that can keep and store the instructions used by the instruction execution device.Computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device or any suitable combination thereof.More specific examples (non-exhaustive list) of computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, for example, a punch card or a convex structure in a groove having instructions stored thereon, and any suitable combination thereof.Computer-readable storage medium used herein is not interpreted as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagated by waveguides or other transmission media (for example, light pulses by fiber optic cables), or electrical signals transmitted by wires.
[0077] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.
[0078] The computer program instructions for performing the disclosed operation can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data or source code or the object code written in any combination of one or more programming languages, programming languages include object-oriented programming languages, and conventional procedural programming languages.Computer-readable program instructions can be performed completely on a user's computer, partially on a user's computer, performed as an independent software package, partly on a user's computer and partly on a remote computer, or performed completely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer by any type of network-including local area network (LAN) or wide area network (WAN), or can be connected to an external computer (such as utilizing an internet service provider to connect by the internet). In certain embodiments, by utilizing the state information of computer-readable program instructions to carry out personalized customization electronic circuits, such as programmable logic circuits, field programmable gate arrays (FPGAs) or programmable logic arrays (PLA), this electronic circuit can perform computer-readable program instructions, thereby realizing various aspects of the present disclosure.
[0079] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0080] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0081] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart, can be implemented by a special hardware-based system that performs the prescribed function or action, or can be implemented by a combination of special hardware and computer instructions.
[0082] While various embodiments of the present disclosure have been described above, the above descriptions are intended to be illustrative and non-exhaustive, and are not intended to limit the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the various embodiments, their practical applications, or technical improvements to existing technologies, or to enable others skilled in the art to understand the various embodiments disclosed herein.
Claims
1. A method for a data warehouse, comprising: Acquire metadata information and business knowledge of a data table, wherein the metadata information includes table association data indicating association relationships between a plurality of data tables corresponding to a data warehouse; Based on the table associated data, obtaining the upstream indicator caliber of the upstream data table associated with the data table; as well as The indicator caliber of the data table is inferred by feeding the metadata information, the business knowledge, and the upstream indicator caliber into a language model to describe the indicators of the data warehouse associated with the data table.
2. The method according to claim 1, wherein the metadata information further includes at least one of the following: Table structure metadata, including at least one of the following: the relationship between the table and the field, table comments, field comments, or field data type; or Table processing task metadata includes at least one of the following: table production code or code comments.
3. The method according to claim 1, wherein the business knowledge indication comprises at least one of the following: Business tracking information; Business terminology; Business background information; or Business keywords.
4. The method according to claim 1, wherein: The upstream data table includes at least one data table located before the data table on an index processing link indicating an association relationship between the plurality of data tables.
5. The method according to claim 4, wherein obtaining the upstream indicator caliber comprises: determining a target data table in the at least one data table based on the table association data; as well as The target index caliber of the target data table is retrieved from a storage device, wherein the corresponding index caliber inferred by the language model for the corresponding data table is stored in the storage device.
6. The method according to claim 4, further comprising: After the indicator caliber reasoning for all data tables in the upstream data table is completed, the indicator caliber reasoning for the data table is enabled.
7. The method according to claim 1, wherein the metadata information is first metadata information, and the business knowledge is first business knowledge, and after the indicator caliber reasoning for the data table is completed, the method further comprises: Acquire second metadata information and second business knowledge of a downstream data table associated with the data table, wherein the second metadata information includes the table-associated data; Based on the table associated data, obtaining the indicator caliber of the data table; as well as By feeding the second metadata information, the second business knowledge, and the indicator caliber into the language model, the downstream indicator caliber of the downstream data table is inferred to describe the indicators of the data warehouse associated with the downstream data table.
8. The method according to claim 7, wherein: The downstream data table is a data table that is adjacent to the data table on an index processing link indicating an association relationship between the plurality of data tables.
9. The method according to claim 1, further comprising: Before feeding the metadata information, the business knowledge, and the upstream indicator caliber into the language model, the metadata information, the business knowledge, and the upstream indicator caliber are organized into a model input based on a natural language description.
10. The method according to claim 9, wherein the organization of the metadata information and the business knowledge comprises: Based on keyword extraction and keyword relationship pairing, the metadata information and the business knowledge are converted into a key-value pair format.
11. An apparatus for a data warehouse, comprising: a data and knowledge acquisition module configured to acquire metadata information and business knowledge of a data table, wherein the metadata information includes table association data indicating association relationships between a plurality of data tables corresponding to the data warehouse; An upstream caliber acquisition module is configured to acquire the upstream indicator caliber of the upstream data table associated with the data table based on the associated data of the table; The caliber reasoning module is configured to infer the indicator caliber of the data table by feeding the metadata information, the business knowledge, and the upstream indicator caliber into a language model to describe the indicators of the data warehouse associated with the data table.
12. An electronic device comprising: processor; as well as A memory coupled to the processor, wherein instructions are stored in the memory, and when the instructions are executed by the processor, the processor is caused to perform the method according to any one of claims 1 to 10.
13. A machine-readable storage medium having machine-executable instructions stored thereon, wherein when the machine-executable instructions are executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 10.
14. A computer program product comprising a computer program which, when executed by a processor of a computer, causes the processor to perform the method according to any one of claims 1 to 10.