Data processing method, data processing device, electronic device and storage medium

By extracting data related to task topics, user topics and facts from multiple business data, and building a fact data set based on business hierarchical relationships, the problem of low data processing efficiency in the existing technology is solved, and high-quality labeled data is quickly constructed, which improves the training quality of artificial intelligence models and the security and stability of the project.

CN113987086BActive Publication Date: 2025-06-06BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111251602.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-26
Publication Date
2025-06-06
Estimated Expiration
2041-10-26

AI Technical Summary

Technical Problem

The existing technology is less efficient when processing business data, making it difficult to quickly build high-quality labeled data, which affects the training of artificial intelligence models and the security and stability of projects.

Method used

By extracting data related to task topics, user topics and facts from multiple business data, a fact data set is constructed based on business hierarchical relationships, and a statistical data set corresponding to statistical indicators is generated to improve data processing efficiency.

Benefits of technology

It realizes the rapid construction of business data corresponding to statistical indicators, improves data processing efficiency, and enhances the training quality of artificial intelligence models and the security and stability of the project.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113987086B_ABST
    Figure CN113987086B_ABST
Patent Text Reader

Abstract

The present disclosure provides a data processing method, a data processing device, an electronic device and a storage medium, which relate to the fields of computer technology, and in particular to the fields of data warehouse, big data and cloud computing technology. The specific implementation scheme is: obtaining task dimension data corresponding to the task subject, user dimension data corresponding to the user subject and fact data corresponding to the fact from multiple business data, the task dimension data represents the business data corresponding to the dimension related to the task subject, and the user dimension data represents the business data corresponding to the dimension related to the user subject; based on the business hierarchical relationship, at least one fact data set is obtained according to the task dimension data, the user dimension data and the fact data, and the business hierarchical relationship represents the hierarchical relationship of the task subject and the hierarchical relationship of the user subject; according to at least one fact data set, at least one statistical data set corresponding to each statistical indicator of at least one statistical indicator is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to data warehouse, big data and cloud computing technology. Specifically, it relates to a data processing method, a data processing device, an electronic device and a storage medium. Background Art

[0002] Using high-quality business data to perform data processing operations can improve the security and stability of the project and promote the implementation of the project.

[0003] For example, high-quality labeled data can be obtained based on high-quality business data. High-quality labeled data can be used as training samples for training models in the field of artificial intelligence, thereby improving the security and stability of related projects based on the model. Summary of the invention

[0004] The present disclosure provides a data processing method, a data processing device, an electronic device, and a storage medium.

[0005] According to one aspect of the present disclosure, a data processing method is provided, comprising: acquiring task dimension data corresponding to a task subject, user dimension data corresponding to a user subject, and fact data corresponding to facts from multiple business data, wherein the task dimension data represents the business data corresponding to the dimension related to the task subject, and the user dimension data represents the business data corresponding to the dimension related to the user subject; based on a business hierarchical relationship, obtaining at least one fact data set according to the task dimension data, the user dimension data, and the fact data, wherein the business hierarchical relationship represents the hierarchical relationship of the task subject and the hierarchical relationship of the user subject; and obtaining at least one statistical data set corresponding to each statistical indicator of at least one statistical indicator according to the at least one fact data set.

[0006] According to another aspect of the present disclosure, a data processing device is provided, including: an acquisition module, used to acquire task dimension data corresponding to a task subject, user dimension data corresponding to a user subject, and fact data corresponding to facts from multiple business data, wherein the task dimension data represents the business data corresponding to the dimension related to the task subject, and the user dimension data represents the business data corresponding to the dimension related to the user subject; a first acquisition module, used to obtain at least one fact data set based on a business hierarchical relationship according to the task dimension data, the user dimension data, and the fact data, wherein the business hierarchical relationship represents the hierarchical relationship of the task subject and the hierarchical relationship of the user subject; and a second acquisition module, used to obtain at least one statistical data set corresponding to each statistical indicator of at least one statistical indicator according to the at least one fact data set.

[0007] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described above.

[0008] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method described above.

[0009] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, and the computer program implements the method described above when executed by a processor.

[0010] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0012] Figure 1 An exemplary system architecture to which the data processing method and device according to an embodiment of the present disclosure can be applied is schematically shown;

[0013] Figure 2 A flowchart schematically shows a data processing method according to an embodiment of the present disclosure;

[0014] Figure 3An example schematic diagram schematically shows a data query process according to an embodiment of the present disclosure;

[0015] Figure 4 A flowchart of storing original business data into a data warehouse according to an embodiment of the present disclosure is schematically shown;

[0016] Figure 5 An example schematic diagram of a process of acquiring multiple original business data according to an embodiment of the present disclosure is schematically shown;

[0017] Figure 6 An example schematic diagram of a process of generating a statistical data set according to an embodiment of the present disclosure is schematically shown;

[0018] Figure 7 A block diagram schematically shows a data processing device according to an embodiment of the present disclosure; and

[0019] Figure 8 A block diagram of an electronic device suitable for implementing a data processing method according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0020] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0021] In order to obtain high-quality annotated data, it is necessary to provide more efficient business data. In the process of obtaining more efficient business data, the calculation of business data is involved. It can be implemented by using scripts to execute structured query statements and then store the business data in a database. For example, after executing structured query statements using PHP (Hypertext Preprocessor), the business data is stored in a database. The above processing method is not efficient.

[0022] To this end, the embodiment of the present disclosure proposes a data processing solution based on a data warehouse, that is, task dimension data corresponding to the task subject, user dimension data corresponding to the user subject, and fact data corresponding to the fact are obtained from multiple business data. The task dimension data represents the business data corresponding to the dimension related to the task subject, and the user dimension data represents the business data corresponding to the dimension related to the user subject. Based on the business hierarchical relationship, at least one fact data set is obtained according to the task dimension data, the user dimension data and the fact data. The business hierarchical relationship represents the hierarchical relationship of the task subject and the hierarchical relationship of the user subject. According to at least one fact data set, at least one statistical data set corresponding to each statistical indicator of at least one statistical indicator is obtained, which realizes the relatively rapid construction of business data corresponding to the statistical indicator and improves the data processing efficiency. In order to facilitate subsequent understanding, the concepts involved in the embodiment of the present disclosure are first explained.

[0023] A data warehouse is a structured data environment for decision support systems and online analytical application data sources. A data warehouse studies and solves the problem of obtaining information from a database. The characteristics of a data warehouse are subject-oriented, integrated, stable, and time-varying.

[0024] A topic can refer to an abstraction that integrates and classifies business data at a higher level and uses it for analysis. Each topic corresponds to a macro analysis field.

[0025] Dimensions can refer to the selection and measurement of facts, and can be seen as a perspective for analyzing business data.

[0026] Facts can refer to data of interest, that is, metrics extracted from business data. Metrics can include cumulative metrics and non-cumulative metrics.

[0027] Figure 1 An exemplary system architecture to which the data processing method and apparatus according to an embodiment of the present disclosure can be applied is schematically shown.

[0028] It should be noted that Figure 1 What is shown is merely an example of a system architecture to which the embodiments of the present disclosure can be applied, in order to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.

[0029] like Figure 1As shown, the system architecture 100 according to this embodiment may include a business system 110, a data warehouse system 120, and an application layer system 130. The data warehouse system 120 may be connected to the business system 110 and the application layer system 130 through a network. The network may be a medium that provides a communication link. The network may include various connection types, such as at least one of a wired communication link and a wireless communication link.

[0030] The business system 110 may include a database 111 , a log system 112 , an event bus management system 113 , and cloud storage 114 .

[0031] The data warehouse system 120 may include metadata management 121, data quality monitoring 122, raw data (Operational Data Store, ODS) layer 123, data warehouse (Data Warehouse, DW) layer 124 and data application (Application Data Service, ADS) layer 125. The data warehouse layer 124 may include a data detail (DataWarehouse Detail, DWD) layer 1240, a common dimension summary (ie, DIM) layer 1241 and a data summary (DataWarehouse Summary, DWS) layer 1242.

[0032] The data warehouse system 120 can be a server that provides various services. For example, the server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability in traditional physical hosts and VPS services (Virtual Private Server, VPS). The server can also be a server of a distributed system, or a server combined with a blockchain.

[0033] The application layer 130 may include a report platform 131 , an experiment platform 132 , a statistical application program interface 133 , and a business module 134 .

[0034] The data warehouse system 120 can obtain original business data from the business system 110. The data warehouse system 120 can obtain task dimension data corresponding to the task subject, user dimension data corresponding to the user subject, and fact data corresponding to the fact from multiple business data. The task dimension data represents the business data corresponding to the dimension related to the task subject, and the user dimension data represents the business data corresponding to the dimension related to the user subject. Based on the business hierarchical relationship, at least one fact data set is obtained according to the task dimension data, the user dimension data, and the fact data. The business hierarchical relationship represents the hierarchical relationship of the task subject and the hierarchical relationship of the user subject. According to the at least one fact data set, at least one statistical data set corresponding to each statistical indicator of the at least one statistical indicator is obtained.

[0035] The application layer system 130 may generate a query request including a query condition, and send the generated request to the data warehouse system 120. The data warehouse system 120 may respond to the query request and determine, according to the query condition, a data query result matching the query condition from at least one statistical data set corresponding to each statistical indicator of the at least one statistical indicator.

[0036] It should be noted that the data processing method provided in the embodiment of the present disclosure can generally be executed by the data warehouse system 120. Accordingly, the data processing device provided in the embodiment of the present disclosure can generally be set in the data warehouse system 120. The data processing method provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the data warehouse system 120 and can communicate with the data warehouse system 120. Accordingly, the data processing device provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the data warehouse system 120 and can communicate with the data warehouse system 120.

[0037] It should be understood that Figure 1 The system architecture in the figure is only for illustration. Other forms of system architecture may be used according to the implementation requirements.

[0038] Figure 2 The flowchart of the data processing method according to the embodiment of the present disclosure is schematically shown.

[0039] like Figure 2 As shown, the method 200 includes operations S210 to S230.

[0040] In operation S210, task dimension data corresponding to the task subject, user dimension data corresponding to the user subject, and fact data corresponding to the fact are obtained from multiple business data. The task dimension data represents the business data corresponding to the dimension related to the task subject, and the user dimension data represents the business data corresponding to the dimension related to the user subject.

[0041] In operation S220, at least one fact data set is obtained based on the business hierarchical relationship, task dimension data, user dimension data and fact data. The business hierarchical relationship represents the hierarchical relationship between the task subject and the hierarchical relationship between the user subject.

[0042] In operation S230, at least one statistical data set corresponding to each statistical indicator of the at least one statistical indicator is obtained according to the at least one fact data set.

[0043] According to an embodiment of the present disclosure, business data may refer to data related to a business. For example, business data may be operation data generated by a user performing a related operation.

[0044] According to an embodiment of the present disclosure, a topic may include a task topic and a user topic. For each business scenario, it can be analyzed from two perspectives: a task topic and a user topic. That is, for each business scenario, the perspective of analyzing the business scenario can be divided into a task topic and a user topic. A task topic may refer to a topic that analyzes business data from a task perspective. A user topic may refer to a topic that analyzes business data from a user perspective. Both a task topic and a user topic may be topics with a hierarchical relationship. That is, a task topic may include multiple dimensions related to the task topic, and the multiple dimensions related to the task topic have a hierarchical relationship. A user topic may include multiple dimensions related to the user topic, and the multiple dimensions related to the user topic have a hierarchical relationship.

[0045] According to an embodiment of the present disclosure, a business hierarchical relationship can be used to characterize the hierarchical relationship of a task subject and the hierarchical relationship of a user subject. The business hierarchical relationship can be used as a basis for constructing a fact data set. Statistical indicators can be used as a basis for business data aggregation. Statistical indicators can include indicators related to projects, indicators related to resources, and indicators related to interactions.

[0046] According to an embodiment of the present disclosure, after obtaining a plurality of business data, task dimension data corresponding to the task subject, user dimension data corresponding to the user subject, and fact data corresponding to the fact can be obtained from the plurality of business data. Then, the dimension related to the task subject and the dimension related to the user subject can be associated based on the hierarchical relationship of the task subject and the hierarchical relationship of the user subject. According to the associated dimension related to the task subject and the dimension related to the user subject, business data corresponding to the associated dimension related to the task subject and the dimension related to the user subject are obtained from the task dimension data, the user dimension data, and the fact data, and the business data corresponding to the associated dimension related to the task subject and the dimension related to the user subject are aggregated to obtain at least one fact data set. The fact data set can be a fact data table. Finally, at least one statistical indicator can be determined. For each statistical indicator in the at least one statistical indicator, business data corresponding to the statistical indicator can be obtained from the at least one fact data set. According to the business data corresponding to the statistical indicator, at least one statistical data set corresponding to the statistical indicator is obtained. The statistical data set can be a statistical data table.

[0047] According to the embodiments of the present disclosure, at least one fact data set is obtained based on task dimension data, user dimension data and fact data based on business hierarchical relationships, and at least one statistical data set corresponding to the statistical indicators is obtained based on the at least one fact data set, thereby achieving relatively rapid construction of business data corresponding to the statistical indicators and improving data processing efficiency.

[0048] Reference below Figure 3 to Figure 6 , the data processing method according to the embodiment of the present disclosure is further explained in combination with specific embodiments.

[0049] According to an embodiment of the present disclosure, operation S210 may include the following operations.

[0050] According to the feature association relationship, the features corresponding to the task subject, the features corresponding to the user subject, and the features corresponding to the fact are determined. According to the features corresponding to the task subject, the task dimension data corresponding to the task subject is obtained from multiple business data. According to the features corresponding to the user subject, the user dimension data corresponding to the user subject is obtained from multiple business data. According to the features corresponding to the fact, the fact data corresponding to the fact is obtained from multiple business data.

[0051] According to an embodiment of the present disclosure, a feature association relationship may refer to an association relationship between an analysis angle and features involved in the analysis angle. The analysis angle may include a subject and a fact. The subject may include a task subject and a user subject.

[0052] According to an embodiment of the present disclosure, the feature corresponding to the analysis angle can be determined according to the association relationship between the analysis angle and the feature involved in the analysis angle. After the feature corresponding to the analysis angle is determined, the business data corresponding to the analysis angle is obtained from multiple business data according to the feature corresponding to the analysis angle.

[0053] For example, if the analysis angle is a task theme, the features involved in the task theme may include features corresponding to dimensions related to the task theme. For example, features corresponding to dimensions related to the task theme may include features corresponding to page dimensions, features corresponding to task dimensions, features corresponding to batch dimensions, and features corresponding to project dimensions.

[0054] For example, if the analysis angle is user theme, the features involved in the user theme may include features corresponding to dimensions related to the user theme. For example, the features corresponding to dimensions related to the user theme may include features corresponding to user dimensions, features corresponding to guild dimensions, and features corresponding to agent dimensions.

[0055] For example, if the analysis angle is facts, the characteristics involved in the facts may include the accuracy of answering questions and the time it takes to answer questions.

[0056] According to an embodiment of the present disclosure, a task topic includes multiple dimensions related to the task topic and having a hierarchical relationship, and a user includes multiple dimensions related to the user topic and having a hierarchical relationship.

[0057] According to an embodiment of the present disclosure, operation S220 may include the following operations.

[0058] A plurality of hierarchically related dimensions related to task topics and a plurality of hierarchically related dimensions related to user topics are associated to obtain a plurality of hierarchically related dimensions. Based on the plurality of hierarchically related dimensions, task dimension data, user dimension data and fact data are aggregated to obtain at least one fact data set.

[0059] According to an embodiment of the present disclosure, a task theme may include multiple dimensions related to the task theme. Multiple dimensions related to the task theme may have a hierarchical relationship. A user theme may include multiple dimensions related to the user theme. Multiple dimensions related to the user theme may have a hierarchical relationship. A dimension may include at least one sub-dimension.

[0060] For example, the dimensions related to the task theme may include page dimension, task dimension, batch dimension, and project dimension. The levels of page dimension, task dimension, batch dimension, and project dimension are increased in sequence. The dimensions related to the user theme may include user dimension, guild dimension, and agent dimension. The levels of user dimension, guild dimension, and agent dimension are increased in sequence.

[0061] According to an embodiment of the present disclosure, an associated dimension may represent a dimension that is associated with both a dimension associated with a task theme and a dimension associated with a user theme. Multiple associated dimensions may be obtained by associating multiple dimensions associated with a task theme having a hierarchical relationship and multiple dimensions associated with a user theme having a hierarchical relationship. Multiple associated dimensions may have a hierarchical relationship.

[0062] For example, a user topic may include M dimensions related to the user topic, namely, the first dimension related to the user topic, the second dimension related to the user topic, ..., the i-th dimension related to the user topic, ..., the (M-1)-th dimension related to the user topic, and the M-th dimension related to the user topic. i∈{1, 2, ..., M-1, M}. M may be an integer greater than or equal to 2.

[0063] For example, the task theme may include N dimensions related to the task theme, i.e., the first dimension related to the task theme, the second dimension related to the task theme, ..., the jth dimension related to the task theme, ..., the (N-1)th dimension related to the task theme, and the Nth dimension related to the task theme. N may be an integer greater than or equal to 2. j∈{1, 2, ..., N-1, N}. The associated dimension may include the i-th dimension related to the user theme - the j-th dimension related to the task theme.

[0064] According to an embodiment of the present disclosure, multiple hierarchical dimensions related to task topics and multiple hierarchical dimensions related to user topics can be associated to form multiple hierarchical associated dimensions. After determining multiple hierarchical associated dimensions, business data related to the associated dimensions can be obtained from task dimension data, user dimension data, and fact data, and the business data related to the associated dimensions can be aggregated to obtain at least one fact data set.

[0065] According to an embodiment of the present disclosure, operation S230 may include the following operations.

[0066] At least one statistical indicator is determined according to the business requirement rule. For each of the at least one statistical indicator, at least one fact data set corresponding to the statistical indicator is associated to obtain at least one statistical data set corresponding to the statistical indicator.

[0067] According to an embodiment of the present disclosure, a business requirement rule can be used as a rule for determining a statistical indicator. At least one statistical indicator corresponding to the business requirement rule can be determined according to the business requirement rule. After determining at least one statistical indicator, a fact data set related to the statistical indicator can be searched from at least one fact data set according to each statistical indicator, and the fact data set related to the statistical indicator can be associated, thereby obtaining one or more statistical data sets corresponding to the statistical indicator.

[0068] For example, a business requirement may be a requirement for collecting user role information. The statistical indicators determined according to the business requirement rules include type, time, and participating projects. After determining these three statistical indicators, a fact data set associated with each of the three statistical indicators may be searched from at least one fact data set to obtain at least one statistical data set corresponding to each statistical indicator.

[0069] According to an embodiment of the present disclosure, the above data processing method may further include the following operations.

[0070] In response to the query request, according to the query condition included in the query request, a data query result matching the query condition is determined from at least one statistical data set corresponding to each statistical indicator of the at least one statistical indicator.

[0071] According to an embodiment of the present disclosure, a query request may refer to a request for requesting processing of a query and a query condition included in the query request. The query request may be generated when it is detected that a query operation for a query box is triggered. The query box may be used to input a query condition or to select a query condition. The query operation may include a click operation, a slide operation, or a voice-triggered operation. For example, when it is detected that a query operation for a query box is triggered, the query condition entered by the operator in the query box is obtained. A query request is generated according to the query condition. The query condition may refer to a condition that needs to be satisfied for the business data to be obtained.

[0072] According to an embodiment of the present disclosure, a query request may be obtained, and the query request may include a query condition. In response to the query request, business data matching the query condition is determined from at least one statistical data set according to the query condition, and the business data matching the query condition is determined as a data query result.

[0073] According to an embodiment of the present disclosure, in response to a query request, according to the query conditions included in the query request, determining a data query result matching the query conditions from at least one statistical data set corresponding to each statistical indicator in at least one statistical indicator may include the following operations.

[0074] In response to the query request, the data warehouse interface is called. Using the data warehouse interface, according to the query condition included in the query request, a data query result matching the query condition is determined from at least one statistical data set corresponding to each statistical indicator in the at least one statistical indicator.

[0075] According to an embodiment of the present disclosure, a data warehouse interface may be configured. The data warehouse interface may refer to an interface for obtaining data query results that match the query conditions included in the query request. The data warehouse interface may have a RESTful style specification. REST (Representational State Transfer) represents the presentation layer state transfer of Internet resources. RESTful is an Internet software architecture.

[0076] According to an embodiment of the present disclosure, a data warehouse interface can be called in response to a query request, and then the data warehouse interface can be used to query business data matching the query conditions from at least one statistical data set, and data query results can be obtained based on the business data matching the query conditions.

[0077] According to an embodiment of the present disclosure, the data warehouse system may also be connected to a Dashboard (ie, data visualization) module and a Showx module to adapt to business needs.

[0078] According to an embodiment of the present disclosure, after obtaining the data query result, the data query result can also be displayed in the form of a business report. The display form can include at least one of a bar chart, a line chart, a progress chart, and a table.

[0079] According to an embodiment of the present disclosure, the above data processing method may further include the following operations.

[0080] The data query results are processed to obtain data processing results.

[0081] According to an embodiment of the present disclosure, after obtaining the data query result, the data query result can be used for subsequent data processing. For example, the data query result can be used to perform data annotation to obtain a data annotation result. For example, for data annotation in the field of artificial intelligence, the data label can be determined based on the model requirement information. The data query result is annotated according to the data label to obtain a data annotation result.

[0082] According to the embodiments of the present disclosure, other calculations are performed based on the constructed statistical data set corresponding to the statistical indicators, which can improve processing efficiency and meet the needs of generating massive data in a short time.

[0083] Figure 3 An example schematic diagram of a data query process according to an embodiment of the present disclosure is schematically shown.

[0084] like Figure 3 As shown, in 300, in response to a query request, a data warehouse interface 302 is called. Using the data warehouse interface 302, according to the query condition included in the query request, in the data warehouse 301, a data query result 303 matching the query condition is determined from at least one statistical data set corresponding to each statistical indicator of at least one statistical indicator.

[0085] According to an embodiment of the present disclosure, the above data processing method may further include the following operations.

[0086] Data cleaning is performed on multiple original business data to obtain multiple business data.

[0087] According to an embodiment of the present disclosure, data cleaning can be used to filter out business data that does not meet preset requirements and retain business data that meets preset requirements. The types of business data that need data cleaning may include at least one of incomplete business data, erroneous business data, and duplicate business data.

[0088] According to an embodiment of the present disclosure, data cleaning can be performed on a plurality of original business data from a plurality of data sources to obtain a plurality of business data. Each piece of original business data may include one or more pieces.

[0089] According to an embodiment of the present disclosure, the data quality of business data is improved by performing data cleaning on a plurality of original data.

[0090] Figure 4 The flowchart of storing original business data into a data warehouse according to an embodiment of the present disclosure is schematically shown.

[0091] like Figure 4 As shown, the method 400 includes operations S401-S402.

[0092] In operation S401 , based on a data structure specification of a data warehouse and a data structure of each of a plurality of data sources, a storage strategy for original business data corresponding to each data source is determined.

[0093] In operation S402, the original business data is stored in a data warehouse according to a storage policy of the original business data corresponding to each data source.

[0094] According to an embodiment of the present disclosure, a data warehouse has a corresponding data structure specification. The data structure specification of a data warehouse may refer to a data structure that is used to specify the data stored in the data warehouse. Each data source has a data structure corresponding to the data source. The data source may include at least one of a summary-level database, a business configuration data source, a fine-grained distributed database, a system service, a system log, and a system bus.

[0095] According to an embodiment of the present disclosure, the storage strategy may refer to a strategy for how to achieve that, when the original business data corresponding to each data source is stored in a data warehouse, the data structure of the original business data corresponding to each data source is a data structure that matches the data structure specification of the data warehouse, that is, the storage strategy may refer to storing the business data of each data source in accordance with the data structure specification of the data warehouse, so that the data structure of the original business data of the data source stored in the data warehouse matches the data structure corresponding to the data structure specification.

[0096] According to an embodiment of the present disclosure, for each of the multiple data sources, a storage strategy for the original business data corresponding to the data source can be determined based on the data structure specification of the data warehouse and the data structure of the data source, that is, a storage strategy for storing the original business data corresponding to the data source is determined. When the storage strategy for the original business data corresponding to the data source is determined, the original business data corresponding to the data source can be stored in the data warehouse according to the storage strategy for the original business data corresponding to the data source.

[0097] According to an embodiment of the present disclosure, according to the storage strategy of the original business data corresponding to the data source, storing the original business data corresponding to the data source in the data warehouse may include: obtaining the original business data from the data source. According to the storage strategy of the original business data corresponding to the data source, storing the original business data in the data warehouse so that the data structure of the original business data corresponding to the data source stored in the data warehouse matches the data structure corresponding to the data structure specification of the data warehouse. Storing the original business data corresponding to the data source in the data warehouse may include: storing the original business data corresponding to the data source in the original data layer of the data warehouse.

[0098] According to an embodiment of the present disclosure, a storage strategy for original business data corresponding to each data source is determined by utilizing a data structure specification based on a data warehouse and a data structure of each of a plurality of data sources, and the original business data corresponding to each data source is stored in the data warehouse, effectively ensuring that the data structure of the original business data from different data sources can match the data structure corresponding to the data structure specification of the data warehouse, thereby solving the problem of heterogeneity of the data structure and reducing development complexity.

[0099] According to an embodiment of the present disclosure, operation S402 may include the following operations.

[0100] According to the storage strategy of the original business data corresponding to each data source, determine the data import tool that matches the storage strategy. Use the data import tool that matches the storage strategy to store the original business data in the data warehouse.

[0101] According to an embodiment of the present disclosure, for each of the multiple data sources, after determining the storage strategy of the original business data corresponding to the data source, a data import tool that matches the business requirements and storage strategy can be determined based on the business requirements and storage strategy of the original business data corresponding to the data source, so as to obtain the original business data from the data source using the data import tool. According to the storage strategy of the original business data corresponding to the data source, the original business data is stored in the data warehouse, so that the data structure of the original business data corresponding to the data source stored in the data warehouse matches the data structure corresponding to the data structure specification of the data warehouse, and can meet the business requirements.

[0102] By adapting to different business needs, development costs can be reduced. Business needs may include at least one of real-time requirements, data volume requirements, and data storage feature requirements. For example, for real-time requirements, if the original business data has high real-time requirements, you can select a data import tool that supports a relatively fast synchronization function or a data import tool that supports a connection import function to perform storage operations. For example, a data import tool that supports a relatively fast synchronization function may include DTS (Data Transmission Service). A data import tool that supports a connection import function may include Flink. For quantity requirements and data storage feature requirements, if the original business data has a large amount of data and the data storage is relatively scattered, you can select a data import tool that supports a connection aggregation function to perform storage operations. For example, a data import tool that supports a connection aggregation function may include SPARK.

[0103] Figure 5 An example schematic diagram of a process of acquiring multiple original business data according to an embodiment of the present disclosure is schematically shown.

[0104] like Figure 5 As shown, in 500, the data sources include a summary-level database, a business configuration data source, a fine-grained distributed database, a system service, a system log, and a system bus.

[0105] For summary-level databases, the data import tool DTS can be used to synchronize the original business data in the summary-level database to the original data layer in the data warehouse according to the configuration information. For example, DTS can be used to call the Stream Load mechanism provided by Doris to synchronize the original business data in the summary-level database to the original data layer in the data warehouse. In addition, the data import tool Doris can be used to store the original business data in the summary-level database to the original data layer in the data warehouse based on Doris's Mapping mechanism. Doris can be an interactive SQL (Structured Query Language) data warehouse based on the MPP (Massively Parallel Processor) architecture.

[0106] For the business configuration data source, the data import tool Doris can be used to store the original business data in the business configuration data source into the original data layer in the data warehouse based on Doris's Mapping mechanism.

[0107] For fine-grained distributed databases, system services, and system logs, you can use the data import tool SPARK to execute connection scripts and store the original business data in the fine-grained distributed databases, system services, and system logs into the original data layer in the data warehouse based on the connection (i.e. Connector) mechanism.

[0108] For the system bus, at least one of the data import tool Flink and the data import tool Doris can be used to store the original business data in the system bus in the original data layer in the data warehouse. For example, the original business data in the system bus can be stored in the original data layer in the data warehouse based on the connection (i.e., Connector) mechanism of Flink or the Routine Load mechanism of Doris.

[0109] Figure 6 An example schematic diagram of a process of generating a statistical data set according to an embodiment of the present disclosure is schematically shown.

[0110] like Figure 6 As shown, in 600, a plurality of original business data are stored in the original data layer of the data warehouse. For example, the plurality of original business data are business data related to the question answering task. The plurality of original business data may include original task information, original user information, ..., original page answering record, original page correctness record, and original working time record, etc.

[0111] The data detail layer of the data warehouse cleans multiple original business data to obtain multiple business data. For example, multiple business data may include task information, user information, ..., page answer records, page correctness records, and working time records.

[0112] The public dimension data layer of the data warehouse determines the features corresponding to the task theme and the features corresponding to the user theme according to the feature association relationship. According to the features corresponding to the task theme, the task dimension data corresponding to the task theme is obtained from multiple business data. According to the features corresponding to the user theme, the user dimension data corresponding to the user theme is obtained from multiple business data. For example, the task theme may include 4 dimensions related to the task theme, namely, page dimension, task dimension, batch dimension and project dimension. The levels of page dimension, task dimension, batch dimension and project dimension are increased in sequence. The user theme may include 4 dimensions related to the user theme, namely, user dimension, guild dimension and agent dimension. The levels of user dimension, guild dimension and agent dimension are increased in sequence. The task dimension may include task status, task type, task stage and participation type, etc. The user dimension may include guild association information, agent, attribute information and time information, etc.

[0113] The data aggregation layer of the data warehouse can obtain fact data corresponding to the facts from multiple business data according to the characteristics corresponding to the facts. Multiple hierarchical dimensions related to task topics and multiple hierarchical dimensions related to user topics are associated to obtain multiple hierarchical associated dimensions. Based on multiple hierarchical associated dimensions, the task dimension data, user dimension data and fact data are aggregated to obtain at least one fact data set. For example, the associated dimensions may include user-task, user-batch, guild-task, guild-project, agent-project, task progress and project progress, etc. Figure 6 The “→” in this part represents the hierarchical relationship, the left side of “→” represents the associated dimension of the lower level, and the back side of “→” represents the associated dimension of the higher level. At least one fact data set may include user-task workload, task accuracy, user-task working hours, guild-task workload, user-batch workload, task progress, guild-project workload, agent-project workload and project progress, etc. Guild-task workload, user-batch workload and task progress can be obtained based on user-task workload. Guild-project workload and agent-project workload can be obtained based on guild-task workload. Project progress can be obtained based on task progress.

[0114] The data application layer of the data warehouse can determine at least one statistical indicator according to the business demand rules. For each statistical indicator in the at least one statistical indicator, at least one fact data set corresponding to the statistical indicator is associated to obtain at least one statistical data set corresponding to the statistical indicator. For example, at least one statistical data set can include object capability labels, object activity within 7 days, ..., object efficiency, object average capacity and object work efficiency ratio, etc.

[0115] The above are merely exemplary embodiments, but are not limited thereto, and may also include other data processing methods known in the art, as long as they can achieve relatively efficient data calculation.

[0116] It should be noted that in the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0117] Figure 7 The block diagram schematically shows a data processing device according to an embodiment of the present disclosure.

[0118] like Figure 7 As shown, the data processing device 700 may include an acquisition module 710 , a first acquisition module 720 , and a second acquisition module 730 .

[0119] The acquisition module 710 is used to acquire task dimension data corresponding to the task subject, user dimension data corresponding to the user subject, and fact data corresponding to the fact from multiple business data. The task dimension data represents the business data corresponding to the dimension related to the task subject, and the user dimension data represents the business data corresponding to the dimension related to the user subject.

[0120] The first obtaining module 720 is used to obtain at least one fact data set based on the business hierarchical relationship according to the task dimension data, the user dimension data and the fact data. The business hierarchical relationship represents the hierarchical relationship between the task subject and the hierarchical relationship between the user subject.

[0121] The second obtaining module 730 is used to obtain at least one statistical data set corresponding to each statistical indicator of the at least one statistical indicator according to at least one fact data set.

[0122] According to an embodiment of the present disclosure, the data processing device 700 may further include a third obtaining module.

[0123] The third acquisition module is used to clean the multiple original business data to obtain multiple business data.

[0124] According to an embodiment of the present disclosure, the above-mentioned data processing device 700 may further include a first determining module and a storage module.

[0125] The first determination module is used to determine a storage strategy for original business data corresponding to each data source based on a data structure specification of the data warehouse and a data structure of each data source among a plurality of data sources.

[0126] The storage module is used to store the original business data in the data warehouse according to the storage strategy of the original business data corresponding to each data source.

[0127] According to an embodiment of the present disclosure, the storage module may include a first determining unit and a storage unit.

[0128] The first determining unit is used to determine a data import tool that matches the storage strategy according to the storage strategy of the original business data corresponding to each data source.

[0129] The storage unit is used to store the original business data into the data warehouse using a data import tool that matches the storage strategy.

[0130] According to an embodiment of the present disclosure, the acquisition module may include a second determining unit, a first acquiring unit, a second acquiring unit, and a third acquiring unit.

[0131] The second determining unit is used to determine the features corresponding to the task theme, the features corresponding to the user theme, and the features corresponding to the facts according to the feature association relationship.

[0132] The first acquisition unit is used to acquire task dimension data corresponding to the task subject from multiple business data according to the characteristics corresponding to the task subject.

[0133] The second acquisition unit is used to acquire user dimension data corresponding to the user topic from multiple business data according to the characteristics corresponding to the user topic.

[0134] The third acquisition unit is used to acquire fact data corresponding to the fact from multiple business data according to the characteristics corresponding to the fact.

[0135] According to an embodiment of the present disclosure, the task subject includes a plurality of hierarchical dimensions related to the task subject, and the user includes a plurality of hierarchical dimensions related to the user subject;

[0136] The first obtaining module 720 may include a first obtaining unit and a second obtaining unit.

[0137] The first obtaining unit is used to associate a plurality of hierarchical dimensions related to task themes with a plurality of hierarchical dimensions related to user themes to obtain a plurality of hierarchical associated dimensions.

[0138] The second obtaining unit is used to aggregate the task dimension data, the user dimension data and the fact data based on a plurality of associated dimensions with a hierarchical relationship to obtain at least one fact data set.

[0139] According to an embodiment of the present disclosure, the second obtaining module may include a third determining unit and a third obtaining unit.

[0140] The third determining unit is used to determine at least one statistical indicator according to the business requirement rule.

[0141] The third obtaining unit is used to associate at least one fact data set corresponding to each statistical indicator of the at least one statistical indicator to obtain at least one statistical data set corresponding to the statistical indicator.

[0142] According to an embodiment of the present disclosure, the data processing device 700 may further include a second determining module.

[0143] The second determination module is used to respond to the query request and determine the data query result matching the query condition from at least one statistical data set corresponding to each statistical indicator of at least one statistical indicator according to the query condition included in the query request.

[0144] According to an embodiment of the present disclosure, the second determining module may include a calling unit and a fourth determining unit.

[0145] The calling unit is used to call the data warehouse interface in response to a query request.

[0146] The fourth determining unit is used to determine, by using the data warehouse interface, a data query result matching the query condition from at least one statistical data set corresponding to each statistical indicator of at least one statistical indicator according to the query condition included in the query request.

[0147] According to an embodiment of the present disclosure, the data processing device 700 may further include a fourth obtaining module.

[0148] The fourth acquisition module is used to process the data query result to obtain the data processing result.

[0149] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0150] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.

[0151] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the method as described above.

[0152] According to an embodiment of the present disclosure, a computer program product includes a computer program, and when the computer program is executed by a processor, the computer program implements the method as described above.

[0153] Figure 8 A block diagram of an electronic device suitable for implementing a data processing method according to an embodiment of the present disclosure is schematically shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0154] like Figure 8 As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0155] Multiple components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0156] The computing unit 801 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as data processing methods. For example, in some embodiments, the data processing method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the data processing method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform the data processing method in any other appropriate manner (e.g., by means of firmware).

[0157] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0158] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0159] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0160] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0161] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0162] A computer system may include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises through computer programs running on respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server combined with a blockchain.

[0163] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0164] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A data processing method, include: Acquire task dimension data corresponding to the task subject, user dimension data corresponding to the user subject, and fact data corresponding to the fact from the plurality of business data, wherein the task dimension data represents the business data corresponding to the dimension related to the task subject, and the user dimension data represents the business data corresponding to the dimension related to the user subject; Based on the business hierarchical relationship, at least one fact data set is obtained according to the task dimension data, the user dimension data and the fact data, wherein the business hierarchical relationship represents the hierarchical relationship between the task subject and the hierarchical relationship between the user subject; and Obtaining, according to the at least one fact data set, at least one statistical data set corresponding to each statistical indicator of the at least one statistical indicator; The task theme includes a plurality of hierarchical dimensions related to the task theme, and the user theme includes a plurality of hierarchical dimensions related to the user theme; The obtaining, based on the business hierarchical relationship and according to the task dimension data, the user dimension data and the fact data, at least one fact data set includes: Associating the multiple hierarchical dimensions related to the task theme with the multiple hierarchical dimensions related to the user theme to obtain multiple hierarchical associated dimensions; and Based on the multiple associated dimensions with a hierarchical relationship, the task dimension data, the user dimension data and the fact data are aggregated to obtain the at least one fact data set.

2. The method according to claim 1, further comprising: include: Data cleaning is performed on a plurality of original business data to obtain the plurality of business data.

3. The method according to claim 2, further comprising: include: Determine, based on a data structure specification of the data warehouse and a data structure of each of the multiple data sources, a storage strategy for the original business data corresponding to each of the data sources; as well as According to the storage strategy of the original business data corresponding to each data source, the original business data is stored in the data warehouse.

4. The method according to claim 3, in, The storing the original business data into the data warehouse according to the storage strategy of the original business data corresponding to each data source includes: According to the storage strategy of the original business data corresponding to each data source, determining a data import tool that matches the storage strategy; and The original business data is stored in the data warehouse using a data import tool that matches the storage strategy.

5. The method according to any one of claims 1 to 4, in, The step of acquiring task dimension data corresponding to the task subject, user dimension data corresponding to the user subject, and fact data corresponding to the fact from the plurality of business data includes: Determine, according to the feature association relationship, a feature corresponding to the task theme, a feature corresponding to the user theme, and a feature corresponding to the fact; According to the features corresponding to the task subject, acquiring task dimension data corresponding to the task subject from the plurality of business data; According to the features corresponding to the user theme, acquiring user dimension data corresponding to the user theme from the plurality of business data; and According to the feature corresponding to the fact, fact data corresponding to the fact is acquired from the plurality of business data.

6. The method according to claim 1, in, The step of obtaining, based on the at least one fact data set, at least one statistical data set corresponding to each statistical indicator in the at least one statistical indicator comprises: Determining the at least one statistical indicator according to the business requirement rule; and For each statistical indicator of the at least one statistical indicator, at least one fact data set corresponding to the statistical indicator is associated to obtain at least one statistical data set corresponding to the statistical indicator.

7. The method according to claim 1, further comprising: include: In response to a query request, according to a query condition included in the query request, a data query result matching the query condition is determined from at least one statistical data set corresponding to each statistical indicator of the at least one statistical indicator.

8. The method according to claim 7, in, The step of responding to the query request and determining, according to the query condition included in the query request, a data query result matching the query condition from at least one statistical data set corresponding to each statistical indicator of the at least one statistical indicator comprises: In response to the query request, calling a data warehouse interface; and By utilizing the data warehouse interface, according to the query condition included in the query request, a data query result matching the query condition is determined from at least one statistical data set corresponding to each statistical indicator of the at least one statistical indicator.

9. The method according to claim 7 or 8, further comprising: include: The data query result is processed to obtain a data processing result.

10. A data processing device, include: An acquisition module, used to acquire task dimension data corresponding to the task subject, user dimension data corresponding to the user subject, and fact data corresponding to the fact from multiple business data, wherein the task dimension data represents business data corresponding to the dimension related to the task subject, and the user dimension data represents business data corresponding to the dimension related to the user subject; A first obtaining module is configured to obtain at least one fact data set based on a business hierarchical relationship according to the task dimension data, the user dimension data and the fact data, wherein the business hierarchical relationship represents a hierarchical relationship between the task subject and a hierarchical relationship between the user subject; and A second obtaining module, configured to obtain, based on the at least one fact data set, at least one statistical data set corresponding to each statistical indicator of the at least one statistical indicator; The task theme includes a plurality of hierarchical dimensions related to the task theme, and the user theme includes a plurality of hierarchical dimensions related to the user theme; The first obtaining module comprises: a first obtaining unit, configured to associate the plurality of hierarchical dimensions related to the task theme with the plurality of hierarchical dimensions related to the user theme to obtain a plurality of hierarchical associated dimensions; and The second obtaining unit is used to perform aggregation processing on the task dimension data, the user dimension data and the fact data based on the multiple associated dimensions with a hierarchical relationship to obtain the at least one fact data set.

11. The device according to claim 10, further comprising: include: The third acquisition module is used to perform data cleaning on multiple original business data to obtain the multiple business data.

12. The device according to claim 11, further comprising: include: A first determination module, configured to determine a storage strategy for original business data corresponding to each data source based on a data structure specification of the data warehouse and a data structure of each of the multiple data sources; as well as The storage module is used to store the original business data in the data warehouse according to the storage strategy of the original business data corresponding to each data source.

13. The device according to claim 12, in, The storage module comprises: A first determining unit is used to determine a data import tool that matches the storage strategy of the original business data corresponding to each data source; and The storage unit is used to store the original business data into the data warehouse by using a data import tool that matches the storage strategy.

14. The device according to any one of claims 10 to 13, in, The acquisition module comprises: A second determining unit, configured to determine, according to the feature association relationship, a feature corresponding to the task theme, a feature corresponding to the user theme, and a feature corresponding to the fact; A first acquisition unit, configured to acquire task dimension data corresponding to the task subject from the plurality of business data according to a feature corresponding to the task subject; a second acquisition unit, configured to acquire user dimension data corresponding to the user theme from the plurality of business data according to features corresponding to the user theme; and The third acquisition unit is used to acquire fact data corresponding to the fact from the multiple business data according to the characteristics corresponding to the fact.

15. The device according to claim 10, in, The second obtaining module comprises: A third determining unit is configured to determine the at least one statistical indicator according to a business requirement rule; and The third obtaining unit is used to associate at least one fact data set corresponding to each statistical indicator of the at least one statistical indicator to obtain at least one statistical data set corresponding to the statistical indicator.

16. An electronic device, include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any one of claims 1 to 9.

17. A non-transitory computer-readable storage medium storing computer instructions, in, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 9.

18. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Data aggregation server for managing a multi-dimensional database and database management system having data aggregation server integrated therein

    US20020029207A1

  • Method and Apparatus for Declarative Data Warehouse Definition for Object-Relational Mapped Objects

    US20090271345A1